Data processing method and apparatus, electronic device, and medium
By employing a combination of soft and hard clustering methods in the internet product matrix, clustering results of user features are generated and migrated, solving the problems of large data volume and high resource consumption in user data migration, and achieving efficient data migration and performance improvement.
Patent Information
- Application Number
- CN202210993473.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-18
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-08-18
AI Technical Summary
In the matrix of internet products, user data migration faces challenges such as large data volume, high hardware resource consumption, and long migration time, especially when migrating a large number of user feature representations, which is difficult to do efficiently.
A method combining soft and hard clustering is adopted to generate soft and hard clustering results of user feature representations. Only these clustering results are transferred, rather than directly transferring user feature representations. The soft clustering model is used to generate probability distribution clustering results of users, and the hard clustering model is used to generate cluster partitioning results of users.
It reduces the amount of data migration, lowers hardware resource consumption and data migration time, while maintaining clustering effects and improving the performance of the target application, such as recommendation accuracy.
Smart Images

Figure CN115358313B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to the field of computers, and more specifically, to data processing methods, apparatus, electronic devices, and media. Background Technology
[0002] With the development of internet technology, the types of internet products have become very diverse, forming an internet product matrix. Examples include news products, music products, social media products, and film and television products. Users can share the same account when using various products within this matrix.
[0003] Meanwhile, with the development of deep learning technology, many internet products have adopted deep learning models. In some internet products (e.g., mobile applications), the number of users is enormous, and the amount of data associated with each user is also vast. Due to the large volume and diverse data types, migrating this massive amount of data between internet product matrices often presents numerous challenges. Summary of the Invention
[0004] Embodiments of this disclosure provide a data processing method, apparatus, electronic device, and computer-readable storage medium.
[0005] According to a first aspect of this disclosure, a data processing method is provided. The method includes obtaining a set of feature representations of multiple users of a first application. The method further includes determining a first clustering result and a second clustering result of the feature representations based on the set of feature representations. The method also includes sending the first clustering result and the second clustering result to a second application different from the first application.
[0006] According to a second aspect of this disclosure, a data processing apparatus is provided. The apparatus includes a feature representation acquisition module configured to acquire a set of feature representations of a plurality of users of a first application. The apparatus further includes a clustering result determination module configured to determine a first clustering result and a second clustering result of the feature representation set based on the feature representation set. The apparatus also includes a clustering result sending module configured to send the first clustering result and the second clustering result to a second application different from the first application.
[0007] According to a third aspect of this disclosure, an electronic device is provided. The electronic device includes a processor and a memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to the first aspect.
[0008] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the method according to the first aspect.
[0009] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 A schematic diagram of an example environment in which data processing methods according to certain embodiments of the present disclosure can be implemented is shown;
[0012] Figure 2 A flowchart of a data processing method according to certain embodiments of the present disclosure is shown;
[0013] Figure 3 A flowchart illustrating a process for determining soft clustering results according to certain embodiments of the present disclosure is shown;
[0014] Figure 4 A schematic diagram comparing soft clustering results using different parameters according to certain embodiments of the present disclosure is shown;
[0015] Figure 5A A flowchart illustrating a process for determining hard clustering results according to certain embodiments of the present disclosure is shown;
[0016] Figure 5B A schematic diagram showing the start of a process 500 for determining hard clustering results according to certain embodiments of the present disclosure is shown;
[0017] Figure 6 A schematic diagram showing a comparison of hard clustering results according to certain embodiments of the present disclosure is illustrated;
[0018] Figure 7 A block diagram of a data processing apparatus according to certain embodiments of the present disclosure is shown; and
[0019] Figure 8 A block diagram of an apparatus for data processing according to certain embodiments of the present disclosure is shown.
[0020] In all the accompanying figures, the same or similar reference numerals denote the same or similar elements. Detailed Implementation
[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0022] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0023] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in words. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0025] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0026] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0027] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0028] Furthermore, all specific values in this article are examples only, intended to aid understanding, and are not intended to limit any range.
[0029] In deep learning systems, user features can be mapped to word embeddings, hence the term "feature representation." There is a one-to-one correspondence between a user's feature and its feature representation. The distance between different feature representations indicates the similarity between the features of different users. Therefore, user data migration in an internet product matrix can essentially be understood as migrating these feature representations.
[0030] It's easy to understand that, given the sheer volume of users in reality, directly migrating these massive feature representations might be impractical. When migrating user data across different applications, to reduce the amount of data being migrated and the overhead on hardware resources, only the clustering results of the user data can be migrated, without migrating the user data itself. This way, the aforementioned objective can be achieved without impacting application performance.
[0031] The inventors noticed that when clustering user data, grouping each user data point into only one cluster (i.e., ensuring that a user data point belongs entirely to one cluster, also known as hard clustering) can be overly absolute. This is because certain characteristics of the data may cause it to belong to multiple clusters. Therefore, a user data point can be divided into multiple clusters, and the probability of the user data point belonging to each cluster can be represented by a probability distribution (also known as soft clustering).
[0032] To address the aforementioned drawbacks, embodiments of this disclosure provide a data processing scheme. This scheme acquires word vectors, i.e., feature representation sets, from user data in an application. The feature representation sets are then clustered, yielding the two clustering results described above. During data migration, only these two clustering results need to be sent to other applications, without directly sending the user's feature representations. Upon receiving these clustering results, other applications can use them to train their own deep learning systems to achieve better performance in areas such as recommendation systems. It can be seen that the method proposed in this disclosure reduces the amount of data being migrated, thereby reducing hardware resource consumption and the time required for data migration.
[0033] In the following description, some embodiments will be discussed with reference to the migration process of user data in video products. However, it should be understood that this is merely to enable those skilled in the art to better understand the principles and ideas of the embodiments of this disclosure, and is not intended to limit the scope of this disclosure in any way.
[0034] Figure 1 A schematic diagram of an example environment 100 in which data processing methods according to certain embodiments of the present disclosure can be implemented is shown. Figure 1As shown, example environment 100 may include video application 110 (also referred to as the first application) and news application 120 (also referred to as the second application). Video application 110 and news application belong to the same internet matrix, and users may have cross-application needs. For example, browsing content of interest in the video application might lead to searching for keywords in the news application to obtain more information, and so on.
[0035] Example environment 100 may also include electronic device 180, which may be a computer, computing system, single server, distributed server, or cloud-based server. Electronic device 180 may obtain user feature representation set 130 from news application 110, obtain user feature representation set 130 from the server of news application 110, or the user feature representation set 130 may be stored in electronic device 180 itself.
[0036] In the electronic device 180, two clustering models can be configured: a soft clustering model 140 (also referred to as the first clustering model) and a hard clustering model 150 (also referred to as the second clustering model). Based on the acquired feature representation set, the soft clustering model 140 can generate a soft clustering result 160 (also referred to as the first clustering result). The hard clustering model 150 can generate a hard clustering result 170 (also referred to as the second clustering result).
[0037] The soft clustering results 160 and hard clustering results 170 can be sent to the news application 120 (also known as the second application). The news application 120 can train its own deep learning model based on the soft clustering results 160 and hard clustering results 170, thereby achieving better performance, such as better recommendation accuracy.
[0038] It should be understood that the architecture and functionality in example environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure. Embodiments of this disclosure can also be applied to other environments with different structures and / or functionalities.
[0039] The following will combine Figures 2 to 6 The process according to embodiments of this disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and not intended to limit the scope of this disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or actions shown may be omitted, and the scope of this disclosure is not limited in this respect.
[0040] Figure 2A flowchart of a data processing method 200 according to certain embodiments of the present disclosure is shown. At block 202, a set of feature representations of multiple users of a first application is obtained. For example, for video application 110, a certain number of users may be selected, each of whom has a feature representation, and these feature representations may form a set, referred to as a feature representation set. The number of users selected may be determined based on needs or the number that hardware resources can process at one time.
[0041] At box 204, based on the feature representation set, a first clustering result and a second clustering result for the feature representation set are determined. As an example, feature representation set 130 is processed by soft clustering model 140 to generate soft clustering result 160. As an example, the soft clustering result can consist of vectors indicating the probability distribution of a user's feature representation. Suppose soft clustering model 140 can divide users into 5 clusters, then the probability of user A belonging to cluster 1 is 0.1, to cluster 2 is 0.1, to cluster 3 is 0.5, to cluster 4 is 0.2, and to cluster 5 is 0.1. This clustering result for user A can then be represented as the vector [0.1 0.1 0.5 0.2 0.1]. Similarly, soft clustering result 160 can consist of such vectors representing a selected number of users.
[0042] The feature representation set 130 is further processed by the hard clustering model 150 to generate a hard clustering result 170. As an example, similarly, the hard clustering result 170 consists of vectors indicating which cluster a user's feature representation belongs to. Suppose the hard clustering model can divide users into 6 clusters, and user A belongs to cluster 6, then this clustering result for user A can be represented as the vector [0 0 00 0 1]. Similarly, the hard clustering result 170 can consist of such vectors for a selected number of users.
[0043] It's understandable that the number of clusters in hard and soft clustering needs to satisfy the following condition: a large number of users are divided into different classes or clusters according to a specific criterion, maximizing the similarity of users within the same cluster and maximizing the differences between users in different clusters. In other words, similar users should be grouped together as much as possible according to a specific criterion, while users of different classes should be separated as much as possible. Of course, in actual clustering, the objects being clustered are feature representations.
[0044] Therefore, choosing the right parameters is crucial for both soft and hard clustering. This will be discussed further in the following description. Figures 3 to 6 This section describes how to implement soft and hard clustering, and how to choose clustering parameters. For the sake of brevity, this section will not go into detail here.
[0045] At box 206, the first and second clustering results are sent to a second application, which is different from the first application. As an example, soft clustering result 160 and hard clustering result 170 are sent to news application 120.
[0046] In this way, by avoiding the direct migration of feature representations from a large number of users, and instead clustering specific representation sets before migrating the clustering results, the amount of data that needs to be migrated is effectively reduced. Simultaneously, during clustering, two methods are utilized to effectively group feature representations with similar characteristics into a single cluster, while ensuring sufficient differences between different clusters. Therefore, this disclosure reduces the amount of data to be migrated while maintaining sufficiently good clustering results. This reduces hardware resource consumption and data migration time while meeting the performance requirements of the second application.
[0047] Figure 3 A flowchart of a process 300 for determining soft clustering results according to certain embodiments of the present disclosure is shown. Figure 2 The soft clustering results mentioned above can be determined in Procedure 300. Traditional clustering methods (e.g., K-means) cannot cluster two samples with the same mean (same cluster centroid), and Procedure 300 proposed in this paper addresses this shortcoming. Procedure 300 performs clustering by selecting "components" to maximize posterior probabilities. The posterior probability of each sample represents its likelihood of belonging to various clusters, rather than determining whether it belongs to a particular cluster. Procedure 300 may be more suitable than k-means clustering when the sizes of individual samples and clusters differ and there are correlations between clusters.
[0048] The main idea of Process 300 is to approximate an arbitrary probability distribution with multiple pre-defined probability density distributions (e.g., Gaussian distribution functions). Therefore, Process 300 consists of multiple individual probability density distributions, each called a "component." The weighted sum of these "components" constitutes the probability density function of Process 300. The samples to be clustered are considered as sampling points of the distribution. The parameters of the probability distribution are estimated using the expectation-maximization (EM) algorithm (e.g., maximum likelihood estimation) based on these sampling points. Finding the parameters yields the probability distribution of the data points for classification. To illustrate Process 300, a Gaussian probability density function is used as an example below.
[0049] The probability distribution of process 300 can be found in formula (1):
[0050]
[0051] Where x represents the sample; k represents the number of clusters (i.e., the number of clusters); α iLet represent the probability of the k-th component (also called the prior distribution), which must be greater than zero, and for a sample x, k α values... i The sum equals 1; p(x|μ) i ,∑ i ) represents the probability density function of an n-dimensional random vector that follows a Gaussian distribution, and can be found in formula (2).
[0052]
[0053] Where n represents the number of feature representations to be clustered; μ represents the mean; ∑ represents the covariance matrix; and T represents the transpose operator.
[0054] In some embodiments, process 300 may be the following process:
[0055] At box 302, set the value of k, which is the number of clusters in initialization process 300. Randomly initialize the Gaussian distribution parameters (mean and covariance) for each cluster. Alternatively, observe the data to obtain a relatively accurate mean and covariance.
[0056] At box 304, calculate the probability that each sample belongs to each Gaussian model, i.e., calculate the posterior probability. The closer a sample is to the center of the Gaussian distribution, the higher its probability, meaning it is more likely to belong to that cluster.
[0057] At box 306, calculate the values of α, μ, and ∑ to maximize the probability of the sample. Use a weighted sum of the sample probabilities to calculate these new parameters, where the weights are the probabilities that the sample belongs to the cluster.
[0058] At box 308, repeat the iterations for boxes 304 and 306 until convergence.
[0059] The process 300 is mainly determined by two parameters: covariance and mean. Different learning mechanisms for learning covariance and mean will directly affect the stability, accuracy and convergence of the model.
[0060] The covariance type used when estimating parameters using the EM algorithm, such as 'full', 'tied', 'diag', 'spherical'.
[0061] `full` represents the complete covariance matrix (all elements are non-zero). Each cluster can have any independent location and shape.
[0062] "Tie" means that each cluster shares the same complete covariance. Each cluster has the same shape, but the shape can be any shape.
[0063] diag represents the diagonal covariance matrix (off-diagonal values are zero, diagonal values are not zero). Each cluster profile axis is along the coordinate axis.
[0064] spherical represents the spherical covariance matrix (zero off-diagonal and identical diagonal). Each cluster has a circular outline.
[0065] Figure 4 A schematic diagram comparing soft clustering results using different parameters according to certain embodiments of this disclosure is shown. It can be seen that process 300 should select parameters to adapt to the transfer of feature representations. As an example, the number of clusters can be selected from 16, 32, 64, and 128 depending on the number of users. When transferring clustering results, process 300 selects 'tied' as the covariance matrix type (each cluster shares the same complete covariance matrix. Each cluster has the same shape, but the shape can be any shape), where the clustering is the most non-uniform, and such results are the most discriminative. When selecting parameters, the selected parameters are also adjusted to adapt to different hardware resources, such as memory space.
[0066] Process 300 avoids the problem of a user belonging 100% to only one cluster, meaning the intersection of any two clusters is an empty set. This type of clustering algorithm makes user affiliation too absolute. In contrast, Process 300 allows a sample to be divided into multiple clusters, representing user affiliation as a probability distribution. This avoids overly absolute divisions, allowing user feature representations to be differentiated to a certain extent while preserving more individual characteristics and correlations between features.
[0067] Figure 5A A flowchart of a process 500 for determining hard clustering results according to certain embodiments of the present disclosure is shown. Figure 5B A schematic diagram is shown illustrating the start of a process 500 for determining hard clustering results according to certain embodiments of the present disclosure. Figure 2 The hard clustering results mentioned can be determined in procedure 500.
[0068] The main idea of Process 500 is that the set of samples with the highest density connectivity derived from density reachability relations constitutes a category, or cluster, in the final clustering. A cluster can contain one or more core objects. If there is only one core object, all other non-core object samples in the cluster are within the ∈-neighborhood of this core object. If there are multiple core objects, then any core object in the ∈-neighborhood of any given core object must contain another core object; otherwise, these two core objects are not density-reachable. The set of all samples within the ∈-neighborhoods of these core objects forms a cluster.
[0069] Specifically, Process 500 can arbitrarily select a core object without a category as a seed, and then find the set of all samples that are density-reachable from this core object, which constitutes a cluster. Then, it continues to select another core object without a category and find the set of samples that are density-reachable, thus obtaining another cluster. This process continues until all core objects have a category. That is, how to calculate the distance between a sample and samples from core objects. In Process 500, the nearest neighbor concept is generally used, employing a distance metric to measure the sample distance, such as Euclidean distance, Manhattan distance, etc.
[0070] In some embodiments, reference Figure 5B Process 500 can begin at sample set 520 and proceed as follows:
[0071] At box 502, obtain the sample set D = {x1, x2, ..., x}. m}, neighborhood parameters (eps, min_sampels).
[0072] At box 504, initialize the core object collection, where... Initialize category k = 0.
[0073] At box 506, iterate through the elements of D and determine whether an element in the sample set is a core object.
[0074] At box 508, if it is a core object, add it to the core object collection Ω. If all elements in the core object collection Ω have been accessed, the process ends; otherwise, proceed to box 510.
[0075] At box 510, in the core object set Ω, randomly select an unvisited core object o, first mark o as visited, then mark o as category k, and finally store the unvisited data in o's ∈-neighborhood into the seed set Seeds.
[0076] At box 512, if the seed set Then the current cluster C k Once generation is complete, set k = k + 1 and proceed to box 506. Otherwise, select a seed point from the seed set Seeds, first mark it as visited and label it with category k, then determine if the seed is a core object. If it is, add the unvisited seed points in the seed to the seed set and proceed to box 510.
[0077] It's important to note that eps represents the distance threshold between the ∈-neighborhood. If eps is too large, more points will fall into the ∈-neighborhood of the core object, potentially reducing the number of clusters and causing samples that shouldn't be in the same class to be grouped together. Conversely, if eps is too small, the number of clusters may increase, separating samples that should be in the same class.
[0078] Similarly, this refers to the threshold number of samples in the ∈-neighborhood required for a sample point to become a core object. It's usually tuned in conjunction with eps. With a fixed eps, if min_samples is too large, there will be too few core objects. In this case, samples within a cluster that are actually of the same class might be labeled as noise points, and the number of clusters will increase. Conversely, if min_samples is too small, a large number of core objects will be generated, potentially leading to too few clusters.
[0079] Figure 6 A schematic diagram comparing hard clustering results 600 according to certain embodiments of the present disclosure is shown. It can be seen that process 500 achieves better clustering results, enabling the clustering of dense data of arbitrary shapes, such as the annular dataset 602 and the concave dataset 604. However, other clustering methods cannot handle annular datasets well, instead dividing them into semicircles. For example, the annular dataset 606 is divided into semicircles, and the concave dataset 608 is divided into upper and lower parts, thus losing the data's characteristics. Similarly, traditional clustering cannot handle concave datasets.
[0080] It's understandable that the choice between parameters eps and min_samples is essentially a trade-off. If the current number of clusters is too small, then eps needs to be reduced or min_samples increased to enhance the discriminative power of the current clustering. As an example, in real-world applications with hundreds of millions of users, eps between [0.1, 0.5] and min_samples between [5, 10] can achieve good clustering results. It's worth noting that the number of clusters in process 300 and process 500 may not be the same; the specific number needs to be selected by the algorithm or predetermined when executing the process.
[0081] Therefore, the hard clustering method of Process 500 can handle datasets of various shapes, thus achieving one-to-one hard clustering with better discriminative power, reducing the loss of user characteristics due to clustering errors, and providing more reliable migration data when migrating data.
[0082] In some embodiments, both soft clustering results 160 and hard clustering results 170 can be sent to a target application, such as a news application 120. Upon receiving these clustering results, the news application 120 or its server can use them to train its own deep learning model to improve performance. As an example, the received clustering results can be used as additional sample data for training the deep learning model, combined with labeled user data and recommended data based on sample labels, to train the deep learning model of the news application 120. This can result in better performance during cold starts, such as better recommendation accuracy for new users.
[0083] comprehensive Figures 2 to 6 As can be seen from the description, this disclosure can reduce the amount of data to be migrated, thereby reducing hardware resource consumption and the time required for data migration. Simultaneously, during clustering, two methods are utilized to effectively group feature representations with similar characteristics into a single cluster, while ensuring sufficient difference between different clusters. Because the two clustering methods have different characteristics, sending both clustering results to other applications can achieve complementarity, maximizing the migration of user features. After receiving these clustering results, other applications can use them to train their own deep learning systems to achieve better performance in areas such as recommendation systems.
[0084] Figure 7 A block diagram of an apparatus 700 for data processing according to certain embodiments of the present disclosure is shown. Figure 7 As shown, the device 700 includes a feature representation acquisition module 702, configured to acquire a set of feature representations of multiple users of a first application. The device 700 also includes a clustering result determination module 704, configured to determine a first clustering result and a second clustering result of the feature representation set based on the feature representation set. The device 700 further includes a clustering result sending module 706, configured to send the first clustering result and the second clustering result to a second application different from the first application.
[0085] In some embodiments, the clustering result determination module 704 may also be configured to implement Figure 3 and / or the process described in Figure 5. These can be referenced. Figure 3 The description in Figure 5 and / or Figure 5 can be used to understand this, and will not be repeated here.
[0086] The apparatus 700 of this disclosure can achieve at least one of the many advantages achievable by the methods or processes described above. For example, it reduces the amount of data to be migrated while maintaining sufficiently good clustering results. This reduces hardware resource overhead and the time required for data migration while improving the performance of the second application.
[0087] Figure 8A block diagram of a device 800 for user data processing according to certain embodiments of the present disclosure is shown. Device 800 may be the device or apparatus described in the embodiments of the present disclosure. Figure 8 As shown, device 800 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 802 or loaded from storage unit 808 into random access memory (RAM) 803. The RAM 803 can also store various programs and data required for the operation of device 800. The CPU / GPU 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804. Although not shown in... Figure 8 As shown, device 800 may also include a coprocessor.
[0088] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0089] The various methods or processes described above can be executed by CPU / GPU 801. For example, in some embodiments, the methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by CPU / GPU 801, one or more steps or actions in the methods or processes described above can be performed.
[0090] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.
[0091] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0092] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network, to an external computer or external storage device. The network may include copper cables, fiber optic cables, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0093] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0094] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0095] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0097] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0098] The following are some example implementations of this disclosure.
[0099] Example 1. A data processing method, comprising:
[0100] Obtain a set of feature representations of multiple users of the first application;
[0101] Based on the feature representation set, determine the first clustering result and the second clustering result of the feature representation set; and
[0102] The first clustering result and the second clustering result are sent to a second application that is different from the first application.
[0103] Example 2. According to the method described in Example 1, determining the first clustering result includes:
[0104] For each of the plurality of users, a probability distribution of the feature representation is determined, wherein the probability distribution indicates the probability that the feature representation belongs to a predetermined plurality of clusters; and
[0105] The first clustering result is determined based on the probability distribution of each of the multiple users.
[0106] Example 3. The method according to any one of Examples 1-2, wherein determining the probability distribution of the feature representation comprises:
[0107] Obtain multiple probability density distributions for the feature representation;
[0108] Based on the plurality of probability density distributions, determine the weighted sum of the plurality of probability density distributions; and
[0109] Based on the weighted sum, the probability distribution of the feature representation is determined.
[0110] Example 4. The method according to any one of Examples 1-3, wherein determining the plurality of probability density distributions comprises:
[0111] Determine the shape and number of clusters; and
[0112] Based on the determined shape and number of the clusters, the plurality of probability density distributions are determined.
[0113] Example 5. The method according to any one of Examples 1-4, wherein determining the shape of the cluster includes:
[0114] The shape of the clusters is adjusted by determining different covariance matrices.
[0115] Example 6. The method according to any one of Examples 1-5, wherein determining the second clustering result includes:
[0116] A certain number of feature representations are determined from the set of feature representations to serve as cluster centers;
[0117] For each cluster center:
[0118] Search for feature representations within a threshold distance of the cluster centers; and
[0119] The cluster represented by the cluster center includes the searched feature representation;
[0120] Based on the clustering, the second clustering result is determined.
[0121] Example 7. The method according to any one of Examples 1-6, wherein searching for feature representations within a threshold distance of the cluster centers comprises:
[0122] In response to the number of feature representations searched within a threshold distance of the cluster center exceeding a threshold number, feature representations are searched in the region outside the threshold distance.
[0123] Example 8. The method according to any one of Examples 1-7, wherein the cluster represented by the cluster center includes the searched feature representation comprising:
[0124] In response to the number of searched feature representations being greater than a threshold number, the searched feature representations are included in the cluster represented by the cluster center.
[0125] Example 9. The method according to any one of Examples 1-8 further includes:
[0126] Adjust the threshold distance and the threshold number; and
[0127] The number of clusters and the number of feature representations included in each cluster are optimized using the adjusted threshold distance and the threshold number.
[0128] Example 10. According to any one of Examples 1-9, in response to the second application receiving the first clustering result and the second clustering result, a deep learning model of the second application is trained based on the first clustering result and the second clustering result.
[0129] Example 11. A data processing apparatus, comprising:
[0130] The feature representation acquisition module is configured to acquire a set of feature representations for multiple users of the first application;
[0131] The clustering result determination module is configured to determine a first clustering result and a second clustering result of the feature representation set based on the feature representation set; and
[0132] The clustering result sending module is configured to send the first clustering result and the second clustering result to a second application that is different from the first application.
[0133] Example 12. The apparatus according to Example 11, wherein the clustering result determination module comprises:
[0134] A first probability module is configured to determine a probability distribution of the feature representations for each of the plurality of users, wherein the probability distribution indicates the probability that the feature representation belongs to a predetermined plurality of clusters; and
[0135] The first clustering module is configured to determine the first clustering result based on the probability distribution of each of the multiple users.
[0136] Example 13. The apparatus according to any one of Examples 11-12, wherein the first probability module comprises:
[0137] The first probability density distribution module is configured to acquire multiple probability density distributions for the feature representation; and
[0138] A weighted sum module is configured to determine a weighted sum of the plurality of probability density distributions based on the plurality of probability density distributions; and
[0139] The second probability distribution module is configured to determine the probability distribution of the feature representation based on the weighted sum.
[0140] Example 14. The apparatus according to any one of Examples 11-13, wherein the first probability density distribution module comprises:
[0141] The clustering parameter module is configured to determine the shape and number of clusters; and
[0142] The third probability distribution module is configured to determine the plurality of probability density distributions based on the determined shape and number of the clusters.
[0143] Example 15. The apparatus according to any one of Examples 11-14, wherein the clustering parameter module comprises:
[0144] The covariance adjustment module is configured to adjust the shape of the clusters by determining different covariance matrices.
[0145] Example 16. The apparatus according to any one of Examples 11-15, wherein the clustering result determination module further comprises:
[0146] The cluster center module is configured to determine a certain number of feature representations in the set of feature representations as cluster centers;
[0147] A clustering search module is configured to, for each cluster center, search for a feature representation within a threshold distance of the cluster center; and include the searched feature representation in the cluster represented by the cluster center; and
[0148] The second clustering module is configured to determine the second clustering result based on the clustering.
[0149] Example 17. The apparatus according to any one of Examples 11-16, wherein the clustering search module comprises:
[0150] The second clustering search module is configured to search for feature representations in regions outside the threshold distance in response to the number of feature representations found within a threshold distance of the cluster center exceeding a threshold number.
[0151] Example 18. The apparatus according to any one of Examples 11-17, wherein the second clustering search module comprises:
[0152] The third clustering search module is configured to include the searched feature representations in the cluster represented by the cluster center in response to the number of searched feature representations being greater than a threshold number.
[0153] Example 19. The apparatus according to any one of Examples 11-18 further includes a parameter adjustment module configured to:
[0154] Adjust the threshold distance and the threshold number; and
[0155] The number of clusters and the number of feature representations included in each cluster are optimized using the adjusted threshold distance and the threshold number.
[0156] Example 20. The apparatus according to any one of Examples 11-19 further includes a training module configured to:
[0157] In response to the second application receiving the first clustering result and the second clustering result, the second application trains a deep learning model based on the first clustering result and the second clustering result.
[0158] Example 21. An electronic device comprising:
[0159] Processor; and
[0160] A memory coupled to the processor, the memory having instructions stored therein, the instructions which, when executed by the processor, cause the electronic device to perform actions, the actions including:
[0161] Obtain a set of feature representations of multiple users of the first application;
[0162] Based on the feature representation set, determine the first clustering result and the second clustering result of the feature representation set; and
[0163] The first clustering result and the second clustering result are sent to a second application that is different from the first application.
[0164] Example 22. The electronic device according to Example 21, wherein determining the first clustering result includes:
[0165] For each of the plurality of users, a probability distribution of the feature representation is determined, wherein the probability distribution indicates the probability that the feature representation belongs to a predetermined plurality of clusters; and
[0166] The first clustering result is determined based on the probability distribution of each of the multiple users.
[0167] Example 23. An electronic device according to any one of Examples 21-22, wherein determining the probability distribution of the feature representation comprises:
[0168] Obtain multiple probability density distributions for the feature representation;
[0169] Based on the plurality of probability density distributions, determine the weighted sum of the plurality of probability density distributions; and
[0170] Based on the weighted sum, the probability distribution of the feature representation is determined.
[0171] Example 24. An electronic device according to any one of Examples 21-23, wherein determining the plurality of probability density distributions includes:
[0172] Determine the shape and number of clusters; and
[0173] Based on the determined shape and number of the clusters, the plurality of probability density distributions are determined.
[0174] Example 25. An electronic device according to any one of Examples 21-24, wherein determining the shape of the cluster includes:
[0175] The shape of the clusters is adjusted by determining different covariance matrices.
[0176] Example 26. An electronic device according to any one of Examples 21-25, wherein determining the second clustering result includes:
[0177] A certain number of feature representations are determined from the set of feature representations to serve as cluster centers;
[0178] For each cluster center:
[0179] Search for feature representations within a threshold distance of the cluster centers; and
[0180] The cluster represented by the cluster center includes the searched feature representation;
[0181] Based on the clustering, the second clustering result is determined.
[0182] Example 27. An electronic device according to any one of Examples 21-26, wherein the search feature representation within a threshold distance of the cluster centers includes:
[0183] In response to the number of feature representations searched within a threshold distance of the cluster center exceeding a threshold number, feature representations are searched in the region outside the threshold distance.
[0184] Example 28. An electronic device according to any one of Examples 21-27, wherein the cluster represented by the cluster center includes the searched feature representation comprising:
[0185] In response to the number of searched feature representations being greater than a threshold number, the searched feature representations are included in the cluster represented by the cluster center.
[0186] Example 29. The electronic device according to any one of Examples 21-28, wherein the operation further includes:
[0187] Adjust the threshold distance and the threshold number; and
[0188] The number of clusters and the number of feature representations included in each cluster are optimized using the adjusted threshold distance and the threshold number.
[0189] Example 30. The electronic device according to any one of Examples 21-29, wherein the operation further includes:
[0190] In response to the second application receiving the first clustering result and the second clustering result, the second application trains a deep learning model based on the first clustering result and the second clustering result.
[0191] Example 31. A computer-readable storage medium having stored thereon one or more computer instructions, wherein the one or more computer instructions are executed by a processor to implement the method according to any one of Examples 1 to 10.
[0192] Example 32. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of Examples 1 to 10.
[0193] Although this disclosure has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A data processing method comprising: obtaining a set of feature representations of a plurality of users of a first application; determining a first clustering result and a second clustering result of the set of feature representations based on the set of feature representations, and wherein determining the first clustering result comprises: determining, for each feature representation of the plurality of users, a probability distribution of the feature representation, wherein the probability distribution indicates respective probabilities of the feature representation belonging to a predetermined plurality of clusters; determining the first clustering result based on the probability distributions of the plurality of users respectively; and wherein determining the second clustering result comprises: determining a number of feature representations in the set of feature representations as cluster centers; for each cluster center, searching for feature representations within a threshold distance of the cluster center; including the searched feature representations in a cluster represented by the cluster center; determining the second clustering result based on the clusters; and sending the first clustering result and the second clustering result to a second application different from the first application.
2. The method of claim 1, wherein determining the probability distribution of a feature representation comprises: obtaining a plurality of probability density distributions for the feature representation; determining a weighted sum of the plurality of probability density distributions based on the plurality of probability density distributions; and determining the probability distribution of the feature representation based on the weighted sum.
3. The method of claim 2, wherein determining the plurality of probability density distributions comprises: determining a shape of clusters and a number of clusters; and determining the plurality of probability density distributions based on the determined shape of clusters and the number of clusters.
4. The method of claim 3, wherein determining the shape of clusters comprises: adjusting the shape of clusters by determining different covariance matrices.
5. The method of claim 1, wherein searching for feature representations within a threshold distance of a cluster center comprises: in response to a number of searched feature representations within the threshold distance of the cluster center exceeding a threshold number, searching for feature representations in a region outside the threshold distance.
6. The method of claim 1, wherein including searched feature representations in a cluster represented by a cluster center comprises: in response to a number of the searched feature representations being greater than a threshold number, including the searched feature representations in the cluster represented by the cluster center.
7. The method of claim 6, further comprising: adjusting the threshold distance and the threshold number; and optimizing a number of clusters and a number of feature representations included in each cluster with the adjusted threshold distance and the threshold number.
8. The method of claim 1, further comprising: in response to the second application receiving the first clustering result and the second clustering result, training a deep learning model of the second application based on the first clustering result and the second clustering result.
9. A data processing apparatus comprising: a feature representation obtaining module configured to obtain a set of feature representations of a plurality of users of a first application; a clustering result determination module configured to determine, based on the set of feature representations, a first clustering result and a second clustering result of the set of feature representations, and wherein determining the first clustering result comprises: determining, for each user in the plurality of users, a probability distribution of the feature representation of the user, wherein the probability distribution indicates respective probabilities of the feature representation belonging to a predetermined plurality of clusters; determining the first clustering result based on the probability distributions of the plurality of users; and wherein determining the second clustering result comprises: determining, among the set of feature representations, a number of feature representations as cluster centers; for each cluster center, searching for feature representations within a threshold distance of the cluster center; including the searched feature representations in a cluster represented by the cluster center; determining the second clustering result based on the clusters; and a clustering result sending module configured to send the first clustering result and the second clustering result to a second application different from the first application.
10. An electronic device, comprising: a processor; and a memory coupled with the processor, the memory having stored therein instructions which, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1 to 8.
11. A computer readable storage medium having stored thereon computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
User information processing method and device, APP servers, and terminal equipment
CN107959757A
Abnormal session text detection method and device
CN113127639A
Model training method and device
CN113850384A
Method and device for clustering data, electronic equipment and storage medium
CN114548276A