Data Processing Method, Apparatus, Computer Storage Medium, and Electronic Device

By performing multi-dimensional processing of social e-commerce users' sharing data and reducing feature data dimensionality, combined with the matching results of the algorithm model, accurate prediction of user sharing behavior is achieved, and the problem of inability to accurately analyze user sharing intentions and effects in the existing technology is solved, and the accuracy and efficiency of operations are improved.

CN111325565BActive Publication Date: 2025-05-27BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201811531160.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-12-14
Publication Date
2025-05-27
Estimated Expiration
2038-12-14

AI Technical Summary

Technical Problem

When evaluating the communication situation and communication effect of users, existing social e-commerce requires a lot of manpower and time, and cannot accurately analyze the willingness of users and the effect of sharing, resulting in the inability to achieve precise operations, especially the inability to cover all users.

Method used

The user shared data is obtained through the multi-dimensional raw data, and processed using the first algorithm model to obtain the first target data and the target dimension number. At the same time, multi-dimensional feature data is obtained and dimensionality reduction is performed through the second algorithm model to obtain the first feature value with the target dimension number. Then, the first target data is matched with the second target data, and the user's sharing behavior is predicted based on the matching result.

Benefits of technology

It realizes automatic data analysis, saves manpower and statistical time, reduces costs, can accurately analyze data in all dimensions, improves the accuracy of analysis results, and provides accurate guidance on product operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111325565B_ABST
    Figure CN111325565B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing method, apparatus, computer-readable medium, and electronic device. The method includes: obtaining user sharing data according to multi-dimensional original data, and processing the user sharing data through a first algorithm model to obtain first target data and a target dimension number, where the quantity of the first target data is the same as the target dimension number; obtaining multi-dimensional feature data, and performing dimensionality reduction processing on the feature data through a second algorithm model to obtain first eigenvalues with the target dimension number; obtaining second target data according to the first eigenvalues, where the type of the second target data is the same as that of the first target data; matching the first target data with the second target data, and predicting the sharing behavior of the user according to the matching result. The present disclosure can save a large amount of manpower and statistical time; and can predict the sharing situation of the user according to the matching result, so as to perform precise operation.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the development of science and technology, the traditional commodity trading mode has been gradually replaced by the e-commerce model. E-commerce is a business activity centered on commodity exchange using information network technology, which realizes online shopping for consumers, online transactions between merchants, online electronic payments, and various business activities, trading activities, financial activities, and related comprehensive service activities. Social e-commerce is a type of e-commerce, which is based on interpersonal networks and uses Internet social tools to engage in the sale of goods or services.

[0003] The most important indicator for operating social e-commerce is to evaluate the sharing and dissemination of users and the dissemination effect. At present, social e-commerce is to collect statistics on e-commerce data, and then the operators manually analyze the statistical results and output statistical reports. However, the existing statistical methods require a lot of manpower and time, and rely on the prediction of product operations. Therefore, it is impossible to accurately analyze the user's willingness to share and the effect after sharing, that is, it is impossible to achieve accurate operation. In addition, the statistical method is only effective for old users and cannot cover all users.

[0004] In view of this, there is an urgent need in the art to develop a new data processing method and device.

[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0006] The purpose of the present disclosure is to provide a data processing method, a data processing device, a computer-readable storage medium and an electronic device, thereby saving manpower and statistical time at least to a certain extent, realizing targeted and accurate operations, and being able to predict the sharing intention and sharing effect of all types of users.

[0007] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.

[0008] According to one aspect of the present disclosure, a data processing method is provided, the method comprising:

[0009] Acquire user sharing data according to the multi-dimensional original data, and process the user sharing data through a first algorithm model to acquire first target data and a target number of dimensions, wherein the number of the first target data is the same as the target number of dimensions;

[0010] Acquire multi-dimensional feature data, and perform dimensionality reduction processing on the feature data through a second algorithm model to obtain a first feature value having the target number of dimensions, wherein the feature data is data in the original data that has a dominant effect on the user's sharing behavior;

[0011] Acquire second target data according to the first characteristic value, where the second target data is of the same type as the first target data;

[0012] The first target data is matched with the second target data, and the user's sharing behavior is predicted according to the matching result.

[0013] In an exemplary embodiment of the present disclosure, the first target data includes a first expectation and a first variance;

[0014] Acquiring user sharing data according to multi-dimensional original data, and clustering the user sharing data through a first algorithm model to acquire first target data and a target number of dimensions, including:

[0015] Clustering the user sharing data by using the first algorithm model to obtain multiple groups of user sharing sub-data that conform to Gaussian distribution;

[0016] Iteratively using a maximum expectation algorithm to obtain a first expectation and a first variance corresponding to each group of user-shared sub-data;

[0017] The target number of dimensions is determined according to the first variance.

[0018] In an exemplary embodiment of the present disclosure, determining the target number of dimensions according to the first variance includes:

[0019] Determine a standard deviation corresponding to each group of the user sharing sub-data according to the first variance;

[0020] Sum the standard deviations corresponding to the first N groups of user-shared sub-data to obtain a first standard deviation sum, and sum the standard deviations corresponding to the first N+1 groups of user-shared sub-data to obtain a second standard deviation sum, where N>1 and is a positive integer;

[0021] The sum of the first standard deviations is compared with the sum of the second standard deviations, and the target number of dimensions is determined according to the comparison result.

[0022] In an exemplary embodiment of the present disclosure, comparing the sum of the first standard deviations with the sum of the second standard deviations, and determining the target number of dimensions according to the comparison result, includes:

[0023] Subtracting the sum of the first standard deviations from the sum of the second standard deviations to obtain an increment of the sum of the first standard deviations relative to the sum of the second standard deviations;

[0024] Comparing the increment with a preset threshold to determine whether the increment is less than the preset threshold;

[0025] If the increment is smaller than the preset threshold, the number of groups of user-shared sub-data corresponding to the sum of the first standard deviation is the target number of dimensions.

[0026] In an exemplary embodiment of the present disclosure, multi-dimensional feature data is acquired, and dimension reduction processing is performed on the feature data through a second algorithm model to obtain a first feature value having the target number of dimensions, including:

[0027] Extracting the feature data from the original data according to preset conditions;

[0028] Redundant data in the feature data is removed by the second algorithm model to obtain a first feature value having the target number of dimensions.

[0029] In an exemplary embodiment of the present disclosure, the target dimension number is multi-dimensional, and the second target data includes a second expectation and a second variance;

[0030] Acquiring second target data according to the first characteristic value, where the second target data is of the same type as the first target data, includes:

[0031] Calculate the expectation and variance of each of the first eigenvalues ​​in the first eigenvalues ​​having the target number of dimensions, and use the expectation and the variance as the second expectation and the second variance, respectively; or

[0032] An expectation and a variance of a combined eigenvalue formed by combining the first eigenvalues ​​having the target number of dimensions are obtained, and the expectation and the variance are used as the second expectation and the second variance, respectively.

[0033] In an exemplary embodiment of the present disclosure, it is characterized in that the first target data is matched with the second target data, and the user's sharing behavior is predicted according to the matching result, including:

[0034] Matching a first expectation and a first variance in the first target data with a second expectation and a second variance in the second target data, respectively;

[0035] If the first expectation matches the second expectation, and the first variance matches the second variance, the first expectation is used as the sharing probability of the user, and the sharing behavior of the user is predicted according to the sharing probability.

[0036] In an exemplary embodiment of the present disclosure, matching the first target data with the second target data, and predicting the user's sharing behavior according to the matching result, includes:

[0037] If the first expectation does not match the second expectation and / or the first variance does not match the second variance, re-performing dimensionality reduction processing on the feature data to obtain a second feature value having the target number of dimensions;

[0038] Acquire third target data according to the second characteristic value, wherein the third target data includes a third expectation and a third difference;

[0039] matching the third expectation and the third variance with the first expectation and the first variance respectively;

[0040] If the first expectation does not match the third expectation and / or the first variance does not match the third variance, the above steps are repeated until a target expectation and a target variance matching the first expectation and the first variance are obtained.

[0041] According to one aspect of the present disclosure, there is provided a data processing device, comprising:

[0042] A first data processing module is used to obtain user sharing data according to the multi-dimensional original data, and process the user sharing data through a first algorithm model to obtain first target data and a target number of dimensions, wherein the number of the first target data is the same as the target number of dimensions;

[0043] A second data processing module is used to obtain multi-dimensional feature data, and perform dimensionality reduction processing on the feature data through a second algorithm model to obtain a first feature value having the target number of dimensions, wherein the feature data is data in the original data that has a dominant effect on the user's sharing behavior;

[0044] a third data processing module, configured to obtain second target data according to the first characteristic value, wherein the second target data is of the same type as the first target data;

[0045] A matching module is used to match the first target data with the second target data, and predict the user's sharing behavior based on the matching result.

[0046] According to one aspect of the present disclosure, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the data processing method described above is implemented.

[0047] According to one aspect of the present disclosure, there is provided an electronic device, including:

[0048] Processor; and

[0049] A memory, configured to store executable instructions of the processor;

[0050] The processor is configured to perform the data processing method as described above by executing the executable instructions.

[0051] It can be seen from the above technical solutions that the data processing method and device, computer-readable storage medium, and electronic device in the exemplary embodiments of the present disclosure have at least the following advantages and positive effects:

[0052] The present disclosure processes user sharing data through a first algorithm model to obtain first target data and target number of dimensions, wherein the user sharing data is obtained based on multi-dimensional original data; at the same time, the multi-dimensional feature data extracted from the original data is subjected to dimensionality reduction processing through a second algorithm model to obtain a first eigenvalue with a target number of dimensions; then, the second target data of the same type as the first target data is obtained based on the first eigenvalue; finally, the first target data and the second target data are matched, and the user's sharing behavior is predicted based on the matching result. On the one hand, the data processing method of the present disclosure can automatically analyze data and predict user sharing situations, saving manpower and statistical time and reducing costs; on the other hand, it can analyze data of all dimensions, improve the accuracy of the analysis results, and provide accurate guidance for product operations.

[0053] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.

[0055] Figure 1 A schematic diagram showing a flow chart of a data processing method in an exemplary embodiment of the present disclosure;

[0056] Figure 2 An example diagram showing an application scenario of a data processing method in an exemplary embodiment of the present disclosure;

[0057] Figure 3 A schematic diagram showing a process of determining the target dimension number in an exemplary embodiment of the present disclosure is shown;

[0058] Figure 4 A schematic diagram showing a probability density function of a Gaussian mixture model in an exemplary embodiment of the present disclosure;

[0059] Figure 5A schematic diagram showing a flow chart of determining a target number of dimensions according to a first variance in an exemplary embodiment of the present disclosure;

[0060] Figure 6 A schematic diagram showing a flow chart of determining the target dimension number according to the sum of the first standard deviation and the sum of the second standard deviation in an exemplary embodiment of the present disclosure;

[0061] Figure 7 A schematic diagram showing a flow chart of matching first target data with second target data in an exemplary embodiment of the present disclosure;

[0062] Figure 8 A schematic diagram showing the structure of a data processing device in an exemplary embodiment of the present disclosure is shown;

[0063] Fig. 9 A schematic diagram showing the structure of a computer storage medium in an exemplary embodiment of the present disclosure;

[0064] Fig.10 A schematic structural diagram of an electronic device in an exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0065] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concepts of the example embodiments are fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0066] The terms "a", "an", "the" and "said" are used in this specification to indicate the presence of one or more elements / components / etc.; the terms "including" and "having" are used to express an open-ended inclusion and mean that additional elements / components / etc. may exist in addition to the listed elements / components / etc.; the terms "first" and "second" etc. are used only as labels and are not intended to limit the quantity of their objects.

[0067] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and their repeated descriptions will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities.

[0068] In the relevant technologies in this field, taking social e-commerce as an example, when evaluating the sharing and dissemination of users and the dissemination effect, it is necessary to make statistics on e-commerce data, and then the operation personnel manually analyze the statistical results and output statistical reports. However, there are the following problems in the statistical process: (1) There are too many dimensions and dimension combinations of e-commerce data, and manual statistics require a lot of manpower and statistical time of operations, product, and technical personnel; (2) Manual statistics rely on the prediction of product operations. Since the prediction of product operations is very difficult, it is impossible to accurately analyze the sharing intentions and the effects of users in each dimension after sharing, and thus it is impossible to carry out targeted and precise operations; (3) The statistical method is only effective for old users and cannot predict the sharing situation of new users.

[0069] Based on the problems existing in the related art, a data processing method is proposed in one embodiment of the present disclosure to optimize the above problems. Figure 1 As shown, the data processing method can be executed by a server, and at least includes the following steps:

[0070] Step S110: acquiring user sharing data according to the multi-dimensional original data, and processing the user sharing data through a first algorithm model to acquire first target data and a target number of dimensions, wherein the number of the first target data is the same as the target number of dimensions;

[0071] Step S120: Acquire multi-dimensional feature data, and perform dimensionality reduction processing on the feature data through a second algorithm model to obtain a first feature value having the target number of dimensions, wherein the feature data is data in the original data that has a dominant effect on the user's sharing behavior;

[0072] Step S130: acquiring second target data according to the first characteristic value, where the second target data is of the same type as the first target data;

[0073] Step S140: Match the first target data with the second target data, and predict the user's sharing behavior based on the matching result.

[0074] On the one hand, the data processing method in the disclosed embodiment can process the user sharing data and feature data respectively through the first algorithm model and the second algorithm model, thereby avoiding manual statistics, saving a lot of manpower and statistical time, and reducing costs; on the other hand, by matching the first target data and the second target data, the user's sharing situation is predicted according to the matching results, thereby avoiding the prediction of product operation.

[0075] In order to make the technical solution of the present disclosure clearer, the following takes group purchase sharing prediction as an example. Figure 2 The structure shown is used to describe in detail each step of the data processing method in the present disclosure.

[0076] In step S110, user sharing data is obtained based on multi-dimensional original data, and the user sharing data is processed through a first algorithm model to obtain first target data and a target number of dimensions, wherein the number of the first target data is the same as the target number of dimensions.

[0077] In an exemplary embodiment of the present disclosure, a user conducts group shopping through an e-commerce platform in a terminal device 201. The user can send the link of the product to be group-purchased to his friends through a chat tool such as WeChat or QQ, and then the friend can join the group to group-purchase with the user, or send the product link to his friends to attract more people to participate in the group purchase; the server 202 can obtain the original data related to sharing of all users who use the e-commerce platform to group-purchase, such as the product category, the number of times the product link is shared, the number of times the product page is browsed, the number of people who group-purchase, the user name, the user age, the user occupation, etc. In an embodiment of the present disclosure, the user sharing data can be obtained according to the multi-dimensional original data obtained by the server 202. For example, if the user sharing data is the sharing rate of the product, then the sharing rate of the product can be obtained according to the number of times the product link is shared and the number of times the product page is browsed in the original data. The sharing rate of the product is only for the specific product, and does not consider the influence of factors such as users and scenarios. After obtaining the first data, the first data can be processed by the first algorithm model to obtain the first target data and the target dimension number.

[0078] In an exemplary embodiment of the present disclosure, the first target data includes a first expectation and a first variance. Figure 3 A schematic diagram of the process of determining the target number of dimensions is shown, such as Figure 3As shown, in step S301, the user sharing data is clustered by the first algorithm model to obtain multiple groups of user sharing sub-data that conform to the Gaussian distribution; in the process of group buying, the sharing rate of the goods is different for different types of users, different scenarios and other conditions. Usually, the sharing situation of a certain type of characteristic user in a certain scenario conforms to the Gaussian distribution. Therefore, the user sharing data can be clustered by the first algorithm model to obtain multiple groups of user sharing sub-data that conform to the Gaussian distribution. In step S302, the maximum expectation algorithm is used to iteratively obtain the first expectation and the first variance corresponding to each group of user sharing sub-data; the maximum expectation algorithm (Expectation MaximizationAlgorithm, referred to as EM algorithm) is an iterative algorithm used in statistics to find the maximum likelihood estimation of parameters in the probability model that depends on unobservable hidden variables. In the embodiment of the present disclosure, the EM algorithm is used to iteratively process the parameters in the probability model corresponding to each group of user sharing sub-data to find the expectation and maximize the expectation until the model converges to obtain the first expectation and the first variance corresponding to the user sharing sub-data. In step S303, the target number of dimensions is determined based on the first variance. When the number of Gaussian distributions K included in the Gaussian mixture model increases by 1, if the variances of multiple Gaussian distributions do not change significantly, it is considered that the target number of dimensions has been found. Therefore, in the present disclosure, the target number of dimensions can be determined based on the first variance corresponding to the sub-data shared by each group of users.

[0079] In an exemplary embodiment of the present disclosure, since the user sharing data is formed by multiple groups of user sharing sub-data that conform to the Gaussian distribution, the sharing rate of the product conforms to the Gaussian mixture model, that is, the distribution of the user sharing data conforms to the Gaussian mixture model. The Gaussian mixture model (Gaussian Mixture Model, GMM for short), which can also be referred to as MOG for short, uses the Gaussian probability density function (normal distribution curve) to accurately quantify things and decomposes one thing into several models based on the Gaussian probability density function (normal distribution curve).

[0080] In an exemplary embodiment of the present disclosure, the first algorithm model may specifically be a Gaussian mixture model, which is used to cluster user sharing data.

[0081] Figure 4 A schematic diagram of the probability density function of the Gaussian mixture model is shown, Figure 4 As shown, the Gaussian mixture model is composed of three Gaussian distribution curves (normal distribution curves) linearly superimposed, where the dotted lines represent the three Gaussian distribution curves, and the solid lines are formed by fitting the three Gaussian distribution curves. The user sharing data obtained in step S110 of the data processing method of the present disclosure can be considered as Figure 4The sample data shown in the solid line part in the figure can be clustered by the Gaussian mixture model to obtain multiple groups of user sharing sub-data that conform to the Gaussian distribution (i.e. Figure 4 The dashed part in the figure), and the first expectation and the first variance corresponding to each group of user sharing sub-data are obtained by iteratively using the maximum expectation algorithm. It is worth noting that the Gaussian mixture model includes but is not limited to the three Gaussian distribution curves mentioned above, and it may include multiple Gaussian distribution curves, which will not be described in detail in this disclosure.

[0082] Furthermore, the Gaussian mixture model considers that the data is generated from several single Gaussian distribution models, and the corresponding expression is shown in formula (1):

[0083]

[0084] Among them, π k is the weight factor, μ k is the first expectation, Σ k is the first variance and K is the number of Gaussian distributions included in the Gaussian mixture model.

[0085] The Gaussian mixture model is a clustering algorithm in which each Gaussian distribution is a cluster center. When only the sample points are known but the sample classification is unknown, the model parameters (π k ,μ k ,Σ k ).

[0086] In an exemplary embodiment of the present disclosure, Figure 5 A schematic diagram of a process for determining the target dimension number according to the first variance is shown, such as Figure 5 As shown, in step S501, the standard deviation corresponding to the first data of each dimension is determined according to the first variance; in step S502, the standard deviations corresponding to the sub-data shared with the first N groups of users are summed to obtain the sum of the first standard deviations, and the standard deviations corresponding to the sub-data shared with the first N+1 groups of users are summed to obtain the sum of the second standard deviations, wherein N>1 and is a positive integer; as K increases, the sum of the standard deviations must continue to decrease, but when the data is more cohesive, the speed at which the sum of the standard deviations decreases will be significantly reduced if the value of K is increased, so the target number of dimensions can be determined according to the sum of the standard deviations; in step S503, the sum of the first standard deviation is compared with the sum of the second standard deviation, and the target number of dimensions is determined according to the comparison result.

[0087] Furthermore, Figure 6 A flow chart of determining the target dimension number according to the sum of the first standard deviation and the sum of the second standard deviation is shown, as shown in Figure 6As shown, in step S601, the sum of the first standard deviation is subtracted from the sum of the second standard deviation to obtain the increment of the sum of the first standard deviation relative to the sum of the second standard deviation; in step S602, the increment is compared with a preset threshold to determine whether the increment is less than the preset threshold; the smaller the preset threshold is set, the better, that is, the closer the sum of the first standard deviation is to the sum of the second standard deviation, the better; in step S603, if the increment is less than the preset threshold, the number of groups of user sharing sub-data corresponding to the sum of the first standard deviation is the target dimension number.

[0088] In step S120, multi-dimensional feature data is obtained, and the feature data is subjected to dimensionality reduction processing through a second algorithm model to obtain a first feature value having the target number of dimensions, wherein the feature data is data in the original data that has a dominant effect on the user's sharing behavior.

[0089] In an exemplary embodiment of the present disclosure, some multi-dimensional feature data can be extracted from the multi-dimensional original data according to preset conditions. The preset conditions can be dimensional features set by the user. For example, data of dimensions such as promotional scenarios, product categories, user gender, user age, and the number of times a product link is shared in the first data can be extracted as feature data. Furthermore, the setting of the preset conditions can be data dimensions that have a dominant role in the user's sharing behavior set according to user experience. For example, according to user experience, user age, user occupation, product category, and the number of times a product link is shared have an important influence on the user's group purchase sharing. In this case, the preset conditions can be set as user age, user occupation, product category, and the number of times a product link is shared. Then, data corresponding to user age, user occupation, product category, and the number of times a product link is shared can be extracted from the multi-dimensional original data, and these multi-dimensional data can be used as feature data.

[0090] In the exemplary embodiment of the present disclosure, since the amount of multi-dimensional feature data obtained is large, correspondingly, there are also many combinations of multi-dimensional feature data. In actual processing, if all data are processed, it takes a lot of time. Therefore, in order to improve the processing efficiency while ensuring the accuracy of the prediction results, the feature data can be processed by the second algorithm model to reduce the dimension, so as to extract the principal components in the feature data to remove relatively unimportant factors, and at the same time extract the independent variables in the feature data to remove the dependent variables. The first eigenvalue can be obtained by reducing the dimension of the feature data, and the eigenvalue has the same number of dimensions as the target number of dimensions determined in step S110.

[0091] In an exemplary embodiment of the present disclosure, the dimension of the feature data is greater than the target dimension. By establishing a covariance matrix relative to the feature data, the dimension of the feature data is gradually reduced to the target dimension, thereby obtaining a first eigenvalue with the target dimension.

[0092] In an exemplary embodiment of the present disclosure, the second algorithm model can specifically be a principal component analysis (PCA) model. Principal component analysis is a statistical method that converts a set of variables that may be correlated into a set of linearly uncorrelated variables through an orthogonal transformation. The converted set of variables is called principal components.

[0093] In step S130, second target data is acquired according to the first characteristic value, and the second target data is of the same type as the first target data.

[0094] In an exemplary embodiment of the present disclosure, after obtaining the first characteristic value, second target data may be acquired according to the first characteristic value, and the type of the second target data is the same as the type of the first target data, so as to facilitate subsequent matching of the two.

[0095] In an exemplary embodiment of the present disclosure, the first target data and the second target data may both include expectations and variances, that is, the first target data includes a first expectation and a first variance, and the second target data includes a second expectation and a second variance.

[0096] In an exemplary embodiment of the present disclosure, the expectation and variance of the first eigenvalue of each dimension in the first eigenvalues ​​having the target number of dimensions can be calculated, or the expectation and variance of the combined eigenvalue formed by combining the first eigenvalues ​​having the target number of dimensions can be calculated, and the obtained expectation and variance can be used as the second expectation and second variance.

[0097] In step S140, the first target data is matched with the second target data, and an output result is determined according to the matching result.

[0098] In an exemplary embodiment of the present disclosure, after obtaining the first target data and the second target data, the first target data and the second target data can be matched to determine whether the user's group purchase sharing is dominated by the corresponding feature value, and the user's sharing behavior can be predicted based on the matching results.

[0099] In an exemplary embodiment of the present disclosure, Figure 7 A schematic diagram of a process of matching the first target data with the second target data is shown, as shown in Figure 7As shown, in step S701, the first expectation and the first variance in the first target data are matched with the second expectation and the second variance in the second target data respectively; in step S702, if the first expectation matches the second expectation and the first variance matches the second variance, the first expectation can be used as the user's sharing probability, and the sharing probability is stored in the database and output for reference by the operator to predict the user's sharing behavior. For example, when it is matched that the group purchase sharing rate of a certain electronic product by male youths aged 25 to 30 during a big promotion follows a standard normal distribution, then when the user characteristics of a new user are also male youths aged between 25 and 30, it can be predicted that he will share the same electronic product during the big promotion. The sharing of products satisfies the standard normal distribution; in step S703, if the first expectation does not match the second expectation and / or the first variance does not match the second variance, the feature data is re-processed for dimensionality reduction to obtain a second eigenvalue with a target number of dimensions; in step S704, the third target data is obtained according to the second eigenvalue, and the third target data includes a third expectation and a third variance; in step S705, the third expectation and the third variance are matched with the first expectation and the first variance respectively; in step S706, if the first expectation does not match the third expectation and / or the first variance does not match the third variance, steps S703 to S705 are repeated until a target expectation and a target variance matching the first expectation and the first variance are obtained.

[0100] In an exemplary embodiment of the present disclosure, based on the matching results of the first target data and the second target data, the user's sharing probability under the characteristic conditions corresponding to the characteristic data can be predicted. When the first target data matches the second target data, the first expectation (second expectation) is the user's sharing probability; when the first target data does not match the second target data, it is necessary to perform multiple dimensionality reduction processes on the characteristic data to obtain appropriate characteristic values, and then obtain the target expectation and target variance that match the first expectation and first variance in the first target data, then the first expectation (target expectation) is the user's sharing probability.

[0101] The data processing method disclosed in the present invention can be applied to predict the group purchase and sharing situation of users. Since the group purchase and sharing data conforms to the Gaussian distribution and multiple groups of mixed sample data are obtained during actual operation, the group purchase and sharing data of users can be processed by the GMM algorithm to obtain the expectation and variance of the multiple groups of sharing data, and the sum of the standard deviations corresponding to the multiple groups of sharing data is calculated according to the variance, and then the target number of dimensions is determined according to the changing trend of the sum of the standard deviations; in addition, feature data that plays a dominant role in the user's sharing behavior can be extracted from the group purchase and sharing data of users based on experience or previous statistical analysis results, and then the feature data is subjected to dimensionality reduction processing through the PCA algorithm model to obtain eigenvalues ​​with the target number of dimensions, and then the expectation and variance of the data corresponding to the eigenvalues ​​are calculated; finally, the expectation and variance of the multiple groups of sharing data are matched with the expectation and variance of the eigenvalues ​​respectively. If they match, it means that the user's sharing behavior is dominated by the eigenvalue, and the expected value can be used as a prediction of the user's sharing probability; if they do not match, it means that the user's sharing behavior is not dominated by the eigenvalue, and then the feature data needs to be re-processed for dimensionality reduction to obtain eigenvalues ​​with the target number of dimensions that are different from the previous eigenvalues, and then match again, repeating the above steps until the expectation and variance that match the multi-dimensional sharing data are obtained, and the user's sharing probability is predicted based on the expectation.

[0102] The data processing method disclosed in the present invention can avoid product operators spending a lot of manpower and time to count the sharing situations in various dimensions. In addition, it can analyze the sharing intentions of users in all dimensions, and then carry out targeted and precise operations. Furthermore, the data processing method disclosed in the present invention can obtain the expectations and variances of feature users' sharing in specific scenarios through algorithms, and thus predict the sharing situations of new feature users.

[0103] The following describes an embodiment of the device of the present disclosure, which can be used to execute the data processing method of the present disclosure. For details not disclosed in the embodiment of the device of the present disclosure, please refer to the embodiment of the data processing method of the present disclosure.

[0104] Figure 8 The block diagram of a data processing device according to an embodiment of the present disclosure is schematically shown. Figure 8 As shown, the data processing device 800 at least includes a first data processing module 801, a second data processing module 802, a third data processing module 803 and a matching module 804. Specifically:

[0105] A first data processing module 801 is used to obtain user sharing data according to the multi-dimensional original data, and process the user sharing data through a first algorithm model to obtain first target data and a target number of dimensions, wherein the number of the first target data is the same as the target number of dimensions;

[0106] The second data processing module 802 is used to obtain multi-dimensional feature data, and perform dimensionality reduction processing on the feature data through a second algorithm model to obtain a first feature value with the target number of dimensions, wherein the feature data is the data in the original data that has a dominant effect on the user's sharing behavior;

[0107] A third data processing module 803 is used to obtain second target data according to the first characteristic value, where the second target data is of the same type as the first target data;

[0108] The matching module 804 is used to match the first target data with the second target data, and predict the user's sharing behavior according to the matching result.

[0109] In an exemplary embodiment of the present disclosure, the first target data includes a first expectation and a first variance; the first data processing module 801 includes a clustering unit, a calculation unit and a first dimension determination unit, specifically:

[0110] A clustering unit, configured to cluster the user sharing data by using the first algorithm model to obtain a plurality of groups of user sharing sub-data conforming to Gaussian distribution;

[0111] A calculation unit, configured to iteratively obtain a first expectation and a first variance corresponding to each group of user-shared sub-data by using a maximum expectation algorithm;

[0112] The first dimension determination unit is used to determine the target number of dimensions according to the first variance.

[0113] In an exemplary embodiment of the present disclosure, the dimension determination unit includes a standard deviation determination unit, a standard deviation summing unit and a comparison unit, specifically:

[0114] a standard deviation determining unit, configured to determine a standard deviation corresponding to each group of the user-shared sub-data according to the first variance;

[0115] a standard deviation summing unit, configured to sum the standard deviations corresponding to the first N groups of user-shared sub-data to obtain a first standard deviation sum, and sum the standard deviations corresponding to the first N+1 groups of user-shared sub-data to obtain a second standard deviation sum, wherein N>1 and is a positive integer;

[0116] A comparing unit is used to compare the sum of the first standard deviations with the sum of the second standard deviations, and determine the target number of dimensions according to the comparison result.

[0117] In an exemplary embodiment of the present disclosure, the comparison unit includes an increment acquisition unit, a judgment unit, and a second dimension determination unit. Specifically:

[0118] an increment obtaining unit, configured to subtract the sum of the first standard deviations from the sum of the second standard deviations to obtain an increment of the sum of the first standard deviations relative to the sum of the second standard deviations;

[0119] A judging unit, configured to compare the increment with a preset threshold value to judge whether the increment is less than the preset threshold value;

[0120] The second dimension determination unit is used to determine that the number of dimensions corresponding to the sum of the first standard deviations is the target number of dimensions when the increment is less than the preset threshold.

[0121] In an exemplary embodiment of the present disclosure, the second data processing module 802 includes a data extraction unit and a data removal unit, specifically:

[0122] A data extraction unit, used for extracting the characteristic data from the original data according to preset conditions;

[0123] A data removal unit is used to remove redundant data in the feature data through the second algorithm model to obtain a first feature value with the target number of dimensions.

[0124] In an exemplary embodiment of the present disclosure, the target dimension is multi-dimensional, and the second target data includes a second expectation and a second variance; the third data processing module 803 includes:

[0125] a first eigenvalue processing unit, configured to obtain the expectation and variance of each of the first eigenvalues ​​having the target number of dimensions, and use the expectation and the variance as the second expectation and the second variance, respectively; or

[0126] Used to obtain the expectation and variance of the combined eigenvalue formed by combining the first eigenvalues ​​with the target number of dimensions, and use the expectation and the variance as the second expectation and the second variance respectively.

[0127] In an exemplary embodiment of the present disclosure, the matching module 804 includes a first matching unit and a probability determination unit, specifically:

[0128] A first matching unit, configured to match a first expectation and a first variance in the first target data with a second expectation and a second variance in the second target data, respectively;

[0129] A probability determination unit is used to use the first expectation as the sharing probability of the user when the first expectation matches the second expectation and the first variance matches the second variance, and predict the sharing behavior of the user according to the sharing probability.

[0130] In an exemplary embodiment of the present disclosure, the matching module 804 further includes a second feature value acquisition unit, a third target data acquisition unit, a second matching unit and a target data determination unit, specifically:

[0131] A second eigenvalue acquisition unit is used for, when the first expectation does not match the second expectation and / or the first variance does not match the second variance, re-performing dimensionality reduction processing on the feature data to obtain a second eigenvalue having the target number of dimensions;

[0132] A third target data acquisition unit, configured to acquire third target data according to the second characteristic value, wherein the third target data includes a third expectation and a third difference;

[0133] A second matching unit, configured to match the third expectation and the third variance with the first expectation and the first variance respectively;

[0134] The target data determination unit is configured to repeat the above steps when the first expectation does not match the third expectation and / or the first variance does not match the third variance until a target expectation and a target variance matching the first expectation and the first variance are obtained.

[0135] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.

[0136] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.

[0137] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.

[0138] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.

[0139] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as "circuits", "modules" or "systems".

[0140] Refer to the following Fig. 9 The electronic device 900 according to this embodiment of the present disclosure is described. Fig. 9 The electronic device 900 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0141] like Fig. 9 As shown, the electronic device 900 is in the form of a general computing device. The components of the electronic device 900 may include but are not limited to: at least one processing unit 910, at least one storage unit 920, and a bus 930 connecting different system components (including the storage unit 920 and the processing unit 910).

[0142] The storage unit stores program codes, which can be executed by the processing unit 910, so that the processing unit 910 executes the steps of various exemplary embodiments of the present disclosure described in the above “Specific Implementation” section of this specification. For example, the processing unit 910 can execute the following steps: Figure 1 Step S110 shown in: obtaining user sharing data based on multi-dimensional original data, and processing the user sharing data through a first algorithm model to obtain first target data and a target number of dimensions, wherein the number of the first target data is the same as the target number of dimensions; step S120: obtaining multi-dimensional feature data, and performing dimensionality reduction processing on the feature data through a second algorithm model to obtain a first feature value having the target number of dimensions, wherein the feature data is data in the original data that has a dominant role in the user's sharing behavior; step S130: obtaining second target data based on the first feature value, the second target data and the first target data are of the same type; step S140: a matching module, used to match the first target data with the second target data, and predict the user's sharing behavior based on the matching result.

[0143] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 9201 and / or a cache storage unit 9202 , and may further include a read-only storage unit (ROM) 9203 .

[0144] The storage unit 920 may also include a program / utility 9204 having a set (at least one) of program modules 9205, such program modules 9205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0145] Bus 930 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0146] The electronic device 900 may also communicate with one or more external devices 1100 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or may communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 950. Furthermore, the electronic device 900 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 960. As shown, the network adapter 960 communicates with other modules of the electronic device 900 via a bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0147] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.

[0148] In an exemplary embodiment of the present disclosure, a computer storage medium is also provided, on which a program product capable of implementing the above method of the present specification is stored. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary implementations of the present disclosure described in the above "Exemplary Method" section of the present specification.

[0149] refer to Fig.10 As shown, a program product 1000 for implementing the above method according to an embodiment of the present disclosure is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.

[0150] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0151] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0152] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.

[0153] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0154] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.

[0155] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.

Claims

1. A data processing method, characterized in that, the method includes: obtaining user sharing data according to multi-dimensional original data, and processing the user sharing data through a first algorithm model to obtain first target data and a target dimension number, wherein the number of the first target data is the same as the target dimension number; the user sharing data is the sharing rate of a commodity; extracting data dimensions that play a dominant role in the user's sharing behavior from the multi-dimensional original data as multi-dimensional feature data, and performing dimensionality reduction processing on the feature data through a second algorithm model to obtain first eigenvalues with the target dimension number; obtaining second target data according to the first eigenvalues, and the second target data and the first target data are of the same type; matching the first target data with the second target data, and predicting the user's sharing behavior according to the matching result; the matching the first target data with the second target data, and predicting the user's sharing behavior according to the matching result includes: matching the first expectation and the first variance in the first target data with the second expectation and the second variance in the second target data respectively, so as to judge whether the group buying sharing behavior is dominated by the corresponding eigenvalues according to the matching result; if the first expectation matches the second expectation, and the first variance matches the second variance, then taking the first expectation as the sharing probability of the user, and predicting the user's sharing behavior according to the sharing probability.

2. The data processing method according to claim 1, characterized in that, obtaining user sharing data according to multi-dimensional original data, and clustering the user sharing data through a first algorithm model to obtain first target data and a target dimension number, includes: clustering the user sharing data through the first algorithm model to obtain multiple groups of user sharing sub-data conforming to the Gaussian distribution; iteratively obtaining the first expectation and the first variance corresponding to each group of the user sharing sub-data by using the maximum expectation algorithm; determining the target dimension number according to the first variance.

3. The data processing method according to claim 2, characterized in that, determining the target dimension number according to the first variance includes: determining the standard deviation corresponding to each group of the user sharing sub-data according to the first variance; summing the standard deviations corresponding to the first N groups of the user sharing sub-data to obtain a first sum of standard deviations, and summing the standard deviations corresponding to the first N + 1 groups of the user sharing sub-data to obtain a second sum of standard deviations, wherein N > 1 and is a positive integer; comparing the first sum of standard deviations with the second sum of standard deviations, and determining the target dimension number according to the comparison result.

4. The data processing method according to claim 3, characterized in that, comparing the first sum of standard deviations with the second sum of standard deviations, and determining the target dimension number according to the comparison result includes: subtracting the first sum of standard deviations from the second sum of standard deviations to obtain the increment of the first sum of standard deviations relative to the second sum of standard deviations; Compare the increment with a preset threshold to determine whether the increment is less than the preset threshold; If the increment is less than the preset threshold, the number of dimensions corresponding to the sum of the first standard deviations is the target number of dimensions.

5. The data processing method according to claim 1, characterized in that, obtaining multi-dimensional feature data, and performing dimensionality reduction processing on the feature data through a second algorithm model to obtain first eigenvalues with the target number of dimensions, including: extracting the feature data from the original data according to preset conditions; removing redundant data in the feature data through the second algorithm model to obtain first eigenvalues with the target number of dimensions.

6. The data processing method according to claim 5, characterized in that, the target number of dimensions is multi-dimensional, and the second target data includes a second expectation and a second variance; obtaining second target data according to the first eigenvalues, and the second target data and the first target data are of the same type, including: calculating the expectation and variance of each of the first eigenvalues in the first eigenvalues with the target number of dimensions, and respectively using the expectation and the variance as the second expectation and the second variance; or calculating the expectation and variance of the combined eigenvalues formed by combining the first eigenvalues with the target number of dimensions, and respectively using the expectation and the variance as the second expectation and the second variance.

7. The data processing method according to claim 1, characterized in that, matching the first target data with the second target data, and predicting the sharing behavior of the user according to the matching result, including: if the first expectation does not match the second expectation and / or the first variance does not match the second variance, re-performing dimensionality reduction processing on the feature data to obtain second eigenvalues with the target number of dimensions; obtaining third target data according to the second eigenvalues, and the third target data includes a third expectation and a third variance; matching the third expectation and the third variance with the first expectation and the first variance respectively; if the first expectation does not match the third expectation and / or the first variance does not match the third variance, repeat the above steps until target expectations and target variances matching the first expectation and the first variance are obtained.

8. A data processing device, characterized in that, comprising: a first data processing module, configured to obtain user sharing data according to multi-dimensional original data, and process the user sharing data through a first algorithm model to obtain first target data and a target number of dimensions, wherein the number of the first target data is the same as the target number of dimensions; the user sharing data is the sharing rate of a commodity; a second data processing module, configured to extract data dimensions that play a dominant role in the sharing behavior of the user from the multi-dimensional original data as multi-dimensional feature data, and perform dimensionality reduction processing on the feature data through a second algorithm model to obtain first eigenvalues with the target number of dimensions; A third data processing module, configured to obtain second target data according to the first eigenvalue, where the type of the second target data is the same as that of the first target data; A matching module, configured to match the first target data with the second target data and predict the sharing behavior of the user according to the matching result; The matching the first target data with the second target data and predicting the sharing behavior of the user according to the matching result includes: Matching the first expectation and the first variance in the first target data with the second expectation and the second variance in the second target data respectively, so as to judge whether the group buying sharing behavior is dominated by the corresponding eigenvalue according to the matching result; If the first expectation matches the second expectation and the first variance matches the second variance, then use the first expectation as the sharing probability of the user, and predict the sharing behavior of the user according to the sharing probability.

9. A computer storage medium, on which a computer program is stored, wherein, the computer program, when executed by a processor, implements the data processing method according to any one of claims 1 to 7.

10. An electronic device, wherein, comprising: a processor; and a memory, configured to store executable instructions of the processor; wherein, the processor is configured to execute the data processing method according to any one of claims 1 to 7 by executing the executable instructions.

Citation Information

Patent Citations

  • User behavior prediction method and device

    CN108121795A