Methods, devices, and electronic equipment for quantifying social media account operation data
By acquiring multi-dimensional data from social media accounts, performing standardized processing and hierarchical feature calculations, the problems of scoring errors and single-dimensionality in the quantification of social media account operation data are solved, enabling more accurate and comprehensive evaluation and guiding users to optimize their operational strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ONE NETWORK INTEROPERABILITY (BEIJING) TECH CO LTD
- Filing Date
- 2023-11-10
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies for quantifying social media account operation data suffer from problems such as high scoring error rates, limited scoring dimensions, and unscientific algorithms, resulting in inaccurate and incomplete scoring.
By acquiring multi-dimensional data of target social media accounts within a preset time period, cleaning and preprocessing are performed using standardized processing functions, a hierarchical feature calculation function is constructed, and quantitative values of social media account operation data are calculated based on hierarchical spatial order to provide multi-dimensional evaluation.
It improves the accuracy and comprehensiveness of ratings, provides a more comprehensive assessment of social media accounts, helps users better understand account performance, and reduces reliance on manually labeled data.
Smart Images

Figure CN117421491B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, and electronic device for quantifying social media account operation data. Background Technology
[0002] Social media account performance data refers to a set of metrics used to evaluate and manage the performance of social media accounts. This data helps social media account operators understand how their accounts are performing on various social media platforms, as well as optimize content and operational strategies to increase exposure and brand awareness.
[0003] Common social media account performance data includes:
[0004] Exposure: Exposure refers to the ranking of a self-media account in search engines and its exposure on social media platforms. By analyzing exposure, one can understand the exposure of a self-media account on different platforms, thereby optimizing content and operational strategies to increase exposure and brand awareness.
[0005] Page views: Page views refer to the number of times a self-media account's pages are viewed in search engines and on social media platforms. By analyzing page views, you can understand the popularity of your self-media account's content and whether your content can attract more users. At the same time, you can also understand the preferences and needs of different user groups, allowing you to better adjust your content.
[0006] Likes: Likes refer to the number of times users like the content of a self-media account. By analyzing likes, we can understand the degree of user approval of the content of a self-media account.
[0007] Comment volume: Comment volume refers to the number of comments users make on the content of a self-media account. By analyzing the comment volume, we can understand users' reactions and feedback to the content of a self-media account.
[0008] Shares: Shares refer to the number of times users share content from a self-media account to other social media platforms. By analyzing shares, one can understand whether the content of a self-media account has virality and influence.
[0009] Follower count: Follower count refers to the number of users who follow a self-media account. By analyzing follower count, one can understand the influence and popularity of a self-media account.
[0010] Interaction rate: The interaction rate refers to the proportion of users who interact with the content of a self-media account. By analyzing the interaction rate, we can understand the degree of interaction and participation of users with self-media accounts.
[0011] This data can help social media account operators understand their account performance, optimize content and operational strategies, and increase exposure and brand awareness.
[0012] The quantification of social media account operation data in existing technologies has the following problems:
[0013] 1. High rating error rate: The limited number of social media account samples leads to inaccurate ratings and a large deviation from reality;
[0014] 2. Limited scoring dimensions: Scores calculated based on a limited number of data dimensions cannot fully reflect the operational performance of social media platforms;
[0015] 3. The algorithm is not scientific enough: the score is obtained by simply adding weights, and the weights are determined by subjective judgment.
[0016] The aforementioned issues have become technical problems that need to be solved. Summary of the Invention
[0017] In view of this, embodiments of the present invention provide a method, apparatus, and electronic device for quantifying social media account operation data, which at least partially solves the problems existing in the prior art.
[0018] In a first aspect, embodiments of the present invention provide a method for quantifying social media account operational data, including:
[0019] After obtaining the target social media account, multi-dimensional data of the target social media account are collected within a preset time period [t1,t2] to form a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions;
[0020] After cleaning and preprocessing the data in the multi-dimensional dataset D using a standardized processing function, the cleaned and preprocessed data is classified and calculated to obtain a classified dataset U={U1,U2,…,Um} containing m dimensions, and the feature vector K=[K1,K2,…,Km] corresponding to the classified dataset K, where m<n;
[0021] Based on the preset hierarchical spatial order and the feature value vector, a hierarchical feature calculation function corresponding to the classified data set is constructed, and the hierarchical value Y={Y1,Y2,Y3} of the classified data set in the three hierarchical levels is calculated through the hierarchical feature calculation function.
[0022] Based on the hierarchical value Y, the quantitative value W of the social media account's operating data is determined, and the quantitative value W is output to the user in numerical or graphical form.
[0023] According to a specific implementation of an embodiment of this disclosure, the step of collecting multi-dimensional data of the target social media account within a preset time interval [t1, t2] to form a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions includes:
[0024] Determine the data query conditions corresponding to n dimensions;
[0025] Using the data query conditions and the preset time period [t1,t2] as the data query request, multi-dimensional data of the target social media account within the preset time period [t1,t2] is collected.
[0026] According to a specific implementation of an embodiment of this disclosure, the step of collecting multi-dimensional data of the target social media account within a preset time interval [t1,t2] to form a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions further includes:
[0027] Set the following metrics as data query requests: number of followers, net increase in followers in the past 7 days, maximum follower growth in the past 7 days, number of posts in the past 7 days, number of interactions per thousand followers in the past 7 days, total number of interactions per post in the past 7 days, maximum number of interactions per post in the past 7 days, usage of hashtags in posts in the past 7 days, and maximum number of interactions per hashtag post.
[0028] According to a specific implementation of this disclosure, after cleaning and preprocessing the data in the multi-dimensional data set D using a standardization processing function, the cleaned and preprocessed data is classified to obtain a classified data set U={U1,U2,…,Um} containing m dimensions, and a feature vector K=[K1,K2,…,Km] corresponding to the classified data set K, including:
[0029] Perform outlier removal, missing data processing, and data format standardization on the data in the multidimensional dataset D.
[0030] According to a specific implementation of this disclosure, after cleaning and preprocessing the data in the multi-dimensional data set D using a standardization processing function, the cleaned and preprocessed data is classified and calculated to obtain a classified data set U={U1,U2,…,Um} containing m dimensions, and an eigenvector K=[K1,K2,…,Km] corresponding to the classified data set K, the method further includes:
[0031] Based on the requirements of the classification task, a preset feature extraction method is selected to extract classification features related to the classification task from the cleaned and preprocessed data. The classification features include numerical features, text features, and image features.
[0032] Based on the extracted classification features, a classifier for classification calculation is constructed, which includes the classification features.
[0033] The trained classifier is used to classify the data obtained after cleaning and preprocessing, resulting in a data set U={U1,U2,…,Um} containing m dimensions, where each dimension represents a category of data.
[0034] For the i-th class of categorized data Ui in the categorized dataset U, calculate the distance value Oi and the dispersion Si of categorized data Ui in the categorized dataset, so as to determine the feature value Ki corresponding to the categorized data Ui based on the distance value and the dispersion:
[0035]
[0036] γ and ρ are the first and second adjustment parameters, respectively.
[0037] According to a specific implementation of this disclosure, the step of constructing a hierarchical feature calculation function corresponding to the categorized data set based on a preset hierarchical spatial order and the feature value vector, and calculating the stratification values Y={Y1,Y2,Y3} of the categorized data set in three strata using the hierarchical feature calculation function, includes:
[0038] The categorized dataset U is segmented to obtain a word vector [w1, w2, ..., wQ] containing Q words;
[0039] Using word embedding matrix Transform sentence s into a vector sequence [e1,e2,…,eQ], where V and D represent the vocabulary size and word embedding dimension, respectively;
[0040] Define the hierarchical value of ei Represented as:
[0041]
[0042] The stratification value Y1 of the categorized dataset in the first stratum is represented as:
[0043]
[0044] in These are the third and fourth adjustment parameters, respectively, and ReLU is a nonlinear activation function.
[0045] According to a specific implementation of this disclosure, the step of constructing a hierarchical feature calculation function corresponding to the categorized data set based on a preset hierarchical spatial order and the feature value vector, and calculating the stratification values Y={Y1,Y2,Y3} of the categorized data set in three strata using the hierarchical feature calculation function, further includes:
[0046] Statistical analysis is performed on the sentences contained in the categorized dataset U to obtain R sentence vector values [x1, x2, ..., xR].
[0047] Then the stratification value Y2 of the categorized dataset in the second stratum is represented as:
[0048]
[0049] in These are the fifth and sixth adjustment parameters, respectively.
[0050] According to a specific implementation of this disclosure, the step of constructing a hierarchical feature calculation function corresponding to the categorized data set based on a preset hierarchical spatial order and the feature value vector, and calculating the stratification values Y={Y1,Y2,Y3} of the categorized data set in three strata using the hierarchical feature calculation function, further includes:
[0051] Statistical analysis is performed on the comment words contained in the categorized dataset U to obtain Z comment vector values [f1, f2, ..., fZ].
[0052] The stratification value Y3 of the categorized dataset in the third stratum is represented as:
[0053]
[0054] Where τ is the seventh adjustment parameter, and f0 is the mean of the comment vector values [f1, f2, ..., fZ].
[0055] Secondly, embodiments of the present invention provide a device for quantifying social media account operational data, comprising:
[0056] The data acquisition module is used to acquire the target social media account and then collect multi-dimensional data of the target social media account within a preset time period [t1,t2] to form a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions.
[0057] The classification module is used to clean and preprocess the data in the multi-dimensional data set D using a standardization processing function, and then perform classification calculations on the cleaned and preprocessed data to obtain a classification data set U={U1,U2,…,Um} containing m dimensions, and the feature vector K=[K1,K2,…,Km] corresponding to the classification data set K, where m<n;
[0058] The calculation module is used to construct the hierarchical feature calculation function corresponding to the classified data set based on the preset hierarchical spatial order and the feature value vector, and to calculate the hierarchical value Y={Y1,Y2,Y3} of the classified data set in three hierarchical layers through the hierarchical feature calculation function.
[0059] The determination module determines the quantized value W of the social media account's operational data based on the hierarchical value Y, and outputs the quantized value W to the user in numerical or graphical form.
[0060] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0061] At least one processor; and,
[0062] The memory is communicatively connected to the at least one processor; wherein,
[0063] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method for quantifying social media account operation data in any of the first aspects or any implementations thereof.
[0064] Fourthly, embodiments of the present invention also provide a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the method for quantifying social media account operation data in the first aspect or any implementation thereof.
[0065] Fifthly, embodiments of the present invention also provide a computer program product, the computer program product including a computing program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the method for quantifying social media account operation data in the first aspect or any implementation thereof.
[0066] The quantification scheme for social media account operation data in this embodiment of the invention includes: after acquiring a target social media account, collecting multi-dimensional data of the target social media account within a preset time period [t1, t2] to form a multi-dimensional data set D={D1, D2, ..., Dn} containing n dimensions; cleaning and preprocessing the data in the multi-dimensional data set D using a standardization processing function, and then classifying and calculating the data obtained after cleaning and preprocessing to obtain a classified data set U={U1, U2, ..., Um} containing m dimensions, and a feature value vector K=[K1, K2, ..., Km] corresponding to the classified data set K, where m < n; constructing a hierarchical feature calculation function corresponding to the classified data set based on a preset hierarchical spatial order and the feature value vector, and calculating the hierarchical values Y={Y1, Y2, Y3} of the classified data set in three hierarchical layers using the hierarchical feature calculation function; determining the quantization value W of the social media account operation data based on the hierarchical value Y, and outputting the quantization value W to the user in numerical or graphical form. The proposed solution provides a more comprehensive evaluation of social media accounts based on multi-dimensional data, helping users better understand the account's operational status. This solution does not rely on limited manually labeled data, but automatically obtains parameters from an expanded pool of social media accounts, improving the fairness and accuracy of the scoring. Attached Figure Description
[0067] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0068] Figure 1 This is a schematic flowchart of a method for quantifying social media account operation data provided in an embodiment of the present invention;
[0069] Figure 2 A schematic diagram of another method for quantifying social media account operation data provided in an embodiment of the present invention;
[0070] Figure 3 A schematic diagram of the structure of the quantification device for social media account operation data provided in an embodiment of the present invention;
[0071] Figure 4 A schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0072] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0073] The following specific examples illustrate the implementation of this disclosure. Those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0074] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.
[0075] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The drawings only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0076] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0077] This disclosure provides a method for quantifying social media account operational data. The method for quantifying social media account operational data provided in this embodiment can be executed by a computing device, which can be implemented as software or a combination of software and hardware. This computing device can be integrated into a server, terminal device, or the like.
[0078] See Figure 1 and Figure 2 This disclosure provides a method for quantifying social media account operation data, including:
[0079] S101, After obtaining the target social media account, collect multi-dimensional data of the target social media account within a preset time period [t1,t2] to form a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions.
[0080] The process of collecting multi-dimensional data can be achieved using APIs provided by social media platforms. This process may include the following steps:
[0081] Obtaining a target social media account: This can be achieved in several ways. Once you have the target social media account, you can input it into a social media API so that the API can retrieve all the account's data. Most social media platforms provide APIs that allow developers to obtain data. These APIs can be used to retrieve data from a target account within a preset time period. For example, Twitter, Facebook, and Instagram all provide such APIs.
[0082] Based on the preset time period [t1, t2], determine the time interval for data collection, such as collecting data within one week or one month.
[0083] Based on the data type and requirements, the collected data will be integrated into a multi-dimensional dataset. For example, it may be necessary to integrate text data, image data, and interactive data together.
[0084] After completing the above steps, you will get a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions, where each dimension Di represents a specific aspect of the target social media account's data within a preset time period (such as the number of posts, likes, comments, etc.).
[0085] For example, multi-dimensional data can be automatically obtained from target social media accounts, including but not limited to the following:
[0086] ① Number of followers: The number of followers or followers of the account at present.
[0087] ② Net increase in followers in the past 7 days: The net increase in the number of followers in the past 7 days, that is, new followers minus followers who unfollowed.
[0088] ③ Maximum fan growth in the past 7 days: The maximum increase in the number of fans in the past 7 days, that is, the maximum change in the number of fans in 7 days.
[0089] ④ Number of posts in the past 7 days: The number of posts published by the account in the past 7 days.
[0090] ⑤ Number of interactions per thousand followers in the past 7 days: The number of interactions per thousand followers of the account in the past 7 days, usually used to measure the account's interaction rate.
[0091] ⑥ Total interactions with posts in the past 7 days: The total number of interactions with all posts in the past 7 days, including likes, comments, shares, etc.
[0092] ⑦ Maximum number of interactions for posts in the past 7 days: The maximum number of interactions for a single post in the past 7 days.
[0093] ⑧ Hashtag Usage in Posts in the Past 7 Days: The number of hashtags used by the account in the past 7 days.
[0094] ⑨ Maximum interaction count for hashtag posts: The post with the most interactions among those that used hashtags.
[0095] S102, after cleaning and preprocessing the data in the multi-dimensional data set D using the standardization processing function, the cleaned and preprocessed data is classified and calculated to obtain a classified data set U={U1,U2,…,Um} containing m dimensions, and the feature vector K=[K1,K2,…,Km] corresponding to the classified data set K, where m<n.
[0096] Data cleaning and preprocessing are crucial steps in data mining and analysis, improving data accuracy and consistency, making it more suitable for subsequent analysis and model training. Data cleaning and preprocessing methods can include:
[0097] Delete invalid or missing data: If there is invalid, missing or anomalous data in the dataset, consider deleting it to avoid interfering with subsequent analysis.
[0098] Data standardization: Data standardization is a common data preprocessing method that transforms data to a uniform scale, eliminating scale differences between data points. This can be achieved by subtracting the mean from the data in each dimension and then dividing by the standard deviation. This ensures that the data in each dimension are distributed around 0 and have the same scale.
[0099] Data normalization: Data normalization scales data proportionally to fit it into a small, specific range. Normalization can speed up computation and avoid subtle differences in magnitude. For example, it scales data uniformly to the range [0,1].
[0100] Handling missing values: For data with missing values, interpolation, deletion, or regression can be used to handle them. For example, methods such as mean imputation, regression imputation, or multiple imputation can be used to fill missing values.
[0101] Noise Removal: If there is noise in the dataset, you can consider using some noise reduction methods, such as filtering and smoothing, to reduce the impact of noise.
[0102] Data transformation: Converting data into a more suitable format or form can improve model performance. For example, text data can be converted into numerical data, or high-dimensional data can be converted into low-dimensional data.
[0103] These steps are used to clean and preprocess the data in the multidimensional dataset D, preparing it for subsequent data analysis and model training.
[0104] After cleaning and preprocessing, the data is classified and calculated to obtain a classified data set U={U1,U2,…,Um} containing m dimensions.
[0105] In this process, each dimension Ui (1≤i≤m) may represent a specific aspect of the data. The values within each dimension Ui are classified and calculated, and then divided into several categories, thus forming a dataset containing multiple categories.
[0106] Classification calculations can typically be implemented using clustering algorithms, such as K-means clustering, hierarchical clustering, and DBSCAN. These algorithms can divide data into different categories based on similarity or distance.
[0107] After obtaining the categorized dataset U, we can further analyze information such as the category distribution of each dimension and the differences between categories to better understand the structure and characteristics of the data. At the same time, the categorized dataset also provides a simpler and more intuitive way to represent data for subsequent data analysis and modeling.
[0108] For the i-th class of categorized data Ui in the categorized dataset U, calculate the distance value Oi and the dispersion Si of categorized data Ui in the categorized dataset, so as to determine the feature value Ki corresponding to the categorized data Ui based on the distance value and the dispersion:
[0109]
[0110] γ and ρ are the first and second adjustment parameters, respectively.
[0111] S103, based on the preset hierarchical spatial order and the feature value vector, construct the hierarchical feature calculation function corresponding to the classified data set, and calculate the hierarchical value Y={Y1,Y2,Y3} of the classified data set in the three hierarchical layers through the hierarchical feature calculation function.
[0112] Based on a predefined hierarchical spatial order and feature value vectors, a hierarchical feature calculation function corresponding to a categorized dataset can be constructed. This function calculates and evaluates the feature values of each dimension according to the predefined hierarchical spatial order.
[0113] Specifically, the hierarchical feature calculation function can be performed according to the following steps:
[0114] Obtain a preset hierarchical spatial order, which can be set according to needs and data characteristics. For example, the hierarchical spatial order can be: {word, sentence, attention}, or it can be {word, sentence, emotion}.
[0115] For each dimension of the feature vector K, the feature values can be sorted and layered according to the preset hierarchical spatial order. The elements in the feature vector K can be sorted in descending order.
[0116] Construct a hierarchical feature calculation function that accepts a feature value vector as input and calculates and evaluates the feature values according to a predefined hierarchical spatial order. The specific calculation method can be selected according to the needs and data characteristics, such as calculating the quantity, proportion, and center position of each category.
[0117] The stratified feature calculation function is used to process the feature vector Ki of each dimension Ui in the categorized dataset U to obtain the stratified feature values for each dimension. These stratified feature values can be further used for analysis and modeling.
[0118] By constructing hierarchical feature calculation functions, we can more flexibly process and parse the feature value vectors of categorized data sets, and perform deeper analysis and mining of the data according to the preset hierarchical spatial order.
[0119] Specifically, the categorized dataset U can be segmented to obtain a word vector [w1, w2, ..., wQ] containing Q words;
[0120] Using word embedding matrix Transform sentence s into a vector sequence [e1,e2,…,eQ], where V and D represent the vocabulary size and word embedding dimension, respectively;
[0121] Define the hierarchical value of ei Represented as:
[0122]
[0123] The stratification value Y1 of the categorized dataset in the first stratum is represented as:
[0124]
[0125] in These are the third and fourth adjustment parameters, respectively, and ReLU is a nonlinear activation function.
[0126] Furthermore, statistics can be performed on the sentences contained in the categorized dataset U to obtain R sentence vector values [x1, x2, ..., xR].
[0127] Then the stratification value Y2 of the categorized dataset in the second stratum is represented as:
[0128]
[0129] in These are the fifth and sixth adjustment parameters, respectively.
[0130] Furthermore, the comment words contained in the categorized dataset U can be statistically analyzed to obtain Z comment vector values [f1, f2, ..., fZ].
[0131] The stratification value Y3 of the categorized dataset in the third stratum is represented as:
[0132]
[0133] Where τ is the seventh adjustment parameter, and f0 is the mean of the comment vector values [f1, f2, ..., fZ].
[0134] S104, based on the layer value Y, determine the quantization value W of the social media account operation data, and output the quantization value W to the user in numerical or graphical form.
[0135] Based on the hierarchical feature value Y, the features of each dimension are weighted and summed to obtain the comprehensive feature value of each dimension. This process can determine the weight of each hierarchical feature according to a pre-defined hierarchical spatial order.
[0136] Dividing the composite feature value of each dimension by the sum of the composite feature values of all dimensions yields the quantified value W of the social media account's performance data. This quantified value W reflects the overall performance of the social media account and allows for comparison and analysis across different dimensions.
[0137] The proposed solution provides a more comprehensive social media account evaluation based on multi-dimensional data, helping users better understand their account's performance. Instead of relying on limited manually labeled data, the solution automatically extracts parameters from an expanded pool of social media accounts, improving the fairness and accuracy of the scoring. The final calculation results can guide users in developing future social media operation strategies for better account management and optimization.
[0138] According to a specific implementation of an embodiment of this disclosure, the step of collecting multi-dimensional data of the target social media account within a preset time interval [t1, t2] to form a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions includes:
[0139] Determine the data query conditions corresponding to n dimensions;
[0140] Using the data query conditions and the preset time period [t1,t2] as the data query request, multi-dimensional data of the target social media account within the preset time period [t1,t2] is collected.
[0141] According to a specific implementation of an embodiment of this disclosure, the step of collecting multi-dimensional data of the target social media account within a preset time interval [t1,t2] to form a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions further includes:
[0142] Set the following metrics as data query requests: number of followers, net increase in followers in the past 7 days, maximum follower growth in the past 7 days, number of posts in the past 7 days, number of interactions per thousand followers in the past 7 days, total number of interactions per post in the past 7 days, maximum number of interactions per post in the past 7 days, usage of hashtags in posts in the past 7 days, and maximum number of interactions per hashtag post.
[0143] According to a specific implementation of this disclosure, after cleaning and preprocessing the data in the multi-dimensional data set D using a standardization processing function, the cleaned and preprocessed data is classified to obtain a classified data set U={U1,U2,…,Um} containing m dimensions, and a feature vector K=[K1,K2,…,Km] corresponding to the classified data set K, including:
[0144] Perform outlier removal, missing data processing, and data format standardization on the data in the multidimensional dataset D.
[0145] According to a specific implementation of this disclosure, after cleaning and preprocessing the data in the multi-dimensional data set D using a standardization processing function, the cleaned and preprocessed data is classified and calculated to obtain a classified data set U={U1,U2,…,Um} containing m dimensions, and an eigenvector K=[K1,K2,…,Km] corresponding to the classified data set K, the method further includes:
[0146] Based on the requirements of the classification task, a preset feature extraction method is selected to extract classification features related to the classification task from the cleaned and preprocessed data. The classification features include numerical features, text features, and image features.
[0147] Based on the extracted classification features, a classifier for classification calculation is constructed, which includes the classification features.
[0148] The trained classifier is used to classify the data obtained after cleaning and preprocessing, resulting in a data set U={U1,U2,…,Um} containing m dimensions, where each dimension represents a category of data.
[0149] For the i-th class of categorized data Ui in the categorized dataset U, calculate the distance value Oi and the dispersion Si of categorized data Ui in the categorized dataset, so as to determine the feature value Ki corresponding to the categorized data Ui based on the distance value and the dispersion:
[0150]
[0151] γ and ρ are the first and second adjustment parameters, respectively.
[0152] Calculating the distance value Oi and the dispersion Si of the categorized data Ui in the categorized dataset allows for further analysis of the data distribution and characteristics.
[0153] The distance value Oi can be used to measure the similarity or distance between categorized data Ui and other data. Different distance metrics can be used to calculate the distance value Oi, such as Euclidean distance, Manhattan distance, Mahalanobis distance, etc. The specific calculation method can be chosen based on the characteristics of the data and the specific needs.
[0154] The dispersion Si can be used to measure the differences or degree of dispersion within categorized data Ui. The method for calculating the dispersion Si can be determined according to the type and characteristics of the data; for example, statistical indicators such as variance and standard deviation can be used to calculate the dispersion Si.
[0155] By calculating the distance value Oi and the dispersion Si, we can gain a deeper understanding of the distribution and characteristics of the data.
[0156] According to a specific implementation of this disclosure, the step of constructing a hierarchical feature calculation function corresponding to the categorized data set based on a preset hierarchical spatial order and the feature value vector, and calculating the stratification values Y={Y1,Y2,Y3} of the categorized data set in three strata using the hierarchical feature calculation function, includes:
[0157] The categorized dataset U is segmented to obtain a word vector [w1, w2, ..., wQ] containing Q words;
[0158] Using word embedding matrix Transform sentence s into a vector sequence [e1,e2,…,eQ], where V and D represent the vocabulary size and word embedding dimension, respectively;
[0159] Define the hierarchical value of ei Represented as:
[0160]
[0161] The stratification value Y1 of the categorized dataset in the first stratum is represented as:
[0162]
[0163] in These are the third and fourth adjustment parameters, respectively, and ReLU is a nonlinear activation function.
[0164] According to a specific implementation of this disclosure, the step of constructing a hierarchical feature calculation function corresponding to the categorized data set based on a preset hierarchical spatial order and the feature value vector, and calculating the stratification values Y={Y1,Y2,Y3} of the categorized data set in three strata using the hierarchical feature calculation function, further includes:
[0165] Statistical analysis is performed on the sentences contained in the categorized dataset U to obtain R sentence vector values [x1, x2, ..., xR].
[0166] Then the stratification value Y2 of the categorized dataset in the second stratum is represented as:
[0167]
[0168] in These are the fifth and sixth adjustment parameters, respectively.
[0169] According to a specific implementation of this disclosure, the step of constructing a hierarchical feature calculation function corresponding to the categorized data set based on a preset hierarchical spatial order and the feature value vector, and calculating the stratification values Y={Y1,Y2,Y3} of the categorized data set in three strata using the hierarchical feature calculation function, further includes:
[0170] Statistical analysis is performed on the comment words contained in the categorized dataset U to obtain Z comment vector values [f1, f2, ..., fZ].
[0171] The stratification value Y3 of the categorized dataset in the third stratum is represented as:
[0172]
[0173] Where τ is the seventh adjustment parameter, and f0 is the mean of the comment vector values [f1, f2, ..., fZ].
[0174] For a corresponding method embodiment, see [link to relevant documentation]. Figure 3 The present invention also discloses a device 30 for quantifying social media account operation data, comprising:
[0175] The acquisition module 301 is used to acquire the target social media account and then acquire multi-dimensional data of the target social media account within a preset time period [t1,t2] to form a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions.
[0176] The classification module 302 is used to clean and preprocess the data in the multi-dimensional data set D using a standardization processing function, and then perform classification calculations on the cleaned and preprocessed data to obtain a classification data set U={U1,U2,…,Um} containing m dimensions, and the feature vector K=[K1,K2,…,Km] corresponding to the classification data set K, where m<n;
[0177] The calculation module 303 is used to construct a hierarchical feature calculation function corresponding to the classified data set based on a preset hierarchical spatial order and the feature value vector, and to calculate the hierarchical value Y={Y1,Y2,Y3} of the classified data set in three hierarchical layers through the hierarchical feature calculation function.
[0178] The determination module 304 determines the quantization value W of the social media account operation data based on the layer value Y, and outputs the quantization value W to the user in numerical or graphical form.
[0179] See Figure 4 This invention also provides an electronic device 60, which includes:
[0180] At least one processor; and,
[0181] The memory is communicatively connected to the at least one processor; wherein,
[0182] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method for quantifying social media account running data in the foregoing method embodiments.
[0183] This invention also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the aforementioned method embodiments.
[0184] This invention also provides a computer program product, which includes a computing program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to perform the method for quantifying social media account operation data in the aforementioned method embodiments.
[0185] The following is for reference. Figure 4 The diagram illustrates a structural schematic of an electronic device 60 suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0186] like Figure 4 As shown, electronic device 60 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 60. Processing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0187] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 60 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 60 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0188] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0189] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for quantifying social media account operation data, characterized in that, include: After obtaining the target social media account, multi-dimensional data of the target social media account are collected within a preset time period [t1,t2] to form a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions; After cleaning and preprocessing the data in the multi-dimensional dataset D using a standardized processing function, the cleaned and preprocessed data is classified and calculated to obtain a classified dataset U={U1,U2,…,Um} containing m dimensions, and the feature vector K=[K1,K2,…,Km] corresponding to the classified dataset, where m<n; Based on the preset hierarchical spatial order and the feature value vector, a hierarchical feature calculation function corresponding to the classified data set is constructed, and the hierarchical value Y={Y1,Y2,Y3} of the classified data set in the three hierarchical levels is calculated through the hierarchical feature calculation function. Based on the hierarchical value Y, a quantitative value W for the social media account's operational data is determined, and the quantitative value W is output to the user in numerical or graphical form; wherein Based on the preset hierarchical spatial order and the feature value vector, a hierarchical feature calculation function is constructed corresponding to the categorized data set. The hierarchical feature calculation function is used to calculate the stratification values Y={Y1,Y2,Y3} of the categorized data set in three strata, including: The categorized dataset U is segmented to obtain a word vector [w1, w2, ..., wQ] containing Q words; Using word embedding matrix Transform sentence s into a vector sequence [e1,e2,…,eQ], where V and D represent the vocabulary size and word embedding dimension, respectively; Define the hierarchical value of ei Represented as: The stratification value Y1 of the categorized data set in the first stratum is represented as: , in These are the third and fourth adjustment parameters, respectively, and ReLU is a nonlinear activation function. Statistical analysis is performed on the sentences contained in the categorized dataset U to obtain R sentence vector values [x1, x2, ..., xR]. Then the stratification value Y2 of the categorized dataset in the second stratum is represented as: , in These are the fifth and sixth adjustment parameters, respectively. Statistical analysis is performed on the comment words contained in the categorized dataset U to obtain Z comment vector values [f1, f2, ..., fZ]. The stratification value Y3 of the categorized dataset in the third stratum is represented as: , Where τ is the seventh adjustment parameter, and f0 is the mean of the comment vector values [f1, f2, ..., fZ].
2. The method according to claim 1, characterized in that, The process involves collecting multi-dimensional data from the target social media account within a preset time interval [t1, t2], forming a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions, including: Determine the data query conditions corresponding to n dimensions; Using the data query conditions and the preset time period [t1,t2] as the data query request, multi-dimensional data of the target social media account within the preset time period [t1,t2] is collected.
3. The method according to claim 2, characterized in that, The step of collecting multi-dimensional data of the target social media account within a preset time interval [t1, t2] to form a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions also includes: Set the following metrics as data query requests: number of followers, net increase in followers in the past 7 days, maximum follower growth in the past 7 days, number of posts in the past 7 days, number of interactions per thousand followers in the past 7 days, total number of interactions per post in the past 7 days, maximum number of interactions per post in the past 7 days, usage of hashtags in posts in the past 7 days, and maximum number of interactions per hashtag post.
4. The method according to claim 3, characterized in that, After cleaning and preprocessing the data in the multi-dimensional dataset D using a standardized processing function, the cleaned and preprocessed data is then categorized to obtain a categorized dataset U={U1,U2,…,Um} containing m dimensions, and a feature vector K=[K1,K2,…,Km] corresponding to the categorized dataset, including: Perform outlier removal, missing data processing, and data format standardization on the data in the multidimensional dataset D.
5. The method according to claim 4, characterized in that, After cleaning and preprocessing the data in the multi-dimensional dataset D using a standardized processing function, the cleaned and preprocessed data is categorized to obtain a categorized dataset U={U1,U2,…,Um} containing m dimensions, and the corresponding feature vector K=[K1,K2,…,Km], further comprising: Based on the requirements of the classification task, a preset feature extraction method is selected to extract classification features related to the classification task from the cleaned and preprocessed data. The classification features include numerical features, text features, and image features. Based on the extracted classification features, a classifier for classification calculation is constructed, which includes the classification features. The trained classifier is used to classify the data obtained after cleaning and preprocessing, resulting in a data set U={U1,U2,…,Um} containing m dimensions, where each dimension represents a category of data. For the i-th class of categorized data Ui in the categorized dataset U, calculate the distance value Oi and the dispersion Si of categorized data Ui in the categorized dataset, so as to determine the feature value Ki corresponding to the categorized data Ui based on the distance value and the dispersion: γ and ρ are the first and second adjustment parameters, respectively.
6. A device for quantifying social media account operation data, characterized in that, include: The data acquisition module is used to acquire the target social media account and then collect multi-dimensional data of the target social media account within a preset time period [t1,t2] to form a multi-dimensional data set D={D1,D2,…,Dn} containing n dimensions. The classification module is used to clean and preprocess the data in the multi-dimensional data set D using a standardization processing function, and then perform classification calculations on the cleaned and preprocessed data to obtain a classification data set U={U1,U2,…,Um} containing m dimensions, and the feature vector K=[K1,K2,…,Km] corresponding to the classification data set, where m<n; The calculation module is used to construct the hierarchical feature calculation function corresponding to the classified data set based on the preset hierarchical spatial order and the feature value vector, and to calculate the hierarchical value Y={Y1,Y2,Y3} of the classified data set in three hierarchical layers through the hierarchical feature calculation function. The determination module, based on the hierarchical value Y, determines the quantized value W of the social media account's operational data, and outputs the quantized value W to the user in numerical or graphical form; wherein Based on the preset hierarchical spatial order and the feature value vector, a hierarchical feature calculation function is constructed corresponding to the categorized data set. The hierarchical feature calculation function is used to calculate the stratification values Y={Y1,Y2,Y3} of the categorized data set in three strata, including: The categorized dataset U is segmented to obtain a word vector [w1, w2, ..., wQ] containing Q words; Using word embedding matrix Transform sentence s into a vector sequence [e1,e2,…,eQ], where V and D represent the vocabulary size and word embedding dimension, respectively; Define the hierarchical value of ei Represented as: The stratification value Y1 of the categorized data set in the first stratum is represented as: , in These are the third and fourth adjustment parameters, respectively, and ReLU is a nonlinear activation function. Statistical analysis is performed on the sentences contained in the categorized dataset U to obtain R sentence vector values [x1, x2, ..., xR]. Then the stratification value Y2 of the categorized dataset in the second stratum is represented as: , in These are the fifth and sixth adjustment parameters, respectively. Statistical analysis is performed on the comment words contained in the categorized dataset U to obtain Z comment vector values [f1, f2, ..., fZ]. The stratification value Y3 of the categorized dataset in the third stratum is represented as: , Where τ is the seventh adjustment parameter, and f0 is the mean of the comment vector values [f1, f2, ..., fZ].
7. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method for quantifying social media account operation data as described in any one of claims 1-5.