An account classification method and device, electronic equipment and storage medium
By acquiring and analyzing account rating and interaction data, combined with information from reference accounts, and using machine learning to predict account types, the problem of low efficiency in self-media platform account classification has been solved, enabling efficient screening of high-quality accounts.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2021-03-31
- Publication Date
- 2026-05-22
AI Technical Summary
Existing technologies for classifying accounts on self-media platforms are inefficient and resource-intensive, making it difficult to efficiently select accounts with the potential to produce high-quality content.
By acquiring the level data, published content, and interaction data of candidate accounts, and combining them with the level data of associated reference accounts, machine learning is used to predict the probability that an account belongs to a preset type, thereby achieving automated classification.
It reduces manpower and material costs, improves the efficiency of account classification, and can quickly screen out accounts with the potential to produce high-quality content.
Smart Images

Figure CN115146136B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and specifically to an account classification method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of computer technology, multimedia applications are becoming increasingly widespread, allowing more users to post comments and create their own content on self-media platforms. For self-media platforms, high-quality content creators have significant commercial value, as their content can attract more users. Therefore, it is necessary to categorize accounts on the platform to select those with the potential to produce high-quality content.
[0003] Currently, the relevant technologies involve manually classifying accounts to determine whether they have the potential to produce high-quality content; however, this method requires a significant amount of manpower and resources and has relatively low classification efficiency when dealing with a massive number of accounts. Summary of the Invention
[0004] This application provides an account classification method, apparatus, electronic device, and storage medium, which can reduce manpower and material costs and improve the efficiency of account classification.
[0005] This application provides an account classification method, including:
[0006] Obtain at least one candidate account in the first application, the level data of the candidate account, the published content of the candidate account, and the interaction data, wherein the interaction data is the interaction data of the user account in the first application with respect to the published content;
[0007] Identify a reference account associated with the candidate account, wherein the reference account is a user account in the second application and the reference account belongs to the same user as the candidate account;
[0008] Based on the candidate account's level data, the published content, the interaction data, and the reference account's level data, predict the probability that the candidate account belongs to a preset account type;
[0009] Based on the probability, the candidate accounts are classified to obtain the target accounts of the preset account type.
[0010] Accordingly, embodiments of this application provide an account classification device, including:
[0011] The acquisition unit is used to acquire at least one candidate account in the first application, the level data of the candidate account, the published content of the candidate account, and the interaction data, wherein the interaction data is the interaction data of the user account in the first application with respect to the published content;
[0012] A determining unit is used to determine a reference account associated with the candidate account, wherein the reference account is a user account in the second application, and the reference account and the candidate account belong to the same user;
[0013] The prediction unit is used to predict the probability that the candidate account belongs to a preset account type based on the candidate account's level data, the published content, the interaction data, and the reference account's level data.
[0014] A classification unit is used to classify the candidate accounts according to the probability to obtain the target account of the preset account type.
[0015] Optionally, in some embodiments of this application, the prediction unit may include a ranking fusion subunit, a quality analysis subunit, an interaction analysis subunit, and a fusion subunit, as follows:
[0016] The grade fusion subunit is used to fuse the grade data of the candidate account and the grade data of the reference account to obtain the account popularity information of the candidate account.
[0017] The quality analysis subunit is used to perform quality analysis on the published content to obtain the content quality information of the published content.
[0018] The interaction analysis subunit is used to analyze the interaction data to obtain user attention information corresponding to the content published by the candidate account. The user attention information represents the attractiveness of the content published by the candidate account to users.
[0019] The fusion subunit is used to fuse the account popularity information, the content quality information, and the user attention information to obtain the probability that the candidate account belongs to a preset account type.
[0020] Optionally, in some embodiments of this application, the level data includes its own level information and fan count level information;
[0021] The level fusion subunit can be used to fuse the candidate account's own level information and the reference account's own level information to obtain the candidate account's own popularity information; fuse the candidate account's follower count level information and the reference account's follower count level information to obtain the candidate account's follower popularity information; and fuse the candidate account's own popularity information and the follower popularity information to obtain the candidate account's account popularity information.
[0022] Optionally, in some embodiments of this application, the quality analysis subunit may specifically be used to perform quantitative statistical processing on published content that meets preset quality conditions to obtain a first statistical result, and calculate the quality information of the published content based on the first statistical result; perform quantitative statistical processing on the original content in the published content to obtain a second statistical result, and calculate the originality information of the published content based on the second statistical result; calculate the professionalism information of the published content based on the category to which the published content belongs; and calculate the content quality information of the published content based on the quality information, the originality information, and the professionalism information.
[0023] Optionally, in some embodiments of this application, the interactive data includes interactive data of at least one interactive type;
[0024] Specifically, the interaction analysis subunit can be used to calculate the growth rate of interaction data for each interaction type within a preset time period; and to merge the growth rates of interaction data for each interaction type within the preset time period to obtain the user attention information corresponding to the content published by the candidate account.
[0025] Optionally, in some embodiments of this application, the step "merging the growth rates of interaction data for each interaction type within a preset time period to obtain user attention information corresponding to the content published by the candidate account" may include:
[0026] Obtain the weights corresponding to the interaction data for each interaction type;
[0027] Based on the weights, the growth rates of interaction data for each interaction type within a preset time period are weighted and fused to obtain the user attention information corresponding to the content published by the candidate accounts.
[0028] Optionally, in some embodiments of this application, the step "obtaining the weights corresponding to the interaction data of each interaction type" may include:
[0029] For each type of interaction data, determine the importance of the interaction data for that interaction type relative to the interaction data for a reference interaction type.
[0030] Based on the aforementioned importance, a comparison matrix of the interaction data for the aforementioned interaction type is constructed;
[0031] Based on the comparison matrix, the weights corresponding to the interactive data of the interactive type are determined.
[0032] Optionally, in some embodiments of this application, the step "determining the weight corresponding to the interaction data of the interaction type based on the comparison matrix" may include:
[0033] A consistency check is performed on the comparison matrix to obtain the consistency ratio of the comparison matrix;
[0034] When the consistency ratio is greater than a preset value, return to the operation of obtaining the importance of the interaction data of the interaction type relative to the interaction data of the reference interaction type;
[0035] The weights corresponding to the interactive data of the interaction type are determined based on a comparison matrix where the consistency ratio is not greater than a preset value.
[0036] An electronic device provided in this application includes a processor and a memory. The memory stores multiple instructions, and the processor loads the instructions to execute the steps in the account classification method provided in this application.
[0037] Furthermore, this application embodiment also provides a storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps in the account classification method provided in this application embodiment.
[0038] This application provides an account classification method, apparatus, electronic device, and storage medium. It can acquire at least one candidate account in a first application, the level data of the candidate account, the published content of the candidate account, and interaction data, wherein the interaction data is the interaction data of a user account in the first application with respect to the published content; determine a reference account associated with the candidate account, wherein the reference account is a user account in a second application, and the reference account belongs to the same user as the candidate account; predict the probability that the candidate account belongs to a preset account type based on the level data of the candidate account, the published content, the interaction data, and the level data of the reference account; and classify the candidate account according to the probability to obtain a target account of the preset account type. This application eliminates the need for manual annotation of account sample data, reducing manpower and material costs and improving the efficiency of account classification. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1a This is a schematic diagram of a scenario illustrating the account classification method provided in an embodiment of this application;
[0041] Figure 1b This is a flowchart of the account classification method provided in the embodiments of this application;
[0042] Figure 1c This is an explanatory diagram of the account classification method provided in the embodiments of this application;
[0043] Figure 1d This is another flowchart of the account classification method provided in the embodiments of this application;
[0044] Figure 2a This is another flowchart of the account classification method provided in the embodiments of this application;
[0045] Figure 2b This is a schematic diagram of the architecture of the account classification method provided in the embodiments of this application;
[0046] Figure 2c This is another flowchart of the account classification method provided in the embodiments of this application;
[0047] Figure 2d This is another flowchart of the account classification method provided in the embodiments of this application;
[0048] Figure 2e This is another flowchart of the account classification method provided in the embodiments of this application;
[0049] Figure 3a This is a schematic diagram of the account classification device provided in the embodiments of this application;
[0050] Figure 3b This is another structural schematic diagram of the account classification device provided in the embodiments of this application;
[0051] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0053] This application provides an account classification method, apparatus, electronic device, and storage medium. Specifically, the account classification apparatus can be integrated into an electronic device, which may be a terminal or server, etc.
[0054] It is understood that the account classification method in this embodiment can be executed on a terminal, on a server, or jointly by a terminal and a server. The above examples should not be construed as limiting this application.
[0055] like Figure 1a As shown, the account classification method is implemented jointly by a terminal and a server as an example. The account classification system provided in this application includes a terminal 10 and a server 11, etc.; the terminal 10 and the server 11 are connected via a network, such as a wired or wireless network, etc., wherein the account classification device can be integrated into the server.
[0056] Server 11 can be used to: acquire at least one candidate account in a first application, the level data of the candidate account, the published content of the candidate account, and interaction data, wherein the interaction data is the interaction data of a user account in the first application with respect to the published content; determine a reference account associated with the candidate account, wherein the reference account is a user account in a second application, and the reference account belongs to the same user as the candidate account; predict the probability that the candidate account belongs to a preset account type based on the level data of the candidate account, the published content, the interaction data, and the level data of the reference account; and classify the candidate account according to the probability to obtain a target account of the preset account type. Server 11 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. In the account classification method or apparatus disclosed in this application, multiple servers can form a blockchain, and the servers are nodes on the blockchain.
[0057] Terminal 10 can receive target accounts of a preset account type sent by server 11 and display the content published by the target accounts on the content recommendation page. Terminal 10 can include smartphones, smart TVs, smart speakers, smartwatches, tablets, laptops, or personal computers (PCs), etc. A client can also be set on terminal 10, which can be an application client or a browser client, etc.
[0058] The step of classifying candidate accounts by the aforementioned server 11 can also be performed by the terminal 10.
[0059] The account classification method provided in this application relates to machine learning in the field of artificial intelligence. This method eliminates the need for manual labeling of account sample data, reducing manpower and material costs and improving the efficiency of account classification.
[0060] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. AI technology is a comprehensive discipline involving a wide range of fields, encompassing both hardware and software technologies. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0061] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0062] The account classification method provided in this application also relates to the big data field in cloud technology.
[0063] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Cloud technology is a general term encompassing network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can only be achieved through cloud computing.
[0064] Big data refers to data sets that cannot be captured, managed, and processed within a certain timeframe using conventional software tools. It represents massive, rapidly growing, and diverse information assets that require new processing models to achieve stronger decision-making, insightful discovery, and process optimization capabilities. With the advent of the cloud era, big data has attracted increasing attention. Big data requires specialized technologies to effectively process large amounts of data within a tolerable timeframe. Technologies suitable for big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the internet, and scalable storage systems.
[0065] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.
[0066] This embodiment will be described from the perspective of an account classification device, which can be integrated into an electronic device, such as a server or a terminal.
[0067] The account classification method of this application can be applied to various scenarios where account classification is required. For example, a video platform needs to select accounts with the potential to produce high-quality content from millions of accounts. The account classification method provided in this embodiment can quickly classify massive numbers of accounts, filter out potential accounts, and save manpower and material costs.
[0068] like Figure 1b As shown, the specific process of this account classification method can be as follows:
[0069] 101. Obtain at least one candidate account in the first application, the level data of the candidate account, the published content of the candidate account, and the interaction data, wherein the interaction data is the interaction data of the user account in the first application with respect to the published content.
[0070] The first application can be a platform or application for content publishing and browsing, and its type is not limited. This first application platform contains a large number of user accounts (i.e., content producers) that produce content. Some user accounts produce low-quality content, while others produce high-quality content. The first application needs to select accounts of a preset account type from this large pool of user accounts. This preset account type can be accounts with growth potential, whose content is of relatively high quality and can attract more user interaction. The first application can cultivate and promote these types of accounts to realize significant commercial value. The produced content can include videos, images, articles, music, etc., and this embodiment does not impose any limitations on this.
[0071] The candidate account's rating data can represent its popularity. This rating data can include the account's own rating and follower count rating. The account's own rating specifically refers to its rating within the first application. The follower count rating can be the rating corresponding to the number of followers. For example, an account with more than 100,000 followers is rated as Level 1, an account with 10,000 to 100,000 followers is rated as Level 2, and an account with 1,000 to 10,000 followers is rated as Level 3. Optionally, in some embodiments, the account's rating can be related to its activity level, specifically the frequency of content posting. Higher posting frequency leads to higher account activity and a higher rating, while lower posting frequency leads to lower account activity and a lower rating.
[0072] The types of content that candidate accounts can publish are not limited; they can include videos, audio, text, images, and so on.
[0073] The interactive data refers to the interaction data between user accounts (including the candidate account itself and other user accounts besides the candidate account) in the first application and the content published by the candidate account. There are various types of interaction data, and this embodiment does not limit them. For example, interactive data may include sharing, liking, commenting, collecting, bullet comments, number of views, and user browsing time, etc.
[0074] In some embodiments, candidate accounts can be obtained by initially screening all accounts in the first application. There are various methods for initial screening, and this embodiment does not limit this one. For example, accounts in the first application with more than a preset value of followers can be identified as candidate accounts. This preset value can be determined according to the actual situation.
[0075] 102. Determine a reference account associated with the candidate account, wherein the reference account is a user account in the second application, and the reference account and the candidate account belong to the same user.
[0076] The second application can be an application of a different type from the first application, or it can be an application of the same type as the first application. This embodiment does not impose any restrictions on this.
[0077] In this embodiment, for each candidate account in the first application, platform data of the corresponding account (i.e., reference account) on other platforms (i.e., the second application) can be obtained. Based on this platform data, the quality of the content produced by the reference account can be judged. If the quality of the content produced by the reference account is high, it indicates that the candidate account associated with the reference account is more likely to become a high-quality content producer in the first application, i.e., it may have great potential. Therefore, when screening candidate accounts, the first application can comprehensively consider the platform data of the associated reference account on other platforms, as well as the platform data of the candidate account on this platform (i.e., the first application), to determine whether the candidate account is an account of a preset account type, so as to further determine the cultivation strategy for the candidate account.
[0078] 103. Based on the candidate account's level data, the published content, the interaction data, and the reference account's level data, predict the probability that the candidate account belongs to a preset account type.
[0079] Among them, the reference account's level data can represent the popularity information of the reference account in the second application. The reference account's level data can include the reference account's own level information and the number of followers level information.
[0080] Optionally, in this embodiment, the step "predicting the probability that the candidate account belongs to a preset account type based on the candidate account's level data, the published content, the interaction data, and the reference account's level data" may include:
[0081] The level data of the candidate accounts and the level data of the reference accounts are merged to obtain the account popularity information of the candidate accounts;
[0082] The published content is subjected to quality analysis to obtain content quality information.
[0083] The interaction data is analyzed to obtain user attention information corresponding to the content published by the candidate accounts. The user attention information represents the attractiveness of the content published by the candidate accounts to users.
[0084] The probability that a candidate account belongs to a preset account type is obtained by fusing the account popularity information, the content quality information, and the user attention information.
[0085] Specifically, account popularity information, content quality information, and user follower information are the primary indicators for predicting the account type of candidate accounts. Optionally, in this embodiment, the probability of a candidate account belonging to a preset account type can be comprehensively considered based on these three dimensions: account popularity information, content quality information, and user follower information.
[0086] Account popularity information can be determined based on the candidate account's rating data in the first application and the rating data of its associated reference account in the second application. Account popularity information represents the candidate account's basic potential.
[0087] The step "integrating the account popularity information, the content quality information, and the user follow information to obtain the probability that the candidate account belongs to a preset account type" may include:
[0088] Determine the weights corresponding to the account popularity information, the content quality information, and the user attention information;
[0089] Based on the weights, the account popularity information, the content quality information, and the user attention information are weighted and calculated to obtain the probability that the candidate account belongs to a preset account type.
[0090] The Analytic Hierarchy Process (AHP) can be used to determine the weights of account popularity, content quality, and user attention information. Specifically, the importance of each primary indicator relative to the others can be determined first. A comparison matrix for each primary indicator can then be constructed based on this importance. A consistency check can be performed on the comparison matrix. If the comparison matrix passes the consistency check, the weight of that primary indicator can be determined based on the comparison matrix.
[0091] Specifically, account popularity, content quality, and user engagement information can be quantified, and their probabilities can be represented by a quantified score. Finally, candidate accounts can be ranked based on the probability (quantified score) of belonging to a preset account type. In some embodiments, multiple score ranges can be defined, each corresponding to a different cultivation strategy. The target cultivation strategy for a candidate account is determined based on its quantified score, with higher scores indicating greater cultivation efforts.
[0092] In a specific embodiment, account popularity information can be denoted as S1, content quality information as S2, and user attention information as S3. The weights corresponding to account popularity information, content quality information, and user attention information are denoted as α, β, and γ, respectively. Then, the probability (quantitative score) S of a candidate account belonging to a preset account type can be calculated as shown in formula (1):
[0093] S=α*S1+β*S2+γ*S3(1)
[0094] The values of α, β, and γ can be set according to actual conditions, and this embodiment does not impose any restrictions on them. For example, α can be set to 0.2, β to 0.4, and γ to 0.4.
[0095] Optionally, in this embodiment, the level data includes its own level information and fan count level information;
[0096] The step "merging the level data of the candidate accounts and the level data of the reference accounts to obtain the account popularity information of the candidate accounts" may include:
[0097] The self-level information of the candidate account and the self-level information of the reference account are fused to obtain the self-popularity information of the candidate account.
[0098] The follower count level information of the candidate account and the follower count level information of the reference account are merged to obtain the follower popularity information of the candidate account.
[0099] The self-popularity information and the fan popularity information are merged to obtain the account popularity information of the candidate account.
[0100] The step of "merging the self-level information of the candidate account and the self-level information of the reference account to obtain the self-popularity information of the candidate account" may include:
[0101] Determine the weights corresponding to the candidate accounts and the reference accounts;
[0102] Based on the weights, the self-level information of the candidate account and the self-level information of the reference account are weighted and calculated to obtain the self-popularity information of the candidate account.
[0103] In some embodiments, assuming the application platform has five levels from 1 to 5, the level of the candidate account and reference account in the application platform can be converted into a percentage system. Then, the level information S of the candidate account or reference account... level =level*(100 / 5).
[0104] The step of "merging the follower count level information of the candidate account and the follower count level information of the reference account to obtain the follower popularity information of the candidate account" may include:
[0105] Determine the weights corresponding to the candidate accounts and the reference accounts;
[0106] Based on the weights, the follower count level information of the candidate account and the follower count level information of the reference account are weighted and calculated to obtain the follower popularity information of the candidate account.
[0107] The weights corresponding to candidate accounts and reference accounts specifically refer to the weights corresponding to the first application and the second application. In this embodiment, a candidate account can be associated with one or more reference accounts, and the candidate account and all the reference accounts belong to the same user. If a candidate account corresponds to n-1 reference accounts, then the user has joined one first application and n-1 second applications. For these n application platforms, the weight 'a' corresponding to each application platform can be determined according to the different application types. Optionally, in a specific embodiment, the weight of each application platform can be set to... .
[0108] The follower count level information of candidate or reference accounts can be denoted as S. fans Let Sp be the popularity information of the candidate account itself, and Sf be the popularity information of the candidate account's followers. The calculation process of the popularity information S1 of the candidate account can be shown in formula (2):
[0109] (2)
[0110] In one specific embodiment, the follower count level information of the candidate account or reference account can be denoted as S. fans The calculation of the number of followers for each account can be shown in equations (3) and (4):
[0111] (3)
[0112] (4)
[0113] in, This represents the number of followers of account i at the current moment, where i is a positive integer not greater than n, and n represents the number of application platforms. Specifically, for... Perform logarithmic transformation The goal is to transform a right-skewed distribution into a symmetrical distribution, thereby reducing the impact of orders-of-magnitude differences.
[0114] In this embodiment, the step "to perform quality analysis on the published content and obtain content quality information of the published content" may include:
[0115] The published content is subjected to quality analysis in multiple dimensions to obtain the quality analysis results of the published content in the multiple dimensions.
[0116] The quality analysis results of the published content across various dimensions are integrated to obtain the content quality information of the published content.
[0117] The dimensions of quality analysis may include information on the quality, originality, and professionalism of the published content, as well as various negative indicators (such as low quality rate, report rate, and negative feedback ratio), etc. This embodiment does not impose any restrictions on these.
[0118] Specifically, the quality information can be the proportion of high-quality content among the content published within a preset time period. A higher proportion indicates higher content quality. The preset time period can be one month, but this embodiment does not limit this. In a specific scenario, published content can be scored, with a score range of 1-5. Content scoring 4-5 can be considered high-quality content. Optionally, the total number of contents published by the candidate account within the preset time period can be denoted as T, where the number of high-quality contents published by the candidate account within that preset time period is denoted as Tq, and the quality information of the candidate account's published content is denoted as Sq. Then, Sq = Tq / T * 100.
[0119] The originality information can be the proportion of original content among the content published within a preset time period. The higher the proportion, the higher the originality and the greater the growth potential of the candidate account. The preset time period can be one month. Optionally, the total number of contents published by the candidate account within the preset time period can be denoted as T, where the number of original contents published by the candidate account within the preset time period is denoted as To, and the originality information of the candidate account's published content is denoted as So. Then, So = To / T * 100.
[0120] Among them, professionalism information, also known as verticality information, can be used to represent the domain focus of a user account and measure the stability of the category of content published by the account. The category to which the published content belongs can include finance, sports, technology, entertainment, and society, etc. The more times the published content belongs to a certain category, the higher the domain focus of the user account in that category; conversely, the fewer the published content belongs to a certain category, the lower the domain focus of the user account in that category. In this embodiment, for the professionalism information of candidate accounts, specifically, candidate accounts with a relatively concentrated category of published content can be considered as accounts with higher professionalism information, which have higher growth potential. Optionally, the professionalism information of the published content of a candidate account can be denoted as Hp, and its calculation process is shown in formula (5):
[0121] (5)
[0122] Where n represents the number of categories of content published by candidate accounts within a preset time period, i.e., n represents the total number of vertical categories of content published by candidate accounts within a preset time period, and i represents the i-th vertical category. This represents the proportion of content published in the i-th vertical category. Specifically, the calculation method for Hp is to obtain the content published by candidate accounts within a preset time period (such as one month), use the primary classification results of the published content to determine the category (i.e., vertical category) to which the published content belongs, and then statistically analyze the professionalism information of the candidate accounts.
[0123] Specifically, the low-quality rate can be the proportion of low-quality content among the content published within a preset time period. The higher the proportion, the worse the content quality. Low-quality content can be obtained through manual or machine review, and the preset time period can be one month. Optionally, the total number of contents published by the candidate account within the preset time period can be denoted as T, where the number of low-quality contents published by the candidate account within the preset time period is denoted as Tl, and the low-quality rate of the candidate account's published content is denoted as Sl. Then, Sl = Tl / T.
[0124] The reporting rate can be the proportion of reported instances of content published within a preset time period to the total number of consumptions. A higher proportion indicates lower content quality. Consumptions can refer to browsing, saving, etc. Optionally, the number of consumptions of content published by a candidate account within a preset time period is denoted as Tx, the number of reports of content published by the candidate account within the preset time period is denoted as Tr, and the reporting rate of the candidate account's published content is denoted as Sr. Then, Sr = Tr / Tx.
[0125] The negative feedback ratio can be the proportion of negative feedback instances among the content published within a preset time period to the total number of consumptions. A higher ratio indicates lower content quality. Optionally, the number of consumptions of content published by a candidate account within a preset time period is denoted as Tx, the number of negative feedback instances of content published by the candidate account within the preset time period is denoted as Tn, and the negative feedback ratio of the content published by the candidate account is denoted as Sn. Then, Sn = Tn / Tx.
[0126] The step "integrating the quality analysis results of the published content across various dimensions to obtain the content quality information of the published content" may include:
[0127] Determine the weights corresponding to the quality analysis results of the published content in each dimension;
[0128] Based on the weights, the quality analysis results of the published content in each dimension are weighted and calculated to obtain the content quality information of the published content.
[0129] The quality analysis results for each dimension can be information on high quality, originality, professionalism, low quality rate, reporting rate, and negative feedback ratio, with corresponding weights denoted as a1, a2, a3, a4, a5, and a6, respectively. The content quality information of the published content is denoted as S2. Then, S2 = a1*Sq + a2*So + a3*Hp + a4*(1 / Sl) + a5*(1 / Sr) + a6*(1 / Sn). The weights can be set according to the actual situation.
[0130] Optionally, in this embodiment, the step "performing quality analysis on the published content to obtain content quality information of the published content" may include:
[0131] The published content that meets the preset quality conditions is subjected to quantity statistical processing to obtain a first statistical result, and the quality information of the published content is calculated based on the first statistical result.
[0132] The original content in the published content is statistically analyzed to obtain a second statistical result, and the originality information of the published content is calculated based on the second statistical result.
[0133] Based on the category to which the published content belongs, calculate the professionalism information of the published content;
[0134] Based on the quality information, the originality information, and the professionalism information, the content quality information of the published content is calculated.
[0135] Specifically, the preset quality conditions can be that the content meets the standards for high-quality content. In one specific embodiment, published content can be scored from 1 to 5 points. Content scoring 4 to 5 points can be considered high-quality content. The number of high-quality published content is counted, and this number is the first statistical result. Based on the first statistical result and the total number of published content, the quality information of the published content can be calculated. Similarly, the originality information of the published content can be calculated based on the second statistical result and the total number of published content.
[0136] The calculation process for the professionalism information can be referred to the description in the above embodiments.
[0137] Optionally, in this embodiment, the interactive data includes interactive data of at least one interactive type;
[0138] The step "analyzing the interaction data to obtain user follow information corresponding to the content published by the candidate accounts" may include:
[0139] For each type of interaction data, calculate the growth rate of the interaction data for that type of interaction within a preset time period;
[0140] By merging the growth rates of interaction data for each interaction type within a preset time period, user attention information corresponding to the content published by the candidate account is obtained.
[0141] Interactive data can be categorized into various types, such as sharing, liking, commenting, favorites, bullet comments, number of views, and user browsing time.
[0142] The step "calculating the growth rate of interaction data of the interaction type within a preset time period" may include:
[0143] Divide the preset time period into at least one sub-time period;
[0144] Calculate the growth rate of interaction data for the aforementioned interaction type within each sub-time period;
[0145] The growth rates of each sub-time period are merged to obtain the growth rate of the interaction data of the interaction type within the preset time period.
[0146] Optionally, in this embodiment, the step "merging the growth rates of interaction data for each interaction type within a preset time period to obtain user attention information corresponding to the content published by the candidate account" may include:
[0147] Obtain the weights corresponding to the interaction data for each interaction type;
[0148] Based on the weights, the growth rates of interaction data for each interaction type within a preset time period are weighted and fused to obtain the user attention information corresponding to the content published by the candidate accounts.
[0149] For example, interaction types include likes, comments, shares, and consumption time. The growth rates of these four within a preset time period are denoted as Pz, Pc, Pf, and Pt, respectively, and their corresponding weights are denoted as q1, q2, q3, and q4. Then, the user attention information S3 corresponding to the content published by the candidate account is calculated as S3 = q1*Pz + q2*Pc + q3*Pf + q4*Pt. The weights q1, q2, q3, and q4 can be set according to the actual situation. This embodiment does not impose any restrictions on this. For example, q1 can be set to 0.15, q2 can be set to 0.25, q3 can be set to 0.25, and q4 can be set to 0.35.
[0150] Optionally, in one embodiment, the preset time period is divided into D sub-time periods, and the growth rate of interactive data of a certain type within the preset time period is denoted as P. The calculation process can be shown in formula (6):
[0151] (6)
[0152] in, This indicates the number of interactive data points corresponding to the rightmost time point on the timeline within the sub-time period. This indicates the number of interactive data points corresponding to the leftmost time point on the timeline within the sub-time period. This represents the rightmost time point on the timeline within a sub-time period. P represents the leftmost point of time on the timeline within a sub-time period. Specifically, P can represent dividing the interaction data of a candidate account for a certain interaction type within a preset time period into several segments according to time, and calculating the average growth rate of the interaction data of that interaction type within each sub-time period.
[0153] Weight refers to the degree of importance of a factor or indicator relative to the whole; the weights of interactive data for each type of interaction can be determined using the analytic hierarchy process (AHP).
[0154] The Analytic Hierarchy Process (AHP) is a systematic and hierarchical analytical method that combines qualitative and quantitative approaches. It can mathematically represent the decision-making process with limited quantitative information, thus providing a convenient decision-making method for complex decision problems with multiple objectives, multiple criteria, or unstructured characteristics. It is a model and method for making decisions on complex systems that are difficult to quantify completely. According to the nature of the problem and the overall goal to be achieved, the AHP decomposes the problem into different components and groups them into different levels according to the interrelationships and hierarchical relationships between the components, forming a multi-level analytical structure model. Ultimately, the problem is reduced to determining the relative importance weights or relative superiority / inferiority of the lowest level (solutions, measures, etc. for decision-making) relative to the highest level (overall goal). The AHP can be divided into the following four steps: (1) establishing a hierarchical structure model; (2) constructing a judgment (pairwise comparison) matrix; (3) hierarchical single ranking and its consistency test; (4) hierarchical overall ranking and its consistency test.
[0155] Optionally, in this embodiment, the step "obtaining the weights corresponding to the interaction data of each interaction type" may include:
[0156] For each type of interaction data, determine the importance of the interaction data for that interaction type relative to the interaction data for a reference interaction type.
[0157] Based on the aforementioned importance, a comparison matrix of the interaction data for the aforementioned interaction type is constructed;
[0158] Based on the comparison matrix, the weights corresponding to the interactive data of the interactive type are determined.
[0159] Specifically, the reference interaction type can be other interaction types. Optionally, in some embodiments, the reference interaction type can be all interaction types, that is, the reference interaction type includes its own interaction type.
[0160] Specifically, "the importance of the interaction data of the interaction type relative to the interaction data of the reference interaction type" can measure the relative importance of the interaction data of the interaction type (factor A) to account classification compared to the interaction data of the reference interaction type (factor B). That is, it is a comparison of the influence of the two factors on the correct prediction result. For example, it can be a comparison between the importance of the interaction data of the interaction type to the account classification result and the importance of the interaction data of the reference interaction type to the account classification result (such as a ratio).
[0161] For example, if the interaction type is a comment and the reference interaction type is a like, then "the importance of the interaction data of the interaction type relative to the interaction data of the reference interaction type" can specifically be: the importance of comment interaction data relative to like interaction data in predicting the account type.
[0162] Optionally, in this embodiment, the step "determining the weight corresponding to the interaction data of the interaction type based on the comparison matrix" may include:
[0163] A consistency check is performed on the comparison matrix to obtain the consistency ratio of the comparison matrix;
[0164] When the consistency ratio is greater than a preset value, return to the operation of obtaining the importance of the interaction data of the interaction type relative to the interaction data of the reference interaction type;
[0165] The weights corresponding to the interactive data of the interaction type are determined based on a comparison matrix where the consistency ratio is not greater than a preset value.
[0166] The consistency test refers to testing the averages or variances calculated from different samples. This preset value can be set according to actual conditions; this embodiment does not impose any restrictions on it, such as setting it to 0.1.
[0167] Specifically, when the consistency ratio is not greater than the preset value, the comparison matrix passes the consistency test; when the consistency ratio is greater than the preset value, the consistency test fails, and the comparison matrix needs to be reconstructed until the consistency ratio of the newly constructed comparison matrix is not greater than the preset value.
[0168] It should be noted that the weights involved in the above embodiments can all be determined using the analytic hierarchy process. For example, the process of determining the weights of the quality analysis results in each dimension can be referred to the process of determining the weights of the interactive data above, which will not be repeated here.
[0169] 104. Based on the probability, classify the candidate accounts to obtain the target accounts of the preset account type.
[0170] Among them, the target accounts of the preset account type are accounts with growth potential. They are neither high-quality accounts nor low-quality accounts. They can be considered as accounts that are in the middle of the ranking in the account system, but they are the fastest improving accounts and may eventually grow into platform-featured or top accounts.
[0171] Specifically, the probability can be a quantified score. Optionally, in some embodiments, it can be further transformed to obtain the total potential score of the candidate account, such as... Figure 1c As shown, curve 1 represents the probability of a candidate account before conversion (i.e., the quantified score), and curve 2 represents the mapping relationship between the candidate account's total potential score (y) and quantified score (x). Specifically, the total potential score is obtained by taking the square root of the quantified score and multiplying it by 10. Both the quantified score and the total potential score range from 0 to 100. By converting the quantified score, candidate accounts with lower scores can receive more compensation, encouraging account growth, without changing the overall ranking of the candidate accounts.
[0172] The step "classifying the candidate accounts according to the probability to obtain the target accounts of the preset account type" may specifically include: identifying candidate accounts with a probability greater than a preset probability threshold as target accounts of the preset account type. This preset probability threshold can be set according to actual circumstances.
[0173] Optionally, in this embodiment, different potential score ranges can be set for different potential levels. For example, candidate accounts with a total potential score of 80-100 are in the first potential level, candidate accounts with a total potential score of 60-80 are in the second potential level, and candidate accounts with a total potential score below 60 are in the third potential level. Different strategies are used for candidate accounts with different potential levels. Specifically, the potential level of candidate accounts can be used for review scheduling and ranking, with accounts with lower potential levels placed at the end of the review queue.
[0174] Alternatively, in one embodiment, as Figure 1dThe flowchart shown illustrates the process of selecting target accounts with growth potential. First, we can determine the primary indicators for judging the potential of candidate accounts. These primary indicators can be divided into basic potential (specifically, the platform's ranking data), prior potential (specifically, the quality of the account's published content), growth potential (specifically, the interaction data of the published content), and other supplementary information (such as negative indicators in the content quality). These primary indicators are considered comprehensively because accounts with high growth potential generally have a large and rapidly growing base of followers. Furthermore, the corresponding reference accounts on other platforms are also of relatively high quality. Accounts with high growth potential typically publish high-quality, original, and vertically focused content with fewer negative indicators, exhibit rapid growth in interaction data related to their published content, and make a significant contribution to the platform.
[0175] Specifically, as shown in Figure 1d, for each candidate account, the basic potential of the candidate account can be determined based on the number of followers of the candidate account on this platform 1 (the first application) and the number of followers of the reference accounts associated with it on other application platforms (i.e., the second application, such as platform a, platform b, and platform c). The prior potential of the candidate account can be determined based on the quality, originality, and verticality information of the content published by the candidate account. The growth potential of the candidate account can be determined by comprehensively considering the growth rate of likes, comments, shares, and duration of its published content. In addition, the negative information of the candidate account's published content can be determined based on the low quality rate, report rate, and negative feedback ratio of the candidate account's published content. Finally, the total potential score of the candidate account can be determined based on the basic potential, prior potential, growth potential, and negative information.
[0176] The account classification system in this application can update and refresh the potential level ranking of accounts at any time based on changes in data.
[0177] The account classification method proposed in this application can select target accounts with growth potential based on account level data, published content, and interaction data. This allows the content of target accounts to be prioritized for distribution in the recommendation pool, significantly reducing the turnover rate of potential accounts. At the same time, it can concentrate traffic on truly high-quality and promising content creators, reducing the waste of platform traffic. Authors who are truly needed by the platform can receive the greatest incentive, which is more conducive to long-term development and reputation. Furthermore, the unsupervised modeling method does not require manual annotation of account sample data, reducing costs and improving the timeliness of potential account discovery and processing, attracting more potential accounts to join and grow rapidly.
[0178] As can be seen from the above, this embodiment can obtain at least one candidate account in a first application, the level data of the candidate account, the published content of the candidate account, and interaction data, wherein the interaction data is the interaction data of a user account in the first application with respect to the published content; determine a reference account associated with the candidate account, wherein the reference account is a user account in a second application, and the reference account and the candidate account belong to the same user; predict the probability that the candidate account belongs to a preset account type based on the level data of the candidate account, the published content, the interaction data, and the level data of the reference account; classify the candidate account according to the probability to obtain the target account of the preset account type. This application does not require manual annotation of account sample data, which can reduce manpower and material costs and improve the efficiency of account classification.
[0179] Based on the method described in the preceding embodiments, the following will provide a more detailed explanation by taking the specific integration of the account classification device into the server as an example.
[0180] This application provides an account classification method, such as... Figure 2a As shown, the specific process of this account classification method can be as follows:
[0181] 201. The server obtains at least one candidate account in the first application, the level data of the candidate account, the published content of the candidate account, and the interaction data, wherein the interaction data is the interaction data of the user account in the first application with respect to the published content.
[0182] Among them, the candidate account's level data can represent the candidate account's popularity information. The level data can include its own level information and the number of followers level information. The own level information is specifically the user account's own account level in the first application, and the number of followers level information is specifically the level corresponding to the number of followers of the user account.
[0183] The types of content that candidate accounts can publish are not limited; they can include videos, audio, text, images, and so on.
[0184] The interactive data refers to the interaction data between user accounts (including the candidate account itself and other user accounts besides the candidate account) in the first application and the content published by the candidate account. There are various types of interaction data, and this embodiment does not limit them. For example, interactive data may include sharing, liking, commenting, collecting, bullet comments, number of views, and user browsing time, etc.
[0185] 202. The server determines a reference account associated with the candidate account, wherein the reference account is a user account in the second application, and the reference account and the candidate account belong to the same user.
[0186] The second application can be an application of a different type from the first application, or it can be an application of the same type as the first application. This embodiment does not impose any restrictions on this.
[0187] 203. The server merges the level data of the candidate account and the level data of the reference account to obtain the account popularity information of the candidate account.
[0188] The step "merging the level data of the candidate accounts and the level data of the reference accounts to obtain the account popularity information of the candidate accounts" may include:
[0189] The self-level information of the candidate account and the self-level information of the reference account are fused to obtain the self-popularity information of the candidate account.
[0190] The follower count level information of the candidate account and the follower count level information of the reference account are merged to obtain the follower popularity information of the candidate account.
[0191] The self-popularity information and the fan popularity information are merged to obtain the account popularity information of the candidate account.
[0192] 204. The server performs quality analysis on the published content to obtain content quality information of the published content.
[0193] In this embodiment, the step "to perform quality analysis on the published content and obtain content quality information of the published content" may include:
[0194] The published content is subjected to quality analysis in multiple dimensions to obtain the quality analysis results of the published content in the multiple dimensions.
[0195] The quality analysis results of the published content across various dimensions are integrated to obtain the content quality information of the published content.
[0196] The dimensions of quality analysis may include information on the quality, originality, and professionalism of the published content, as well as various negative indicators (such as low quality rate, report rate, and negative feedback ratio), etc. This embodiment does not impose any restrictions on these.
[0197] Optionally, in this embodiment, the step "performing quality analysis on the published content to obtain content quality information of the published content" may include:
[0198] The published content that meets the preset quality conditions is subjected to quantity statistical processing to obtain a first statistical result, and the quality information of the published content is calculated based on the first statistical result.
[0199] The original content in the published content is statistically analyzed to obtain a second statistical result, and the originality information of the published content is calculated based on the second statistical result.
[0200] Based on the category to which the published content belongs, calculate the professionalism information of the published content;
[0201] Based on the quality information, the originality information, and the professionalism information, the content quality information of the published content is calculated.
[0202] 205. The server analyzes the interaction data to obtain user attention information corresponding to the content published by the candidate account. The user attention information represents the attractiveness of the content published by the candidate account to users.
[0203] Optionally, in this embodiment, the interactive data includes interactive data of at least one interactive type;
[0204] The step "analyzing the interaction data to obtain user follow information corresponding to the content published by the candidate accounts" may include:
[0205] For each type of interaction data, calculate the growth rate of the interaction data for that type of interaction within a preset time period;
[0206] By merging the growth rates of interaction data for each interaction type within a preset time period, user attention information corresponding to the content published by the candidate account is obtained.
[0207] Interactive data can be categorized into various types, such as sharing, liking, commenting, favorites, bullet comments, number of views, and user browsing time.
[0208] The step "calculating the growth rate of interaction data of the interaction type within a preset time period" may include:
[0209] Divide the preset time period into at least one sub-time period;
[0210] Calculate the growth rate of interaction data for the aforementioned interaction type within each sub-time period;
[0211] The growth rates of each sub-time period are merged to obtain the growth rate of the interaction data of the interaction type within the preset time period.
[0212] 206. The server integrates the account popularity information, the content quality information, and the user attention information to obtain the probability that the candidate account belongs to a preset account type.
[0213] The step "integrating the account popularity information, the content quality information, and the user follow information to obtain the probability that the candidate account belongs to a preset account type" may include:
[0214] Determine the weights corresponding to the account popularity information, the content quality information, and the user attention information;
[0215] Based on the weights, the account popularity information, the content quality information, and the user attention information are weighted and calculated to obtain the probability that the candidate account belongs to a preset account type.
[0216] 207. The server classifies the candidate accounts according to the probability to obtain the target accounts of the preset account type.
[0217] The step "classifying the candidate accounts according to the probability to obtain the target accounts of the preset account type" may specifically include: identifying candidate accounts with a probability greater than a preset probability threshold as target accounts of the preset account type. This preset probability threshold can be set according to actual circumstances.
[0218] Based on the methods described in the above embodiments, the following examples will provide further detailed explanations.
[0219] This embodiment uses an account classification device integrated server as an example, which is a cluster server. This cluster server may include uplink / downlink content interface servers, content storage servers, a scheduling center server, a deduplication server, an auditing server, a recommendation and distribution server, a content distribution exit server, a statistical reporting interface server, a statistical data storage server, as well as servers for potential account identification and potential account feature mining models. The connection relationships between the servers in the cluster server can be as follows: Figure 2b As shown, the cluster server can communicate with the content producer through the uplink / downlink content interface server, and with the content consumer through the uplink / downlink content interface server, the statistics reporting interface server, or the content distribution export server. The content producer can be a client that produces content to be published, and the content consumer can be a client that receives and displays the content to be published pushed by the cluster server. There can be one or more content producers and one or more content consumers.
[0220] Please see Figure 2c , Figure 2c This is a flowchart illustrating the account classification method provided in an embodiment of this application. The method flow may include:
[0221] S10. The upstream and downstream content interface server receives the content to be published uploaded by the content producer.
[0222] Content producers can generate content to be published through user accounts associated with professionally generated content (PGC), user-generated content (UGC), multi-channel networks (MCN), or professionally generated user content (PUGC). For example, they can use the content producer (mobile app) or backend application programming interface (API) to upload text and image content or video content (including short videos and micro-videos) provided by a local or web (World Wide Web) publishing system, awaiting publication. The content producer can establish a communication connection with the upstream and downstream content interface servers, obtain their server interface addresses, and then send the content to be published to the upstream and downstream content interface servers based on those addresses. The upstream and downstream content interface servers then receive the content uploaded by the content producer.
[0223] S11. The upstream and downstream content interface servers write the content to be published and metadata to the content storage server.
[0224] Metadata for all content published by content producers can be stored in a content storage server (i.e., a content database). During content review (which may include manual review), information can be retrieved from the content storage server, and the review results and status can also be sent back to the content storage server for storage. The content storage server can also store the processing results from deduplication servers and potential account identification servers, etc.
[0225] Meta information can include content size, cover image link, title, publication time, account author, source channel, and entry time (i.e., storage time). This meta information can also include the classification of content during the content review process. This classification can include first-level classification, second-level classification, third-level classification, and tag information. For example, a piece of content explaining XX brand mobile phones may have the first-level classification as technology, the second-level classification as smartphones, the third-level classification as domestic mobile phones, and the tag information as XX brand and XX model.
[0226] It's important to note that different content pools can be set up within the content storage server based on different content categories, and content of different categories can be stored in their corresponding content pools. Recommendation and distribution servers, as well as deduplication servers, all need to retrieve content from the content storage server. For example, the deduplication server can load content that has been stored and enabled over a period of time (e.g., one week) based on business needs. Duplicate content that is re-stored will be marked with a filter and will not be provided to the recommendation and distribution server for publishing.
[0227] S12. The uplink and downlink content interface servers write the content to be published to the scheduling center server.
[0228] It should be noted that the execution order of steps S11 and S12 can be flexibly set according to actual needs. For example, steps S11 and S12 can be executed simultaneously, or steps S11 can be executed first and then steps S12, or steps S12 can be executed first and then steps S11, etc.
[0229] The scheduling center server can be used to manage the entire scheduling process of content flow, receiving content stored in the content storage server through the uplink and downlink content interface servers, and obtaining the metadata of the content from the content storage server.
[0230] S13, The scheduling center server calls the deduplication server's content deduplication service.
[0231] The scheduling center server can schedule the deduplication server to mark and filter the content stored on the duplicate content storage server, generate deduplication transaction information, and synchronize the deduplication transaction information to the potential account identification fusion model in the potential account identification server as input.
[0232] The dispatch center server can also dispatch potential account identification servers to evaluate and calculate the potential account score ranking of each self-media posting account, which can be used in practical application scenarios such as subsequent manual review and dispatch or account weighting in the distribution process.
[0233] The deduplication operation of a deduplication server can include title deduplication, cover image deduplication, content text deduplication, video fingerprint deduplication, and audio fingerprint deduplication. For example, simhash (a hash algorithm) and BERT (Bidirectional Encoder Representations from Transformers) algorithms can be used to vectorize titles, cover images, and content text. For video content, video and audio fingerprints are extracted to construct vectors, and then the distance between vectors (such as Euclidean distance) is calculated to determine whether there is a duplicate, and duplicate content is filtered out. During the deduplication process, the deduplication server can obtain the user account's account level from a potential account identification server, and if the content is the same, use the content from the account with the higher potential account ranking.
[0234] Before invoking the content deduplication service of the deduplication server, the scheduling center server can obtain relevant data information of candidate accounts. The process of identifying potential accounts can be as follows: Figure 2d As shown, it includes:
[0235] S20. Obtain content metadata.
[0236] S21. Read statistical data from the statistical data storage server.
[0237] S22. Potential account identification server constructs potential account identification fusion model.
[0238] S23. The dispatch center server calls the potential account identification server and uses the potential account identification fusion model to identify the target account (specifically, the potential account) of the preset account type.
[0239] The statistical data can include the number of times content is viewed, likes, comments, and total interactions (such as reposts, shares, or favorites), as well as the time parameters and quantity of content published. This statistical data can also include the content published by candidate accounts. The potential account identification server can use a potential account identification fusion model to calculate the basic potential, prior potential, and growth potential of candidate accounts based on the statistical data. It then merges these three potentials to determine the probability that a candidate account belongs to a preset account type. The magnitude of this probability can be used for review scheduling and ranking, as well as for weighting high-quality content support, thereby increasing the probability of distributing high-quality content.
[0240] The potential account identification server can send the total potential score of candidate accounts to the scheduling center server. This allows the scheduling center server to mark the total potential score of user accounts, providing a reference for subsequent business scenarios. For example, the review server can prioritize reviewing content from user accounts with high total potential scores, and the recommendation and distribution server can prioritize recommending content from user accounts with high total potential scores.
[0241] The storage process for statistical data pre-stored in the statistical data storage server can be as follows: Figure 2e As shown,
[0242] S30, the upstream and downstream content interface servers obtain content indexes and published content from the content consumer.
[0243] S31. The content consumption end reports the content published by the user account and the statistical information of the end to the statistical reporting interface server and the upstream and downstream content interface server.
[0244] S32. The statistical reporting interface server writes the statistical data obtained from the content published by the user account and the terminal statistical information into the statistical data storage server.
[0245] Among them, published content refers to historical content. The content consumer can communicate with the upstream and downstream content interface servers. The upstream and downstream content interface servers can obtain the content index and published content from the content consumer. For example, the content index can be obtained through the Feeds recommendation distribution.
[0246] The content reporting interface server reports user account-published content to the statistics reporting interface server. This content can include: content to be published obtained from the content producer and published content obtained from the content consumer. The statistics can include user actions such as comments, reposts, shares, favorites, or likes based on the published content, as well as exposure data for Feeds (source messages) content (exposure data can refer to the visual aspects of the content display).
[0247] This embodiment constructs a potential account identification fusion model, which quantifies the total potential score of candidate accounts using unsupervised machine learning algorithms. It comprehensively considers multiple dimensions, including grade data, published content, and interaction information, quantifying each dimension individually and then fusing these features into a single quantitative score to rank accounts. Different thresholds can be set for different grades, allowing for clear distinction between account levels. The potential account identification fusion model can update its ranking daily based on data changes, prioritizing high-quality user accounts for inclusion in the recommendation pool. This concentrates traffic on truly high-quality content creators, reducing waste of valuable traffic and maximizing incentives for the platform's high-quality and active authors. Furthermore, the unsupervised modeling method eliminates the need for manual annotation, reducing costs and improving processing efficiency, thus fostering a virtuous cycle and a healthy content ecosystem.
[0248] S14. The dispatch center server calls the content review service of the review server.
[0249] The content review server can review content using a pre-built review model or through manual review. This review model can be flexibly configured according to actual needs. For example, it can segment text within the content, perform semantic analysis on the segmented words, and analyze for sensitive words or security issues based on the semantic analysis results. If sensitive words or security issues are found, the review will fail, and the content can be prohibited from publication. Similarly, it can identify images within the content to determine if they contain prohibited elements; if so, the review will fail, and the content can be prohibited from publication. When manual review is invoked, the reviewer can read information stored on the content storage server, and the results and status of the manual review can be sent back to the content storage server for storage.
[0250] S15. The dispatch center server updates the metadata in the content storage server based on the audit results.
[0251] For example, when the review is approved, the updated metadata can include information related to the content being approved; when the review is rejected, the updated metadata can include information related to the content being rejected.
[0252] S16. The scheduling center server sends the content to be published to the recommendation distribution server.
[0253] Before sending content to be published to the recommendation distribution server, the dispatch center server can base its decisions on, for example... Figure 2dThe process shown determines the probability that a candidate account belongs to a preset account type, and the probability, along with the content to be published, can be sent to the recommendation distribution server.
[0254] S17. The recommendation distribution server sends the content to be published to the content distribution exit server.
[0255] The recommendation distribution server sends the content to be published to the content distribution exit server based on the probability.
[0256] S18. The content distribution export server pushes the content to be published to the content consumer.
[0257] Among them, the content distribution egress server can be a group of access servers deployed geographically near the content consumption end. The content distribution egress server can obtain the recommendation distribution results and push the content to be published to the content consumption end. When the content consumption end receives the content to be published, it can display the content to be published.
[0258] As can be seen from the above, this embodiment can obtain at least one candidate account in a first application, the level data of the candidate account, the published content of the candidate account, and interaction data through a server. The interaction data refers to the interaction data of user accounts in the first application with respect to the published content. A reference account associated with the candidate account is determined. The reference account is a user account in a second application, and the reference account belongs to the same user as the candidate account. The level data of the candidate account and the level data of the reference account are fused to obtain the account popularity information of the candidate account. The published content is analyzed to obtain the content quality information of the published content. The interaction data is analyzed to obtain the user attention information corresponding to the published content of the candidate account. The user attention information represents the attractiveness of the published content of the candidate account to users. The account popularity information, the content quality information, and the user attention information are fused to obtain the probability that the candidate account belongs to a preset account type. Based on the probability, the candidate account is classified to obtain the target account of the preset account type. This application does not require manual annotation of account sample data, which can reduce manpower and material costs and improve the efficiency of account classification.
[0259] To better implement the above methods, embodiments of this application also provide an account classification device, such as... Figure 3a As shown, the account classification device may include an acquisition unit 301, a determination unit 302, a prediction unit 303, and a classification unit 304, as follows:
[0260] (1) Obtain unit 301;
[0261] The acquisition unit 301 is used to acquire at least one candidate account in the first application, the level data of the candidate account, the published content of the candidate account, and the interaction data, wherein the interaction data is the interaction data of the user account in the first application with respect to the published content.
[0262] (2) Determine unit 302;
[0263] The determining unit 302 is used to determine a reference account associated with the candidate account, wherein the reference account is a user account in the second application and the reference account and the candidate account belong to the same user.
[0264] (3) Prediction unit 303;
[0265] The prediction unit 303 is used to predict the probability that the candidate account belongs to a preset account type based on the candidate account's level data, the published content, the interaction data, and the reference account's level data.
[0266] Optionally, in some embodiments of this application, the prediction unit 303 may include a ranking fusion subunit 3031, a quality analysis subunit 3032, an interaction analysis subunit 3033, and a fusion subunit 3034, see [link to relevant documentation]. Figure 3b ,as follows:
[0267] The level fusion subunit 3031 is used to fuse the level data of the candidate account and the level data of the reference account to obtain the account popularity information of the candidate account.
[0268] The quality analysis subunit 3032 is used to perform quality analysis on the published content to obtain the content quality information of the published content.
[0269] Interaction analysis subunit 3033 is used to analyze the interaction data to obtain user attention information corresponding to the content published by the candidate account, wherein the user attention information represents the attractiveness of the content published by the candidate account to users.
[0270] The fusion subunit 3034 is used to fuse the account popularity information, the content quality information and the user attention information to obtain the probability that the candidate account belongs to a preset account type.
[0271] Optionally, in some embodiments of this application, the level data includes its own level information and fan count level information;
[0272] The level fusion subunit 3031 can be specifically used to fuse the self-level information of the candidate account and the self-level information of the reference account to obtain the self-popularity information of the candidate account; fuse the fan number level information of the candidate account and the fan number level information of the reference account to obtain the fan popularity information of the candidate account; and fuse the self-popularity information and the fan popularity information to obtain the account popularity information of the candidate account.
[0273] Optionally, in some embodiments of this application, the quality analysis subunit 3032 may specifically be used to perform quantitative statistical processing on published content that meets preset quality conditions to obtain a first statistical result, and calculate the quality information of the published content based on the first statistical result; perform quantitative statistical processing on the original content in the published content to obtain a second statistical result, and calculate the originality information of the published content based on the second statistical result; calculate the professionalism information of the published content based on the category to which the published content belongs; and calculate the content quality information of the published content based on the quality information, the originality information, and the professionalism information.
[0274] Optionally, in some embodiments of this application, the interactive data includes interactive data of at least one interactive type;
[0275] The interaction analysis subunit 3033 can be used to calculate the growth rate of the interaction data of each interaction type within a preset time period; and to merge the growth rates of the interaction data of each interaction type within the preset time period to obtain the user attention information corresponding to the content published by the candidate account.
[0276] Optionally, in some embodiments of this application, the step "merging the growth rates of interaction data for each interaction type within a preset time period to obtain user attention information corresponding to the content published by the candidate account" may include:
[0277] Obtain the weights corresponding to the interaction data for each interaction type;
[0278] Based on the weights, the growth rates of interaction data for each interaction type within a preset time period are weighted and fused to obtain the user attention information corresponding to the content published by the candidate accounts.
[0279] Optionally, in some embodiments of this application, the step "obtaining the weights corresponding to the interaction data of each interaction type" may include:
[0280] For each type of interaction data, determine the importance of the interaction data for that interaction type relative to the interaction data for a reference interaction type.
[0281] Based on the aforementioned importance, a comparison matrix of the interaction data for the aforementioned interaction type is constructed;
[0282] Based on the comparison matrix, the weights corresponding to the interactive data of the interactive type are determined.
[0283] Optionally, in some embodiments of this application, the step "determining the weight corresponding to the interaction data of the interaction type based on the comparison matrix" may include:
[0284] A consistency check is performed on the comparison matrix to obtain the consistency ratio of the comparison matrix;
[0285] When the consistency ratio is greater than a preset value, return to the operation of obtaining the importance of the interaction data of the interaction type relative to the interaction data of the reference interaction type;
[0286] The weights corresponding to the interactive data of the interaction type are determined based on a comparison matrix where the consistency ratio is not greater than a preset value.
[0287] (4) Classification unit 304;
[0288] The classification unit 304 is used to classify the candidate accounts according to the probability to obtain the target account of the preset account type.
[0289] As can be seen from the above, this embodiment can acquire at least one candidate account in a first application, the level data of the candidate account, the published content of the candidate account, and the interaction data through the acquisition unit 301. The interaction data refers to the interaction data of the user account in the first application with respect to the published content. The determination unit 302 determines a reference account associated with the candidate account. The reference account is a user account in a second application, and the reference account belongs to the same user as the candidate account. The prediction unit 303 predicts the probability that the candidate account belongs to a preset account type based on the level data of the candidate account, the published content, the interaction data, and the level data of the reference account. The classification unit 304 classifies the candidate account according to the probability to obtain the target account of the preset account type. This application does not require manual annotation of account sample data, which can reduce manpower and material costs and improve the efficiency of account classification.
[0290] This application also provides an electronic device, such as... Figure 4 The diagram shows a structural schematic of an electronic device involved in an embodiment of this application. This electronic device can be a terminal or a server, specifically:
[0291] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0292] The processor 401 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, it performs various functions and processes data, thereby performing overall detection of the electronic device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.
[0293] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0294] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0295] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0296] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:
[0297] The process involves: acquiring at least one candidate account in a first application, the level data of the candidate account, the published content of the candidate account, and interaction data, wherein the interaction data refers to the interaction data of a user account in the first application with respect to the published content; determining a reference account associated with the candidate account, wherein the reference account is a user account in a second application and belongs to the same user as the candidate account; predicting the probability that the candidate account belongs to a preset account type based on the level data of the candidate account, the published content, the interaction data, and the level data of the reference account; and classifying the candidate account according to the probability to obtain a target account of the preset account type.
[0298] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0299] As can be seen from the above, this embodiment can obtain at least one candidate account in a first application, the level data of the candidate account, the published content of the candidate account, and interaction data, wherein the interaction data is the interaction data of a user account in the first application with respect to the published content; determine a reference account associated with the candidate account, wherein the reference account is a user account in a second application, and the reference account and the candidate account belong to the same user; predict the probability that the candidate account belongs to a preset account type based on the level data of the candidate account, the published content, the interaction data, and the level data of the reference account; classify the candidate account according to the probability to obtain the target account of the preset account type. This application does not require manual annotation of account sample data, which can reduce manpower and material costs and improve the efficiency of account classification.
[0300] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0301] Therefore, embodiments of this application provide a storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the account classification methods provided in embodiments of this application. For example, the instructions can execute the following steps:
[0302] The process involves: acquiring at least one candidate account in a first application, the level data of the candidate account, the published content of the candidate account, and interaction data, wherein the interaction data refers to the interaction data of a user account in the first application with respect to the published content; determining a reference account associated with the candidate account, wherein the reference account is a user account in a second application and belongs to the same user as the candidate account; predicting the probability that the candidate account belongs to a preset account type based on the level data of the candidate account, the published content, the interaction data, and the level data of the reference account; and classifying the candidate account according to the probability to obtain a target account of the preset account type.
[0303] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0304] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0305] Since the instructions stored in the storage medium can execute the steps in any of the account classification methods provided in the embodiments of this application, the beneficial effects that any of the account classification methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.
[0306] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations of the account classification described above.
[0307] The above provides a detailed description of an account classification method, apparatus, electronic device, and storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An account classification method, characterized in that, include: The system acquires at least one candidate account in a first application, the level data of the candidate account, the published content of the candidate account, and interaction data, wherein the interaction data is the interaction data of a user account in the first application with respect to the published content; the interaction data includes interaction data of at least one interaction type. Identify a reference account associated with the candidate account, wherein the reference account is a user account in the second application and the reference account belongs to the same user as the candidate account; Based on the candidate account's level data, the published content, the interaction data, and the reference account's level data, predicting the probability that the candidate account belongs to a preset account type includes: fusing the candidate account's level data and the reference account's level data to obtain the candidate account's account popularity information; performing quality analysis on the published content to obtain the published content's content quality information; calculating the growth rate of interaction data for each interaction type within a preset time period; fusing the growth rates of interaction data for each interaction type within the preset time period to obtain the user attention information corresponding to the candidate account's published content; and fusing the account popularity information, the content quality information, and the user attention information to obtain the probability that the candidate account belongs to a preset account type. Based on the probability, the candidate accounts are classified to obtain the target accounts of the preset account type.
2. The method according to claim 1, characterized in that, The level data includes its own level information and the number of followers level information; The process of fusing the level data of the candidate accounts and the level data of the reference accounts to obtain the account popularity information of the candidate accounts includes: The self-level information of the candidate account and the self-level information of the reference account are fused to obtain the self-popularity information of the candidate account. The follower count level information of the candidate account and the follower count level information of the reference account are merged to obtain the follower popularity information of the candidate account. The self-popularity information and the fan popularity information are merged to obtain the account popularity information of the candidate account.
3. The method according to claim 1, characterized in that, The quality analysis of the published content to obtain content quality information includes: The published content that meets the preset quality conditions is subjected to quantity statistical processing to obtain a first statistical result, and the quality information of the published content is calculated based on the first statistical result. The original content in the published content is statistically analyzed to obtain a second statistical result, and the originality information of the published content is calculated based on the second statistical result. Based on the category to which the published content belongs, calculate the professionalism information of the published content; Based on the quality information, the originality information, and the professionalism information, the content quality information of the published content is calculated.
4. The method according to claim 1, characterized in that, The step of integrating the growth rates of interaction data for each interaction type within a preset time period to obtain user attention information corresponding to the content published by the candidate accounts includes: Obtain the weights corresponding to the interaction data for each interaction type; Based on the weights, the growth rates of interaction data for each interaction type within a preset time period are weighted and fused to obtain the user attention information corresponding to the content published by the candidate accounts.
5. The method according to claim 4, characterized in that, The process of obtaining the weights corresponding to the interaction data for each interaction type includes: For each type of interaction data, determine the importance of the interaction data for that interaction type relative to the interaction data for a reference interaction type. Based on the aforementioned importance, a comparison matrix of the interaction data for the aforementioned interaction type is constructed; Based on the comparison matrix, the weights corresponding to the interactive data of the interactive type are determined.
6. The method according to claim 5, characterized in that, Determining the weight corresponding to the interactive data of the interaction type based on the comparison matrix includes: A consistency check is performed on the comparison matrix to obtain the consistency ratio of the comparison matrix; When the consistency ratio is greater than a preset value, return to the operation of obtaining the importance of the interaction data of the interaction type relative to the interaction data of the reference interaction type; The weights corresponding to the interactive data of the interaction type are determined based on a comparison matrix where the consistency ratio is not greater than a preset value.
7. An account classification device, characterized in that, include: The acquisition unit is configured to acquire at least one candidate account in a first application, the level data of the candidate account, the published content of the candidate account, and the interaction data, wherein the interaction data is the interaction data of the user account in the first application in response to the published content; the interaction data includes interaction data of at least one type of interaction. A determining unit is used to determine a reference account associated with the candidate account, wherein the reference account is a user account in the second application, and the reference account and the candidate account belong to the same user; The prediction unit is used to predict the probability that the candidate account belongs to a preset account type based on the candidate account's level data, the published content, the interaction data, and the reference account's level data. A classification unit is used to classify the candidate accounts according to the probability to obtain the target account of the preset account type; The prediction unit includes a level fusion subunit, a quality analysis subunit, an interaction analysis subunit, and a fusion subunit; The grade fusion subunit is used to fuse the grade data of the candidate account and the grade data of the reference account to obtain the account popularity information of the candidate account. The quality analysis subunit is used to perform quality analysis on the published content to obtain the content quality information of the published content. The interaction analysis subunit is used to calculate the growth rate of the interaction data of each interaction type within a preset time period. By merging the growth rates of interaction data for each interaction type within a preset time period, user attention information corresponding to the content published by the candidate account is obtained. The fusion subunit is used to fuse the account popularity information, the content quality information, and the user attention information to obtain the probability that the candidate account belongs to a preset account type.
8. The apparatus according to claim 7, characterized in that, The level data includes its own level information and the number of followers level information; The grade fusion subunit is specifically used to fuse the self-grade information of the candidate account and the self-grade information of the reference account to obtain the self-popularity information of the candidate account; and to fuse the fan count grade information of the candidate account and the fan count grade information of the reference account to obtain the fan popularity information of the candidate account. The self-popularity information and the fan popularity information are merged to obtain the account popularity information of the candidate account.
9. The apparatus according to claim 7, characterized in that, The quality analysis subunit is specifically used to perform quantitative statistical processing on the published content that meets the preset quality conditions, obtain a first statistical result, and calculate the quality information of the published content based on the first statistical result. The original content in the published content is statistically analyzed to obtain a second statistical result, and the originality information of the published content is calculated based on the second statistical result. Based on the category to which the published content belongs, calculate the professionalism information of the published content; Based on the quality information, the originality information, and the professionalism information, the content quality information of the published content is calculated.
10. The apparatus according to claim 7, characterized in that, The step of integrating the growth rates of interaction data for each interaction type within a preset time period to obtain user attention information corresponding to the content published by the candidate accounts includes: Obtain the weights corresponding to the interaction data for each interaction type; Based on the weights, the growth rates of interaction data for each interaction type within a preset time period are weighted and fused to obtain the user attention information corresponding to the content published by the candidate accounts.
11. The apparatus according to claim 10, characterized in that, The process of obtaining the weights corresponding to the interaction data for each interaction type includes: For each type of interaction data, determine the importance of the interaction data for that interaction type relative to the interaction data for a reference interaction type. Based on the aforementioned importance, a comparison matrix of the interaction data for the aforementioned interaction type is constructed; Based on the comparison matrix, the weights corresponding to the interactive data of the interactive type are determined.
12. The apparatus according to claim 11, characterized in that, Determining the weight corresponding to the interactive data of the interaction type based on the comparison matrix includes: A consistency check is performed on the comparison matrix to obtain the consistency ratio of the comparison matrix; When the consistency ratio is greater than a preset value, return to the operation of obtaining the importance of the interaction data of the interaction type relative to the interaction data of the reference interaction type; The weights corresponding to the interactive data of the interaction type are determined based on a comparison matrix where the consistency ratio is not greater than a preset value.
13. An electronic device, characterized in that, It includes a memory and a processor; the memory stores an application program, and the processor runs the application program within the memory to perform the operations in the account classification method according to any one of claims 1 to 6.
14. A storage medium, characterized in that, The storage medium stores a plurality of instructions adapted for loading by a processor to execute the steps of the account classification method according to any one of claims 1 to 6.
15. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; the processor of the computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in the account classification method according to any one of claims 1 to 6.