Content distribution method and apparatus, electronic device, and storage medium
By automatically assessing the originality of content from self-media accounts, the problem of low efficiency in manual review has been solved, achieving efficient content review and distribution.
Patent Information
- Application Number
- CN202010478221.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-29
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2040-09-05
AI Technical Summary
Current technologies for reviewing content on social media accounts are inefficient, and manual review is time-consuming and labor-intensive, making it difficult to efficiently identify the originality of content.
By obtaining the type of content to be distributed and the number of published contents, the system determines account distribution information, collects historical content from reference accounts, calculates content similarity and account identifiers, automatically assesses content originality, and distributes content when it reaches a preset value.
It improved the efficiency of content review, reduced the waste of human resources, accurately identified the originality of content, and improved the efficiency of content distribution.
Smart Images

Figure CN111639291B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically to a content distribution method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of modern technology, media outlets have found it increasingly convenient to disseminate information. These media can register accounts on online platforms and then publish information, such as text, audio, and video messages. This also includes self-media, which refers to the dissemination of facts and news by ordinary people through the internet and other channels. Recent years have seen a surge in content creation, with major internet companies actively entering the content market, leading to a proliferation of self-media platforms. Everyone can create their own self-media through writing. A vast number of self-media accounts create a massive amount of articles daily; however, some self-media accounts may plagiarize content from other self-media platforms. Therefore, it is necessary to verify whether the content published by self-media accounts is historical.
[0003] Currently, the content published by self-media accounts is reviewed manually. However, due to the large number of self-media accounts, it is time-consuming, labor-intensive, and inefficient for operators to review each submission one by one every day. Summary of the Invention
[0004] This application provides a content distribution method, apparatus, electronic device, and storage medium that can improve the efficiency of content review.
[0005] This application provides a content distribution method, including:
[0006] Obtain the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account;
[0007] Based on the content type and the number of contents published by the distribution account, determine the distribution information of the distribution account in terms of content;
[0008] Collect a list of reference accounts and the historical content published by each reference account in the historical time period from the content distribution system to which the distribution account belongs;
[0009] Based on distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account, the originality of the content of the distribution account is calculated;
[0010] When the calculated originality of the content is greater than a preset value, the content to be distributed is distributed.
[0011] Accordingly, this application also provides a content distribution device, including:
[0012] The acquisition module is used to acquire the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account;
[0013] The determination module is used to determine the distribution information of the distribution account on the content based on the content type and the number of contents published by the distribution account;
[0014] The collection module is used to collect a list of reference accounts and the historical content published by each reference account in the historical time period from the content distribution system to which the distribution account belongs.
[0015] The calculation module is used to calculate the originality of the content of the distribution account based on distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account.
[0016] The distribution module is used to distribute the content to be distributed when the calculated originality of the content is greater than a preset value.
[0017] Optionally, in some embodiments of this application, the computing module includes:
[0018] An extraction submodule is used to extract the content authentication information of the historical content and the content authentication information of the content to be distributed, respectively.
[0019] The update submodule is used to update the account identifier of the distribution account based on the content authentication information of the historical content and the content authentication information of the content to be distributed;
[0020] The calculation submodule is used to calculate the originality of the content of the distribution account based on the distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier.
[0021] Optionally, in some embodiments of this application, the update submodule includes:
[0022] The generation unit is used to generate the relevance between the distribution account and each reference account in the reference account list based on the content authentication information of the historical content and the content authentication information of the content to be distributed;
[0023] The update unit is used to update the account identifier of the distribution account based on the relevance between the distribution account and each reference account in the reference account list.
[0024] Optionally, in some embodiments of this application, the generation unit includes:
[0025] The detection subunit is used to detect whether the content identification information of the historical content is consistent with the content identification information of the content to be distributed.
[0026] The first calculation subunit is used to calculate the similarity between the content title information of each historical content and the content title information of the content to be distributed;
[0027] A generation subunit is used to generate the relevance between the distribution account and the reference account based on the similarity between the content title information of the historical content and the content title information of the content to be distributed, as well as the detection results.
[0028] Optionally, in some embodiments of this application, the generating subunit is specifically used for:
[0029] The similarity between the content title information and the content title information of the content to be distributed, as well as the detection results, are fused to obtain the correlation between the distribution account and the reference account.
[0030] Optionally, in some embodiments of this application, the computing module includes:
[0031] The acquisition unit is used to acquire the account authentication level of each reference account in the reference account list in the content distribution system to which it belongs;
[0032] The calculation unit is used to calculate the similarity between the content to be distributed and each historical content based on the account authentication level of each reference account in the reference account list.
[0033] The fusion unit is used to fuse distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier to obtain the content originality of the distribution account.
[0034] Optionally, in some embodiments of this application, the fusion unit is specifically used for:
[0035] The content duplication score of the content to be distributed is generated based on the similarity between the content to be distributed and each historical content.
[0036] By integrating the updated account identifier, distribution information, and the content duplication rate of the content to be distributed, the originality of the content of the distribution account is obtained.
[0037] Optionally, in some embodiments of this application, the computing unit includes:
[0038] The sub-unit is determined to determine the weight of each piece of historical content based on the account authentication level of each reference account in the reference account list.
[0039] The second calculation subunit is used to calculate the similarity between the content to be distributed and each historical content based on the weights corresponding to each historical content.
[0040] Optionally, in some embodiments of this application, the second computing subunit is specifically used for:
[0041] The content to be distributed and each historical content are vectorized to obtain a first vector corresponding to the content to be distributed and a second vector corresponding to each historical content.
[0042] Calculate the distance between the first vector and each of the second vectors respectively;
[0043] Based on the distance between the first vector and each second vector, as well as the weights corresponding to each historical content, the similarity between the content to be distributed and each historical content is determined.
[0044] This application, after obtaining the content to be distributed, the distribution account, and the content types corresponding to the content already published by the distribution account, determines the distribution information of the distribution account based on the content types. Then, it collects a list of reference accounts and the historical content published by each reference account in the historical time period from the content distribution system to which the distribution account belongs. Next, it extracts the content authentication information of the historical content and the content authentication information of the content to be distributed. Then, it updates the account identifier of the distribution account based on the content authentication information of the historical content and the content to be distributed. Finally, based on the distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier, it calculates the content originality of the distribution account. When the calculated content originality is greater than a preset value, the content to be distributed is distributed. Therefore, this scheme can improve the efficiency of content review. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1a This is a schematic diagram illustrating a scenario of the content distribution method provided in this application;
[0047] Figure 1b This is a flowchart illustrating the content distribution method provided in this application;
[0048] Figure 1c This is a schematic diagram illustrating the verticality of article posting in the content distribution method provided in this application;
[0049] Figure 1d This is a schematic diagram illustrating the calculation of the similarity between the content to be distributed and the reference content in the content distribution method provided in this application;
[0050] Figure 2a This is another flowchart illustrating the content distribution method provided in this application;
[0051] Figure 2b This is another scenario illustration of the content distribution method provided in this application;
[0052] Figure 3 This is a schematic diagram of the content distribution device provided in this application;
[0053] Figure 4 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0054] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0055] This application provides a content distribution method, apparatus, electronic device, and storage medium.
[0056] Specifically, the content distribution device can be integrated into a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.
[0057] For example, please see Figure 1aThe content distribution device is integrated into a server. After obtaining the content to be distributed, the distribution account, and the content types corresponding to the content already published by the distribution account (e.g., a user logs into the distribution account via a terminal, the content to be distributed is article A, and the distribution account is a self-media account H in self-media platform X), the server determines the content distribution information of self-media account H based on the content types corresponding to the content already published by self-media account H. Then, it collects a list of reference accounts and the historical content published by each reference account in the historical time period from the content distribution system (self-media platform X) to which the distribution account belongs. Next, the server extracts the content authentication information of the historical content and the content authentication information of the content to be distributed. The content authentication information may include content title information and content identification information. The content identification information may include watermarks in the article, the author of the article, and the publication time of the article, etc. Then, the server updates the account identifier of the distribution account based on the content authentication information of the historical content and the content authentication information of the content to be distributed. Finally, the server calculates the content originality of the distribution account based on the distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier. When the calculated content originality is greater than a preset value, the content to be distributed is distributed.
[0058] The content distribution method provided in this application can update the account identifier of the distribution account based on the content authentication information of historical content and the content to be distributed. Based on distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier, the originality of the content of the distribution account is calculated. When determining whether the content to be distributed is original, the content authentication information of historical content and the content to be distributed is taken into account, making the calculated originality of the content of the distribution account more accurate. Moreover, the entire process does not require manual intervention, reducing the waste of human resources, improving the efficiency of content review, and thus improving the efficiency of content distribution.
[0059] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the priority of the embodiments.
[0060] A content distribution method includes: obtaining content to be distributed, a distribution account, and content types corresponding to the content already published by the distribution account; determining the distribution information of the distribution account on the content based on the content types and the number of published contents by the distribution account; collecting a list of reference accounts and historical content published by each reference account in the historical time period from the content distribution system to which the distribution account belongs; calculating the content originality of the distribution account based on the distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account; and distributing the content to be distributed when the calculated content originality is greater than a preset value.
[0061] Please see Figure 1b , Figure 1b A flowchart illustrating the content distribution method provided in this application. The specific flow of this content distribution method can be as follows:
[0062] 101. Obtain the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account.
[0063] For example, specifically, one can obtain the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account by accessing a network interface. The content to be distributed is the content to be distributed, obtained from content published by the distribution account. The distribution account is an account with content publishing capabilities, and can be a self-media account. Self-media (We Media) can be understood as a general term for new media where private, autonomous communicators use modern, electronic means to transmit standardized and non-standardized information to an unspecified majority or a specific individual. A self-media account can be an account registered on an independent content distribution platform capable of autonomously publishing content (such as a Weibo account), or an account registered on a content distribution platform integrated into a social media platform capable of autonomously publishing content. A content distribution platform integrated into a social media platform can be a content distribution platform integrated into an instant messaging platform.
[0064] 102. Based on the content type and the number of contents published by the distribution account, determine the distribution information of the distribution account in terms of content.
[0065] For example, specifically, the number of articles published by the collection and distribution account. If the distribution account has published a total of 7 articles, of which 2 articles are in the lifestyle category and 5 articles are in the medical category, then the distribution account's content distribution is: lifestyle and medical, with no distribution in other fields.
[0066] It's important to note that some content reposting accounts may have a very diverse content distribution, potentially covering multiple unrelated fields. For example, a reposting account might have content on pharmaceuticals, metal manufacturing, automobile manufacturing, and sports. In this case, the account is very likely a reposting account. Original content accounts, on the other hand, typically focus on specific areas, resulting in a more concentrated content distribution. They tend to publish a large amount of content within their targeted areas. This leads to the concept of "posting verticality," which reflects the degree of focus a reposting account places on its area of expertise. Please refer to [link to relevant documentation]. Figure 1cThe core principle of post verticality can be explained using normal distribution and kurtosis: it represents the distribution of posts by vertical category on an account. The horizontal axis represents the post vertical category (which can be represented by the first-level category of the post content), and the vertical axis represents the proportion of posts in the corresponding vertical category. If we compare it to a normal distribution, then the area of the shaded part is 1 (the sum of the proportions of posts in all vertical categories is 1). That is, post verticality case one (left figure): the smaller the kurtosis of the normal distribution (the smaller the proportion of the vertical category with the most posts), the larger the standard deviation (the more dispersed the post vertical categories) when the area is constant, that is, the posts are not vertical. Then, case two (right figure): the larger the kurtosis of the normal distribution (the larger the proportion of the vertical category with the most posts), the smaller the standard deviation (the more concentrated the post vertical categories) when the area is constant.
[0067] 103. Collect a list of reference accounts and the historical content published by each reference account in the historical time period from the content distribution system to which the distribution account belongs.
[0068] Reference accounts refer to distribution accounts certified by the content distribution system (also known as the content distribution platform). These can include corporate accounts and private accounts. For example, a corporate account can be a distribution account for a news media outlet, while a private account can be a distribution account for a writer. The specifics depend on the actual situation. Specifically, the reference account list and the historical content published by each reference account within a historical time period can be collected from the content distribution platform. The historical time period can be the past month, the past year, or the period from the registration of the reference account to the current time. The specifics depend on the actual situation.
[0069] 104. Based on distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account, calculate the originality of the content of the distribution account.
[0070] For the same content distribution platform, each distribution account is assigned an account identifier to distinguish whether it is a reposting account or an original content account. Furthermore, the account identifier is not static; it can be updated based on the content to be distributed and previously published historical content. Then, based on distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier, the originality of the content of the distribution account is calculated. Optionally, in some embodiments, the step "calculating the originality of the content of the distribution account based on distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account" may specifically include:
[0071] (11) Extract the content authentication information of historical content and the content authentication information of content to be distributed.
[0072] (12) Update the account identifier of the distribution account based on the content authentication information of the historical content and the content authentication information of the content to be distributed.
[0073] (13) Calculate the originality of the content of the distribution account based on the distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier.
[0074] Content authentication information is a type of protection information embedded in the content itself. It is primarily used to identify the source of the content, its creator (i.e., the account distributing the content), and its structural representation. For example, content authentication information may include content title information and content identification information. Content title information includes the text length and semantic information of the title, while content identification information includes watermark information and the time of publication. Content identification information is used to mark the content for easy identification later. Specifically, the account identifier of the distribution account can be updated based on the content authentication information of historical content and the content to be distributed. For instance, when based on historical content authentication information... By comparing the content authentication information of historical content with that of the content to be distributed, it is determined that the distribution account is a copycat account. If so, the account identifier of that distribution account is updated, indicating that it is a copycat account. Then, based on distribution information, the similarity between the content to be distributed and each piece of historical content, and the updated account identifier, the originality of the distribution account's content can be calculated. Originality refers to independently completed creation. Originality does not refer to works that distort, alter, plagiarize, or steal others' creations, nor does it refer to works that adapt, translate, annotate, or organize existing works. In essence, content originality is used to measure the originality of the content on the distribution account. To determine the degree of originality, since the units of measurement for distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier are different, it is necessary to standardize these data to include them in the calculation. This involves mapping their values to a specific range using a function transformation. Therefore, normalization can be performed on the distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier separately. After normalization, the results are weighted and summed according to a pre-defined strategy to obtain the content originality of the distribution account. It should be noted that when measuring the distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier separately... Before normalizing the distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier, values can be assigned to the distribution information and the updated account identifier. The distribution information can be used to determine the content proportion of the distribution account with the largest proportion of the content, and the corresponding proportion can be assigned to the distribution information. For example, if the content type with the largest proportion of the distribution account accounts for 80% of the content, then the value assigned to the distribution information is 80%. In addition, if the account identifier of the distribution account is an original account, then a value of 1 is assigned; if it is a plagiarized account, then a value of 0 is assigned. The preset strategy can be set by the content distribution system according to actual needs, which will not be elaborated here.
[0075] Optionally, in some embodiments, to improve the accuracy of the calculated originality of the content, the relevance between the distribution account and each reference account in the reference account list can be generated based on the content authentication information of historical content and the content authentication information of the content to be distributed. Then, the account identifier is updated based on the relevance. That is, the step "updating the account identifier of the distribution account based on the content authentication information of historical content and the content authentication information of the content to be distributed" can specifically include:
[0076] (21) Based on the content authentication information of historical content and the content authentication information of the content to be distributed, generate the relevance between the distribution account and each reference account in the reference account list;
[0077] (22) Update the account identifier of the distribution account based on the correlation between the distribution account and each reference account in the reference account list.
[0078] The higher the relevance between a distribution account and a reference account, the greater the probability that the distribution account is a copying account. There are many methods to generate the relevance between a distribution account and each reference account in the reference account list. For example, the similarity between the account name of the distribution account and the account name of the reference account can be determined as the relevance between the distribution account and the reference account. However, if the similarity between the account name of the distribution account and the account name of the reference account is very high, such as 90%, but the content distributed by the distribution account is completely unrelated to the content already published by the reference account, determining the similarity between the account name of the distribution account and the account name of the reference account as the relevance between the distribution account and the reference account may lead to inaccurate calculation of the originality of the content of the distribution account. Therefore, optionally, in some embodiments, the relevance between a distribution account and each reference account in the reference account list can be generated based on content identification information and content title information. That is, the step "generating the relevance between a distribution account and each reference account in the reference account list based on the content authentication information of historical content and the content authentication information of the content to be distributed" can specifically include:
[0079] (31) Check whether the content identification information of the historical content is consistent with the content identification information of the content to be distributed;
[0080] (32) Calculate the similarity between the content title information of each historical content and the content title information of the content to be distributed;
[0081] (33) Based on the similarity between the content title information of historical content and the content title information of the content to be distributed, as well as the detection results, generate the correlation between the distribution account and the reference account.
[0082] For example, if a distribution account reposts video content from "Voice of XX," the video content published by "Voice of XX" through the content distribution platform will be marked with "Voice of XX" identifier information, such as a "Voice of XX" watermark. Since "Voice of XX" is in the reference account list, but the distribution account corresponding to the content to be distributed is not "Voice of XX," it can be determined that the distribution account is strongly related to "Voice of XX." Then, a "1" can be recorded in the distribution account's account identifier to indicate that the distribution account reposted a reference account. If a distribution account copies content from two reference accounts, it's recorded as "1, 1". This means the number of times "1" is recorded corresponds to the number of reference accounts the distribution account copied. Additionally, the similarity between the title information of historical content and the title information of the content to be distributed can be calculated. Natural Language Processing (NLP) techniques can be used to process the titles of both historical and the content to be distributed, thus obtaining the similarity. NLP is an important field in computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language, i.e., the language people use in daily life, and thus it is closely related to linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0083] In natural language processing (NLP), text processing typically utilizes machine learning (ML) techniques. Machine learning is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0084] In this application, natural language processing technology is used to detect the similarity between the content titles of historical content and the content titles of the content to be distributed. For example, entities can be extracted from the content titles of historical content and the content titles of the content to be distributed, and then the similarity between entities can be calculated to generate the similarity between the content titles of historical content and the content titles of the content to be distributed. Finally, the similarity between the content title information and the content title information of the content to be distributed, as well as the detection results, are fused to obtain the relevance between the distribution account and the reference account.
[0085] For distribution accounts within the same content distribution system, each account is assigned an account verification level upon joining the system. Initially, this level is typically determined by the content distribution system based on its operational strategy. For example, content distribution system Q might pre-set five account levels: S, A, B, C, and D, and designate all S-level distribution accounts as reference accounts for system Q. That is, the reference account list includes all S-level distribution accounts within system Q. Additionally, accounts with a strong presence in a specific vertical field can be designated as A-level accounts. It's important to note that account levels are not static; the reference account level may vary. Accounts are typically determined by operational strategies. For other distribution accounts, such as A-level and B-level accounts, which are capable of rapid growth, their level in the content distribution system is determined by their originality and the distribution status of their content on the platform. The distribution status of the content includes user complaints and reports. Therefore, in some embodiments, the originality of the distribution account's content can be calculated based on the account authentication level corresponding to the reference account, the similarity between the content to be distributed and various historical content, distribution information, and the updated account identifier. That is, the step "calculate the originality of the distribution account's content based on distribution information, the similarity between the content to be distributed and various historical content, and the updated account identifier" can specifically include:
[0086] (41) Obtain the account authentication level of each reference account in the reference account list in its respective content distribution system;
[0087] (42) Based on the account authentication level of each reference account in the reference account list, calculate the similarity between the content to be distributed and each historical content;
[0088] (43) By integrating distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier, the originality of the content of the distribution account is obtained.
[0089] In practical applications, the content of authoritative media accounts registered in content distribution systems, as well as some well-known industry accounts, is often easily copied and plagiarized. Therefore, for the identification of distribution accounts, the similarity between the content to be distributed and each historical content can be calculated. Since the account authentication levels of different reference accounts may be different, in some embodiments, the weight of each historical content can be determined based on the account authentication level of the reference account, and then the similarity between the content to be distributed and each historical content can be calculated based on the determined weight. That is, the step "calculate the similarity between the content to be distributed and each historical content based on the account authentication level of each reference account in the reference account list" can specifically include:
[0090] (51) Determine the weight of each historical content based on the account authentication level of each reference account in the reference account list;
[0091] (52) Calculate the similarity between the content to be distributed and each historical content based on the weights corresponding to each historical content.
[0092] Please see Figure 1d For example, if a distribution account is classified as Level B in the content distribution system, a record of duplicate content (historical content) with an S-level account will be scored as 1. If a record of duplicate content (historical content) with an A-level account will be scored as 0.5. Records of the same or lower level will be scored as 0.
[0093] It should be noted that there are many ways to determine content deduplication. Content can include text, images, and videos. Therefore, to reduce computational load, the content (including historical content and content to be distributed) can be dimensionality-reduced. Specifically, in some embodiments, the step "calculating the similarity between the content to be distributed and each historical content based on the weights corresponding to each historical content" may include:
[0094] (61) Vectorize the content to be distributed and each historical content separately to obtain the first vector corresponding to the content to be distributed and the second vector corresponding to each historical content.
[0095] (62) Calculate the distance between the first vector and each of the second vectors respectively;
[0096] (63) Based on the distance between the first vector and each second vector and the weights corresponding to each historical content, determine the similarity between the content to be distributed and each historical content.
[0097] For example, for plain text, after word segmentation, the segmented text is converted into feature vectors, and then the distance between the vectors to be compared is calculated, such as Euclidean distance. However, since an article may have a large number of feature vector words, resulting in a high dimensionality of the entire vector, the computational cost is too high. Therefore, the feature vectors corresponding to the content to be distributed and the feature vectors corresponding to the historical content can be hashed separately to transform the high-dimensional feature vectors into a fingerprint. Then, the similarity between the historical content and the content to be distributed is determined by calculating the Hamming distance between the two fingerprints. The smaller the Hamming distance, the lower the similarity.
[0098] As for video content, we can extract the video feature vector and audio feature vector corresponding to the video content, and then calculate the distance between the vectors to determine whether the video is repeated. The specific method is similar to the previous one, so it will not be repeated here.
[0099] Subsequently, the content duplication rate of the content to be distributed can be generated based on the similarity between the content to be distributed and each historical content. Finally, the content originality of the distribution account is obtained by integrating the updated account identifier, distribution information and the content duplication rate of the content to be distributed. The norm calculation method can be used to integrate the updated account identifier, distribution information and the content duplication rate of the content to be distributed to obtain the content originality of the distribution account.
[0100] 105. When the calculated originality of the content is greater than the preset value, distribute the content to be distributed.
[0101] The preset value is set in advance and can be set according to the strategy of the content distribution system. For example, for a content distribution system with abundant original content, the preset value can be set to 1, while for a content distribution system lacking original content (i.e., with a small amount of original content), the preset value can be set to 10. When the originality of the content of the distribution account is greater than the preset value, the content to be distributed will be distributed.
[0102] This application, after obtaining the content to be distributed, the distribution account, and the content types corresponding to the content already published by the distribution account, determines the distribution information of the distribution account based on the content types and the number of published contents by the distribution account. Next, it collects a list of reference accounts and the historical content published by each reference account within a historical time period from the content distribution system to which the distribution account belongs. Finally, based on the distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account, it calculates the originality of the content of the distribution account. When the calculated originality is greater than a preset value, the content to be distributed is distributed. The content distribution method provided in this application calculates the originality of the content of the distribution account based on the distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account. When determining whether the content to be distributed is original, it considers not only the similarity between each historical content and the account identifier of the distribution account, but also the distribution information of the distribution account in the content, making the subsequently calculated originality of the content of the distribution account more accurate. Furthermore, the entire process requires no manual intervention, reducing the waste of human resources, improving the efficiency of content review, and thus improving the efficiency of content distribution.
[0103] The method described in the embodiments will be further described in detail below with examples.
[0104] In this embodiment, the content distribution device will be specifically integrated into the server as an example for illustration.
[0105] Please see Figure 2a A content distribution method, the specific process of which can be as follows:
[0106] 201. The server obtains the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account.
[0107] For example, specifically, a server can obtain the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account by accessing a network interface. The content to be distributed is the content to be distributed, obtained from content published by the distribution account. The distribution account is an account with content publishing capabilities, and can be a self-media account. Self-media (We Media) can be understood as a general term for new media where private, autonomous communicators use modern, electronic means to transmit standardized and non-standardized information to a large, unspecified number of people or a specific individual. A self-media account can be an account registered on an independent content distribution platform capable of autonomously publishing content (such as a Weibo account), or an account registered on a content distribution platform integrated into a social media platform capable of autonomously publishing content. A content distribution platform integrated into a social media platform can be a content distribution platform integrated into an instant messaging platform.
[0108] 202. The server determines the distribution information of the distribution account on the content based on the content type and the number of contents published by the distribution account.
[0109] For example, the server collects the number of articles published by the distribution account. If the distribution account has published a total of 7 articles, of which 2 articles are in the lifestyle category and 5 articles are in the medical category, then the distribution account's content distribution is: lifestyle and medical, with no distribution in other fields.
[0110] It's important to note that for some content-reposting accounts, the content distribution can be very diverse, potentially covering multiple unrelated fields. Therefore, the server can utilize the primary category results of the posted content to calculate the percentage of articles from the most frequently posted vertical category out of all posted vertical categories. Non-vertical posting is defined as the percentage of articles from the most frequently posted vertical category out of all articles from a given account. The smaller the percentage, the less vertical the posting, i.e., the distribution information S. p It can be expressed as follows:
[0111]
[0112] U represents the number of articles in the most popular vertical category published by an account within a certain period of time, and T represents the total number of articles published by an account within a certain period of time, which can be 1 month.
[0113] 203. The server collects a list of reference accounts and the historical content published by each reference account in the historical time period from the content distribution system to which the distribution account belongs.
[0114] 204. The server extracts the content authentication information of historical content and the content authentication information of content to be distributed.
[0115] The content authentication information includes content title information and content identification information. The content title information includes the text length and semantic information of the content title. The content identification information includes watermark information and content publication time information, etc. The content identification information is used to identify the content and facilitate subsequent content identification.
[0116] 205. The server updates the account identifier of the distribution account based on the content authentication information of the historical content and the content authentication information of the content to be distributed.
[0117] For the same content distribution platform, each distribution account is assigned an account identifier to distinguish whether it is a reposting account or an original content account. Therefore, the server can update the account identifier based on the content authentication information of historical content and the content to be distributed. For example, if the platform determines that the distribution account is a reposting account based on the historical and original content authentication information, then the account identifier is updated. The updated account identifier indicates that the distribution account is a reposting account. The updated account identifier S can be represented by the following formula: account The details are as follows:
[0118]
[0119] Among them, X tag : This indicates the match between the corresponding tag of the content posted by this account (i.e., the content to be distributed) and the tag of the account in the whitelist. Each match is recorded as 1. For example, if an account reposts video content from XX Voice, the content itself will usually be tagged with XX Voice in order to distribute the version. In this case, since XX Voice already has this account in the whitelist, and the account that posted the content is not this account, then it matches once.
[0120] X title This indicates the approximate number of times the title of the account's post is similar to the title of a post already published in the whitelist. If there is a match, it is counted as 0.5; otherwise, it is counted as 0.
[0121] 206. The server calculates the originality of the content of the distribution account based on the distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier.
[0122] Specifically, comparing the duplication of published content, i.e., the duplication between the content to be distributed and historical content, can be calculated. This can include calculating the duplication of text and images (using Simhash for text deduplication, image deduplication, and BERT for text deduplication) and video (using video fingerprint vector deduplication and audio fingerprint deduplication). Simhash, proposed by Charikar in 2002, is a hash algorithm that can calculate document similarity. Google uses it for deduplication of massive amounts of text. Simhash is a type of locality-sensitive hash. Its main idea is dimensionality reduction, transforming high-dimensional feature vectors into an f-bit fingerprint. The similarity between two articles is determined by calculating the Hamming distance between the two fingerprints. The smaller the Hamming distance, the lower the similarity (according to the paper "Detecting Near-Duplicates for Web Crawling"). Generally, a Hamming distance of 3 indicates that the two articles are identical. BERT stands for Bidirectional Encoder Representations from Transformers, a novel language model that trains pre-trained deep bidirectional representations by jointly tuning bidirectional Transformers across all layers. Based on the original implementation of the multi-layer bidirectional Transformer encoder described by Vaswani et al. (2017), and the Transformer architecture released by Google in 2017, BERT requires only an additional output layer, unlike typical Transformers which use a set of encoder and decoder networks. Through fine-tuning the pre-trained layer, it can meet various tasks. Video fingerprint and audio vector deduplication involves extracting video and audio features from the video content, vectorizing them, calculating the distance between the vectors to determine if the video is repeated, and finally, expressing the results using a 3-norm aggregation to determine the repetition rate S of the published content. copy as follows:
[0123]
[0124] Among them, X txt This indicates the result of text deduplication, specifically the result of the simhash algorithm on the main text; X pic This indicates the image deduplication results. For example, if most of the images in the main text are the same, it is considered a duplicate. Typically, if more than 50% of the images are the same, it is considered a duplicate. bert This represents the result of BERT text deduplication; X videoThis represents the fingerprint deduplication result of the video content; X voice This represents the audio fingerprint deduplication result of the video content.
[0125] For distribution accounts within the same content distribution system, each account is assigned an account verification level upon joining the system. Initially, this level is typically determined by the system's operational strategy. For example, content distribution system Q might pre-set five account levels: S, A, B, C, and D, designating all S-level accounts as reference accounts for Q. In other words, the reference account list includes all S-level distribution accounts within Q. Additionally, accounts with a strong presence in a specific vertical market might be designated as A-level accounts. It's important to note that account levels are not static; the level of reference accounts is usually determined by operational strategies. For other distribution accounts, such as A-level and B-level accounts, which are capable of rapid growth, their level in the content distribution system is determined by their originality and the distribution status of their content on the platform. The distribution status includes user complaints and reports. Therefore, in some embodiments, the server can calculate the originality of the distribution account's content based on the account authentication level of the reference account, the similarity between the content to be distributed and historical content, distribution information, and the updated account identifier. This can be achieved by using a norm calculation method, integrating the updated account identifier, distribution information, and the content duplication rate of the content to be distributed. The norm is a fundamental concept in mathematics. In functional analysis, it is defined in a normed linear space and satisfies certain conditions: nonnegativity, homogeneity, and the triangle inequality. It is often used to measure the length or size of each vector in a vector space (or matrix). The most commonly used norm is the p-norm. Its specific definition is as follows:
[0126] like ,So,
[0127] It can be verified that the p-norm does indeed satisfy the definition of a norm. The proof of the triangle inequality is not trivial; this conclusion is often called the Minkowski inequality. When p takes the values 1, 2, or... The following are the simplest scenarios:
[0128] 1-norm:
[0129] 2-norm:
[0130] ∞-norm:
[0131] Here, we use the 3-norm, which is equivalent to p taking the value of 3, resulting in:
[0132]
[0133] Among them, the originality S of the content of the distribution account is equal to the updated account identifier S. account Duplicate content (S) copy and distribution information S p Each parameter is multiplied by its corresponding weight. The weight for each parameter can be set according to the actual situation, which will not be elaborated here.
[0134] 207. When the calculated originality of the content is greater than the preset value, the server distributes the content to be distributed.
[0135] The preset value is set in advance and can be set according to the strategy of the content distribution system. For example, for a content distribution system with abundant original content, the preset value can be set to 1, while for a content distribution system lacking original content (i.e., with a small amount of original content), the preset value can be set to 10. When the originality of the content of the distribution account is greater than the preset value, the server will distribute the content to be distributed.
[0136] For a better understanding of the content distribution scheme of this application, please refer to [link / reference]. Figure 2bThe above diagram illustrates a flowchart of a self-media plagiarism account modeling method and system based on unsupervised machine learning. In the main workflow of self-media production and posting, a plagiarism account identification service is invoked to rank the degree of plagiarism by accounts, and different application strategies are adopted according to different scenarios. Different self-media platforms typically have their own target user groups, and each self-media account entering the system is assigned an account level (e.g., S-5, A-4, B-3, C-2, D-1). In the early stages, this level is usually determined by the platform according to its own operational strategy, forming a whitelist of top accounts. The account level is not static; for top-tier, well-known accounts, it is usually determined by operational strategy, while for mid-tier accounts that can grow rapidly, it is determined by their originality, the distribution of their content on the platform, including user complaints and reports. Here, the existing level results are mainly used to identify whether newly posted account content is plagiarized and the quantifiable degree of plagiarism. The identification results of account plagiarism can be used in the following scenarios: (1) When there are no original accounts, the plagiarism accounts are given lower priority or restricted during the recommendation and distribution process, or even the distribution is canceled; (2) In order to protect the interests of the original account authors, if the original account publishes the same content after the distribution has started, the plagiarism account's content will be withdrawn and the traffic will be given to the original account; (3) The granularity of the incentive for plagiarism accounts will be reduced according to the degree of plagiarism, or the subsidies and incentives for plagiarism accounts will be canceled and the posts of plagiarism accounts will be restricted according to the platform's operation strategy; (4) In the content review process, due to the limited review resources, and in order to make the content of the original top accounts be processed and distributed as soon as possible, plagiarism accounts will be placed at the end of the review scheduling. All of the above scenarios require accurate identification and sorting of plagiarism accounts.
[0137] The following is an introduction Figure 2b The main functions of each service module are as follows:
[0138] I. C-end publishing system or web publishing system (production end) and content consumption end
[0139] (1) PGC or UGC, MCN or PUGC content producers provide text and image content or upload video content, including short videos and mini videos, through mobile terminal or backend interface API system. These are the main sources of content for distribution.
[0140] (2) By communicating with the upstream and downstream content interface servers, first obtain the upload server interface address, and then publish the content;
[0141] (3) As a consumer, communicate with the upstream and downstream content interface server to obtain the index information of the accessed content, and then communicate with the upstream and downstream content interface server and the content export service to directly consume the content. The premise of consumption is to obtain the index of the content through Feeds recommendation distribution.
[0142] (4) Feeds and user click behavior and environment reporting module, collects the user's current network environment and the user's click operation behavior on Feeds intermediate information and the exposure data of Feeds content, and reports it to the statistics reporting interface server;
[0143] (5) If the video content is reported, the video playback time is too long, the caching time is too long, and various interactive behaviors of the content such as forwarding, sharing, collecting, liking, etc.
[0144] II. Uplink and downlink content interface servers and content export services
[0145] (1) Communicate directly with the content production end. The content submitted from the front end usually includes the title, publisher, summary, cover image, and publication time. Store the content in the database.
[0146] (2) The content delivery service and the recommendation distribution system communicate to obtain the recommendation distribution results and send them to the consumer end to be displayed in the user's Feeds list;
[0147] (3) Content export services are usually a set of access services that are geographically located near the user;
[0148] (4) When the content is entered into the database, the initial account level is set according to the source of the publisher's account through the operation configuration. This is mainly related to the operation strategy.
[0149] (5) At the same time, report the posting information of each account to the statistics interface server, including the posting time and content type. Also, store the content tagging information provided by the account owner, such as category, tags, selected cover image, and title, as extended information in the content database.
[0150] III. Content Database
[0151] (1) The core database of content, where all the metadata of the content published by producers is stored. The focus is on the metadata of the content itself, such as size, cover image link, title, publication time, account author, source channel, and entry practice. It also includes the classification of content during the manual review process (including first, second, and third level classifications and tag information. For example, an article explaining Huawei mobile phones has the first level category as technology, the second level category as smartphones, the third level category as domestic mobile phones, and the tag information as Huawei and Mate 30).
[0152] (2) During the manual review process, information in the content database will be read, and the results and status of the manual review will also be sent back to the content database for storage. The results of the manual review are also an important basis for evaluating the efficiency of the algorithm filtering model.
[0153] (3) The content processing in the entire business process mainly includes machine processing and manual review processing. Based on different content tags, the content library is divided into different content pools. The recommendation distribution server and the deduplication server, as well as the content feature modeling service, all need to obtain content from the content database. For example, the image and text deduplication server will load content that has been entered and used in the past period (such as a week) according to business needs. For content that is repeatedly entered into the database, a filter tag will be added and it will no longer be provided to the content recommendation service for output to the user.
[0154] (4) Both the deduplication service and the handling account identification service are machine processing processes, and the processing results are stored in the content database;
[0155] IV. Dispatch Center
[0156] (1) Responsible for the entire scheduling process of content flow, receiving the content into the database through the upstream and downstream content interface servers, and then obtaining the content's metadata from the content database;
[0157] (2) Schedule the deduplication server to mark and filter duplicate entries, and simultaneously synchronize the deduplication flow information to the transport feature mining model module as input;
[0158] (3) Schedule the reposting account identification service, evaluate and calculate the reposting score ranking of each posting account (accounts that have been manually marked and certified as original accounts can be exempted from this process), and use it in subsequent manual review scheduling or distribution process demotion and other practical application scenarios;
[0159] (4) For content that cannot be processed by the machine, such as security issues that require manual review, call the manual review system for manual review.
[0160] V. Manual Review Service System
[0161] (1) It is necessary to read the original information of the video content itself in the content database. This is usually a complex web database-based system, mainly to ensure that the pushed content complies with the access permitted by local laws and policies.
[0162] (2) The content reviewed comes from self-media voluntarily published content and web crawlers obtained from public networks;
[0163] In the technical implementation process of this application, the content reviewed came from self-media proactive releases and web crawlers obtained from public networks. This information was obtained through legal means and underwent strict anonymization processing to ensure it does not contain any personally identifiable data (such as names, ID numbers, contact information, etc.). This information is used solely for technical function implementation and data analysis, complying with the requirements of laws and regulations such as the Personal Information Protection Law and the Cybersecurity Law. Furthermore, through anonymization and de-identification techniques, the risk of privacy leaks is fundamentally eliminated. All data processing flows adhere to the principle of minimum necessity, and data security is ensured through security measures such as encrypted storage and access control isolation.
[0164] (3) The review results are finally written into the content database through the dispatch center;
[0165] VI. Deduplication Service
[0166] (1) Communication with the content scheduling server mainly includes title deduplication, cover image deduplication, content text deduplication, and video and audio fingerprint deduplication. Usually, the title and text of the image and text content are vectorized, and simmhash and BERT text vectors are used for image vector deduplication. For video content, video fingerprints and audio fingerprints are extracted to construct vectors, and then the distance between vectors, such as Euclidean distance, is calculated to determine whether there is a duplicate. This will be introduced separately in the invention and solution, and is not the focus of this invention. This invention mainly uses the result of the judgment here.
[0167] (2) Communicate with the transport feature model mining module to provide raw information on the wastewater discharge flow;
[0168] VII. Statistical Reporting Interface Server
[0169] (1) Receive reports from content consumer users on their current network environment, click behavior of users on Feeds intermediate information, and exposure data of Feeds articles;
[0170] (2) Write the reported statistical data results into the statistical database;
[0171] (3) Accept the original account posting data reported by the content production portal;
[0172] VIII. Mining of Transportation Feature Models
[0173] (1) Based on the specific unsupervised model described above, account conflict features, plagiarism features and verticality features are constructed through content processing.
[0174] (2) The content data for modeling is obtained by reading the content metadata in the content database, statistical database and deduplication service.
[0175] 9. Account Removal Identification Service
[0176] (1) The engineering implementation of the above-mentioned transfer feature model mining results is used to conduct quantitative evaluation of transfer accounts. The core is to realize the fusion of transfer account identification.
[0177] (2) Work with the dispatch center staff to identify and mark the reposting level of the posting account;
[0178] 10. Statistical Databases
[0179] (1) Receive statistical data reports from content consumption terminals to provide data support for subsequent statistical analysis and mining;
[0180] (2) Receive the document production log report from the content production end.
[0181] After obtaining the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account, the server determines the distribution information of the distribution account based on the content type. Next, the server collects a list of reference accounts and the historical content published by each reference account within a historical time period from the content distribution system to which the distribution account belongs. Then, the server extracts the content authentication information of the historical content and the content authentication information of the content to be distributed. Following this, the server updates the account identifier of the distribution account based on the content authentication information of the historical content and the content to be distributed. Finally, the server calculates a score based on the distribution information, the similarity between the content to be distributed and each piece of historical content, and the updated account identifier. The content originality score of the distribution account is calculated. When the calculated content originality score is greater than a preset value, the content to be distributed is distributed. The server provided in this application can update the account identifier of the distribution account based on the content authentication information of historical content and the content authentication information of the content to be distributed. Based on the distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier, the content originality score of the distribution account is calculated. When determining whether the content to be distributed is original, the content authentication information of historical content and the content authentication information of the content to be distributed are taken into account, making the content originality score of the distribution account calculated later more accurate. Moreover, the whole process does not require manual intervention, reducing the waste of human resources, improving the efficiency of content review, and thus improving the efficiency of content distribution.
[0182] To facilitate better implementation of the content distribution method of this application, this application also provides a content distribution device (hereinafter referred to as the distribution device) based on the above-described content distribution device. The meanings of these terms are the same as in the content distribution method described above, and specific implementation details can be found in the description of the method embodiments.
[0183] Please see Figure 3 , Figure 3The schematic diagram of the content distribution device provided in this application shows that the distribution device may include an acquisition module 301, a determination module 302, a collection module 303, a calculation module 304, and a distribution module 305, which may be as follows:
[0184] The acquisition module 301 is used to acquire the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account.
[0185] For example, specifically, the acquisition module 301 can obtain the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account by accessing the network interface.
[0186] The determination module 302 is used to determine the distribution information of the distribution account on the content based on the content type and the number of contents published by the distribution account.
[0187] For example, specifically, module 302 collects the number of articles published by the distribution account. If the distribution account has published a total of 7 articles, of which 2 articles are in the lifestyle category and 5 articles are in the medical category, then the distribution account's content distribution is: lifestyle and medical, with no distribution in other fields.
[0188] The collection module 303 is used to collect a list of reference accounts and the historical content published by each reference account in the historical time period from the content distribution system to which the distribution account belongs.
[0189] Among them, reference accounts refer to the distribution accounts certified by the content distribution system (also known as the content distribution platform), which can include corporate accounts and private accounts. For example, a corporate account can be a distribution account of a news media, and a private account can be a distribution account of a writer, depending on the actual situation. Specifically, the collection module 303 can collect the list of reference accounts and the historical content published by each reference account within a historical time period from the content distribution platform. The historical time period can be the past month, the past year, or the period from the registration of the reference account to the current time, depending on the actual situation.
[0190] The calculation module 304 is used to calculate the originality of the content of the distribution account based on the distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account.
[0191] For the same content distribution platform, each distribution account is assigned an account identifier to identify whether it is a reposting account or an original content account. Furthermore, the account identifier is not static; the calculation module 304 can update the distribution account identifier based on the content to be distributed and previously published historical content. Then, based on distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier, the calculation module 304 calculates the content originality of the distribution account. Optionally, in some embodiments, the calculation module 304 may specifically include:
[0192] The extraction submodule is used to extract the content authentication information of historical content and the content authentication information of content to be distributed, respectively.
[0193] The update submodule is used to update the account identifier of the distribution account based on the content authentication information of the historical content and the content authentication information of the content to be distributed;
[0194] The calculation submodule is used to calculate the originality of the content of the distribution account based on the distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier.
[0195] The content authentication information includes content title information and content identification information. The content title information includes the text length and semantic information of the content title. The content identification information includes watermark information and content publication time information, etc. The content identification information is used to identify the content and facilitate subsequent content identification.
[0196] For the same content distribution platform, each distribution account is assigned an account identifier to distinguish whether it is a reposting account or an original content account. Therefore, the update submodule can update the account identifier of the distribution account based on the content authentication information of historical content and the content to be distributed. For example, if the update submodule determines that the distribution account is a reposting account based on the content authentication information of historical content and the content to be distributed, then it updates the account identifier of that distribution account. The updated account identifier indicates that the distribution account is a reposting account. Specifically, the update submodule can generate the relevance between the distribution account and each reference account in the reference account list based on the content authentication information of historical content and the content to be distributed, and then update the account identifier based on the relevance. Optionally, in some embodiments, the update submodule may specifically include:
[0197] The generation unit is used to generate the relevance between the distribution account and each reference account in the reference account list based on the content authentication information of the historical content and the content authentication information of the content to be distributed;
[0198] The update unit is used to update the account identifier of the distribution account based on the relevance between the distribution account and each reference account in the reference account list.
[0199] Optionally, in some embodiments, the generation unit includes:
[0200] The detection subunit is used to detect whether the content identification information of the historical content is consistent with the content identification information of the content to be distributed.
[0201] The first calculation subunit is used to calculate the similarity between the content title information of each historical content and the content title information of the content to be distributed;
[0202] The generation sub-unit is used to generate the relevance between the distribution account and the reference account based on the similarity between the content title information of historical content and the content title information of the content to be distributed, as well as the results of various detections.
[0203] Optionally, in some embodiments, the generating subunit is specifically used to: fuse the similarity between the content title information and the content title information of the content to be distributed, as well as the detection results, to obtain the relevance between the distribution account and the reference account.
[0204] The calculation module 304 can calculate the originality of the content of the distribution account based on the account authentication level corresponding to the reference account, the similarity between the content to be distributed and each historical content, the distribution information, and the updated account identifier. Optionally, in some embodiments, the calculation module 304 may specifically include:
[0205] The acquisition unit is used to obtain the account authentication level of each reference account in the reference account list in its respective content distribution system.
[0206] The calculation unit is used to calculate the similarity between the content to be distributed and each historical content based on the account authentication level of each reference account in the reference account list.
[0207] The fusion unit is used to merge distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier to obtain the content originality of the distribution account.
[0208] Optionally, in some embodiments, the fusion unit is specifically used to: generate the content duplication rate of the content to be distributed based on the similarity between the content to be distributed and each historical content, and fuse the updated account identifier, distribution information and the content duplication rate of the content to be distributed to obtain the content originality of the distribution account.
[0209] Optionally, in some embodiments, the computing unit may specifically include:
[0210] The sub-unit is determined to determine the weight of each piece of historical content based on the account authentication level of each reference account in the reference account list.
[0211] The second calculation subunit is used to calculate the similarity between the content to be distributed and each historical content based on the weight corresponding to each historical content.
[0212] Optionally, in some embodiments, the second calculation subunit is specifically used to: perform vectorization processing on the content to be distributed and each historical content to obtain a first vector corresponding to the content to be distributed and a second vector corresponding to each historical content; calculate the distance between the first vector and each second vector; and determine the similarity between the content to be distributed and each historical content based on the distance between the first vector and each second vector and the weight corresponding to each historical content.
[0213] The distribution module 305 is used to distribute the content to be distributed when the calculated originality of the content is greater than a preset value.
[0214] The preset value is set in advance, and the distribution module 305 can be set according to the strategy in the content distribution system. For example, for a content distribution system with abundant original content, the preset value can be set to 1, while for a content distribution system lacking original content (i.e., with a small amount of original content), the preset value can be set to 10. When the originality of the content of the distribution account is greater than the preset value, the content to be distributed is distributed.
[0215] After the acquisition module 301 acquires the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account, the determination module 302 determines the distribution information of the distribution account based on the content type. Next, the collection module 303 collects a list of reference accounts and the historical content published by each reference account in the historical time period from the content distribution system to which the distribution account belongs. Finally, the calculation module 304 calculates the content originality of the distribution account based on the distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account. When the calculated content originality is greater than a preset value, the distribution module 305 distributes the content to be distributed. The content distribution device provided in this application can update the account identifier of the distribution account based on the content authentication information of historical content and the content to be distributed. Based on the distribution information, the similarity between the content to be distributed and each historical content, and the updated account identifier, it calculates the content originality of the distribution account. When determining whether the content to be distributed is original, it takes into account the content authentication information of historical content and the content authentication information of the content to be distributed, making the content originality of the distribution account calculated later more accurate. Moreover, the whole process does not require manual intervention, reducing the waste of human resources, improving the efficiency of content review, and thus improving the efficiency of content distribution.
[0216] In addition, this application also provides an electronic device, such as Figure 4 As shown, it illustrates the structural diagram of the electronic device involved in this application, specifically:
[0217] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0218] The processor 401 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, it performs various functions and processes data, thereby performing overall detection of the electronic device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.
[0219] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0220] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0221] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0222] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:
[0223] Obtain the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account. Based on the content type and the number of published contents by the distribution account, determine the distribution information of the distribution account in terms of content. From the content distribution system to which the distribution account belongs, collect the list of reference accounts and the historical content published by each reference account in the historical time period. Based on the distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account, calculate the content originality of the distribution account. When the calculated content originality is greater than a preset value, distribute the content to be distributed.
[0224] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0225] This application, after obtaining the content to be distributed, the distribution account, and the content types corresponding to the content already published by the distribution account, determines the distribution information of the distribution account based on the content types and the number of published contents by the distribution account. Next, it collects a list of reference accounts and the historical content published by each reference account within a historical time period from the content distribution system to which the distribution account belongs. Finally, based on the distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account, it calculates the originality of the content of the distribution account. When the calculated originality is greater than a preset value, the content to be distributed is distributed. The content distribution method provided in this application calculates the originality of the content of the distribution account based on the distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account. When determining whether the content to be distributed is original, it considers not only the similarity between each historical content and the account identifier of the distribution account, but also the distribution information of the distribution account in the content, making the subsequently calculated originality of the content of the distribution account more accurate. Furthermore, the entire process requires no manual intervention, reducing the waste of human resources, improving the efficiency of content review, and thus improving the efficiency of content distribution.
[0226] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0227] To this end, this application provides a storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the content distribution methods provided in this application. For example, the instructions can execute the following steps:
[0228] Obtain the content to be distributed, the distribution account, and the content type corresponding to the content already published by the distribution account. Based on the content type and the number of published contents by the distribution account, determine the distribution information of the distribution account in terms of content. From the content distribution system to which the distribution account belongs, collect the list of reference accounts and the historical content published by each reference account in the historical time period. Based on the distribution information, the similarity between the content to be distributed and each historical content, and the account identifier of the distribution account, calculate the content originality of the distribution account. When the calculated content originality is greater than a preset value, distribute the content to be distributed.
[0229] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0230] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0231] Since the instructions stored in the storage medium can execute the steps of any of the content distribution methods provided in this application, the beneficial effects that any of the content distribution methods provided in this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0232] The above provides a detailed description of a content distribution method, apparatus, electronic device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A content distribution method characterized by, The method comprises the following steps: obtaining to-be-distributed content, a distribution account, and a content type corresponding to content published by the distribution account; determining distribution information of the distribution account on content based on the content type and the number of content published by the distribution account; collecting a reference account list and historical content published by each reference account in the reference account list in a historical time period from a content distribution system to which the distribution account belongs; calculating content originality of the distribution account based on the distribution information, similarity between the to-be-distributed content and each historical content, and an account identifier of the distribution account, comprising: extracting content authentication information of the historical content and content authentication information of the to-be-distributed content respectively; generating relevance between the distribution account and each reference account in the reference account list according to the content authentication information of the historical content and the content authentication information of the to-be-distributed content; updating the account identifier of the distribution account based on the relevance between the distribution account and each reference account in the reference account list; obtaining an account authentication level corresponding to each reference account in the reference account list in the content distribution system to which the reference account belongs; calculating the similarity between the to-be-distributed content and each historical content based on the account authentication level corresponding to each reference account in the reference account list; and fusing the distribution information, the similarity between the to-be-distributed content and each historical content, and the updated account identifier to obtain the content originality of the distribution account; when the calculated content originality is greater than a preset value, distributing the to-be-distributed content.
2. The method of claim 1, wherein, The content authentication information comprises content title information and content identifier information, and the generating of the relevance between the distribution account and each reference account in the reference account list according to the content authentication information of the historical content and the content authentication information of the to-be-distributed content comprises: detecting whether the content identifier information of the historical content is consistent with the content identifier information of the to-be-distributed content; calculating similarity between the content title information of each historical content and the content title information of the to-be-distributed content; generating the relevance between the distribution account and the reference account according to the similarity between the content title information of each historical content and the content title information of the to-be-distributed content and each detection result.
3. The method of claim 2, wherein, The generating of the relevance between the distribution account and the reference account according to the similarity between the content title information of each historical content and the content title information of the to-be-distributed content and each detection result comprises: fusing the similarity between the content title information of each historical content and the content title information of the to-be-distributed content and each detection result to obtain the relevance between the distribution account and the reference account.
4. The method of claim 1, wherein, The fusing of the distribution information, the similarity between the to-be-distributed content and each historical content, and the updated account identifier to obtain the content originality of the distribution account comprises: generating content repetition degree of the to-be-distributed content according to the similarity between the to-be-distributed content and each historical content; fusing the updated account identifier, the distribution information, and the content repetition degree of the to-be-distributed content to obtain the content originality of the distribution account.
5. The method of claim 1, wherein, The similarity between the to-be-distributed content and each historical content is calculated based on the account authentication levels corresponding to each reference account in the reference account list. The weight corresponding to each historical content is determined according to the account authentication levels corresponding to each reference account in the reference account list. The similarity between the to-be-distributed content and each historical content is calculated based on the weight corresponding to each historical content.
6. The method of claim 5, wherein, The similarity between the to-be-distributed content and each historical content is calculated based on the weight corresponding to each historical content. The to-be-distributed content and each historical content are respectively subjected to vectorization processing to obtain a first vector corresponding to the to-be-distributed content and a second vector corresponding to each historical content. The distance between the first vector and each second vector is calculated. The similarity between the to-be-distributed content and each historical content is determined based on the distance between the first vector and each second vector and the weight corresponding to each historical content.
7. A content distribution apparatus characterized by comprising: The method comprises the following steps: An acquisition module is configured to acquire to-be-distributed content, a distribution account, and a content type corresponding to content published by the distribution account; A determination module is configured to determine distribution information of the distribution account on content based on the content type and the number of content published by the distribution account; A collection module is configured to collect a reference account list and historical content published by each reference account in the reference account list within a historical time period from a content distribution system to which the distribution account belongs; A calculation module is configured to calculate content originality of the distribution account based on the distribution information, the similarity between the to-be-distributed content and each historical content, and an account identifier of the distribution account, comprising: extracting content authentication information of the historical content and content authentication information of the to-be-distributed content respectively; generating a relevance between the distribution account and each reference account in the reference account list according to the content authentication information of the historical content and the content authentication information of the to-be-distributed content; updating the account identifier of the distribution account based on the relevance between the distribution account and each reference account in the reference account list; acquiring an account authentication level corresponding to each reference account in the reference account list in the content distribution system to which the reference account belongs; calculating the similarity between the to-be-distributed content and each historical content based on the account authentication levels corresponding to each reference account in the reference account list; and fusing the distribution information, the similarity between the to-be-distributed content and each historical content, and the updated account identifier to obtain the content originality of the distribution account; A distribution module is configured to distribute the to-be-distributed content when the calculated content originality is greater than a preset value.
8. The apparatus of claim 7, wherein, The device further comprises: A detection subunit is configured to detect whether content identifier information of the historical content and content identifier information of the to-be-distributed content are consistent. A first calculation subunit is configured to calculate the similarity between content title information of each historical content and content title information of the to-be-distributed content. A generation subunit is configured to generate a relevance between the distribution account and a reference account according to the similarity between the content title information of the historical content and the content title information of the to-be-distributed content and each detection result.
9. The apparatus of claim 8, wherein, The generation subunit is specifically configured to: The similarity between the content title information of the content to be distributed and the content title information of the content to be distributed and each detection result are fused to obtain a correlation between the distribution account and the reference account.
10. The apparatus of claim 7, wherein, The computing module is further configured to: generate a content repetition degree of the content to be distributed according to the similarity between the content to be distributed and each historical content; fuse the updated account identifier, the distribution information, and the content repetition degree of the content to be distributed to obtain a content originality degree of the distribution account.
11. The apparatus of claim 7, wherein, The apparatus further includes: a determining subunit configured to determine a weight corresponding to each historical content according to an account authentication level corresponding to each reference account in the reference account list; a second computing subunit configured to compute the similarity between the content to be distributed and each historical content based on the weight corresponding to each historical content.
12. The apparatus of claim 11, wherein, The second computing subunit is specifically configured to: vectorize the content to be distributed and each historical content respectively to obtain a first vector corresponding to the content to be distributed and a second vector corresponding to each historical content; compute the distance between the first vector and each second vector respectively; determine the similarity between the content to be distributed and each historical content based on the distance between the first vector and each second vector and the weight corresponding to each historical content.
13. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, The processor implements the steps of the content distribution method of any one of claims 1-6 when executing the program.
14. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program, when executed by a processor, implements the steps of the content distribution method of any one of claims 1-6.
Citation Information
Patent Citations
Data excavating method and system for high quality user generation content (UGC)
CN103914491A
Original data protection method, medium, device and computing device
CN108959515A