A Classification Method and Device for Content Publishing Entities

By performing word segmentation, feature vector transformation, aggregation and normalization of content description information, the problems of low accuracy and high cost of content publishing subject classification are solved, and efficient and accurate classification of content publishing subject classification is achieved.

CN114792096BActive Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110105908.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-26
Publication Date
2025-07-08
Estimated Expiration
2041-01-26

AI Technical Summary

Technical Problem

In the prior art, the classification accuracy of the content publishing subject is low, difficult to classify, and high classification cost. Especially when the content contains a large amount of video content, the accuracy of the classification model is affected.

Method used

By obtaining content description information, performing word segmentation and converting it into feature vectors. After aggregation and normalization, the subject vector is determined and the preset subject vector samples are matched to obtain category labels to realize the classification of the content publishing subjects.

Benefits of technology

It improves the classification accuracy and efficiency of the content publishing subjects, while reducing classification costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114792096B_ABST
    Figure CN114792096B_ABST
Patent Text Reader

Abstract

The present application discloses a classification method and device for content publishing entities. The method includes: obtaining content description information corresponding to the content published by the content publishing entity; performing word segmentation on the obtained content description information to obtain the words that make up the content description information, and converting the words into feature vectors of a predetermined dimension; performing aggregation processing on the vectors of the target dimension among the feature vectors corresponding to the words that make up the content description information to obtain a topic vector of the content description information; performing normalization processing on the vectors of the target dimension among the obtained topic vectors to obtain an entity vector of the content publishing entity; determining a preset entity vector sample that matches the entity vector, and obtaining an entity category label corresponding to the preset entity vector sample to classify the content publishing entity. The present application improves the classification accuracy and efficiency of content publishing entities, and at the same time reduces the classification cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of content classification, and particularly relates to a method and device for classifying content publishing entities. Background Art

[0002] A content publishing entity is the entity that carries the content published by a user. For example, a video account is a content publishing entity for a user to publish video content. Currently, the number of content publishing entities on various platforms is huge, and it is very important to classify these content publishing entities in a timely manner. For example, the classified content publishing entities can be used to train a supervised classification model, and then the classification model can be used to automatically classify content publishing entities, or reliable management of content publishing entities can be carried out.

[0003] In the prior art, the category of a content publishing entity is usually determined by the content published therein. However, video content is usually published in a content publishing entity, and at this time, a large amount of content information is carried by video frames, resulting in low classification accuracy and difficulty in classifying content publishing entities. Taking the solution of training a supervised classification model as an example, the content publishing entity samples used to train the supervised classification model usually require manual annotation of the categories of content publishing entities, with high classification costs. Moreover, video content is usually published in a content publishing entity, resulting in low classification accuracy of content publishing entities and affecting the classification accuracy of the classification model. Summary of the Invention

[0004] Embodiments of this application provide a method and device for classifying content publishing entities, aiming to improve the classification accuracy and efficiency of content publishing entities, and at the same time reduce classification costs.

[0005] Embodiments of this application provide the following technical solutions:

[0006] According to an embodiment of this application, a method for classifying a content publishing entity includes: obtaining content description information corresponding to the content published in the content publishing entity, where the content description information is calibrated by the publishing user of the content; performing word segmentation on the obtained content description information to obtain the words that make up the content description information, and converting the words into feature vectors of a predetermined dimension; performing aggregation processing on the vectors of a target dimension in the feature vectors corresponding to the words that make up the content description information to obtain a topic vector of the content description information; performing normalization processing on the vectors of a target dimension in the topic vectors corresponding to the obtained content description information to obtain a main body vector of the content publishing entity; determining a preset main body vector sample that matches the main body vector, and obtaining a main body category label corresponding to the preset main body vector sample as the category label of the content publishing entity, so as to classify the content publishing entity.

[0007] According to an embodiment of the present application, a method for obtaining training samples of a classification model for content publishing entities, the training samples being used to train the classification model, and the training of the classification model taking target information in the content publishing entity as input and the category label corresponding to the content publishing entity as output; the method includes: determining a plurality of content publishing entities included in the target publishing platform; respectively obtaining the category label corresponding to each content publishing entity based on the method in the foregoing embodiment; obtaining the target information and the corresponding category label in each content publishing entity as the positive training samples of the classification model.

[0008] According to an embodiment of the present application, a classification device for content publishing entities includes: a first acquisition module, configured to acquire content description information corresponding to the content published by the content publishing entity, where the content description information is calibrated by the publishing user of the content; a conversion module, configured to perform word segmentation on the acquired content description information to obtain the words constituting the content description information, and convert the words into feature vectors of a predetermined dimension; an aggregation module, configured to perform an aggregation process on the vectors of a target dimension in the feature vectors corresponding to the words constituting the content description information to obtain a topic vector of the content description information; a normalization module, configured to perform a normalization process on the vectors of a target dimension in the topic vector corresponding to the acquired content description information to obtain a main body vector of the content publishing entity; a classification module, configured to determine a preset main body vector sample that matches the main body vector, and acquire the main body category label corresponding to the preset main body vector sample as the category label of the content publishing entity, so as to classify the content publishing entity.

[0009] In some embodiments of the present application, the content description information is an information label with a target symbol; the first acquisition module includes: a label acquisition unit, configured to acquire the information label corresponding to the content published by the content publishing entity to obtain all the information labels included in the content publishing entity; a first frequency calculation unit, configured to calculate the first frequency of each information label in the information labels included in the content publishing entity; a score determination unit, configured to determine the importance score of each information label according to the first frequency corresponding to each information label; a screening unit, configured to use the plurality of information labels with the highest importance scores in the content publishing entity as the acquired content description information.

[0010] In some embodiments of the present application, the content publishing entity is from a target publishing platform, and the target publishing platform further includes other content publishing entities; the score determination unit includes: a second frequency calculation subunit, configured to calculate a second frequency of each of the information tags in the other content publishing entities in the target publishing platform; a score determination subunit, configured to calculate an importance score of each of the information tags according to the first frequency and the second frequency, wherein the importance score is directly proportional to the first frequency and inversely proportional to the second frequency.

[0011] In some embodiments of the present application, the second frequency calculation subunit is configured to: obtain a first number of all content publishing entities included in the target publishing platform; sequentially determine a second number of content publishing entities containing the information tag from all content publishing entities included in the target publishing platform, to obtain the second number corresponding to each of the information tags; take the logarithm of the quotient obtained by dividing the first number by the second number corresponding to each of the information tags as the second frequency corresponding to each of the information tags; the score determination subunit is configured to: calculate the product of the first frequency and the second frequency corresponding to each of the information tags as the importance score of each of the information tags.

[0012] In some embodiments of the present application, the first frequency calculation unit includes: a similarity calculation subunit, configured to calculate the similarity between the information tags included in the content publishing entity; a first frequency calculation subunit, configured to determine the number of occurrences of each information tag in the information tags included in the content publishing entity by determining that the information tags with a similarity greater than a predetermined threshold are the same information tags, and calculate the first frequency.

[0013] In some embodiments of the present application, the aggregation module includes: an accumulation unit, configured to accumulate the feature vectors with the same dimension in the feature vectors corresponding to the words constituting the content description information to obtain a topic vector of the content description information; the normalization module includes: an averaging unit, configured to take the average value of the feature vectors with the same dimension in the topic vector corresponding to the obtained content description information to obtain a subject vector of the content publishing entity.

[0014] In some embodiments of the present application, the conversion module includes: a word segmentation unit, configured to segment the obtained content description information using a target word segmenter to obtain the words constituting the content description information; a vector conversion unit, configured to input the words into a pre-trained vector conversion model to obtain a feature vector with a predetermined dimension corresponding to the words output by the vector conversion model, wherein the word samples for training the vector conversion model are obtained by segmenting content description information samples using the target word segmenter.

[0015] In some embodiments of the present application, the first acquisition module includes: a content type acquisition unit configured to acquire the content type corresponding to the content published by the content publishing entity; a ratio determination unit configured to determine the proportion of target type content in the content type, where the target type content is other types of content other than text type content; and a determination and acquisition unit configured to acquire the content description information corresponding to the content published by the content publishing entity when the proportion of the target type content is higher than a predetermined threshold.

[0016] The proportion of video content in the content of the content publishing entity is higher than a predetermined threshold.

[0017] According to an embodiment of the present application, an apparatus for acquiring training samples of a classification model of a content publishing entity, where the training samples are used to train the classification model, and the training of the classification model takes the target information in the content publishing entity as input and the category label corresponding to the content publishing entity as output; the apparatus includes: a main body determination module configured to determine a plurality of content publishing entities included in a target publishing platform; a classification apparatus for the content publishing entity configured to respectively acquire the category label corresponding to each content publishing entity; and a training sample acquisition module configured to acquire the target information and the corresponding category label in each content publishing entity as the positive training sample of the classification model.

[0018] According to another embodiment of the present application, an electronic device may include: a memory storing computer-readable instructions; and a processor configured to read the computer-readable instructions stored in the memory to execute the exception testing method described in the embodiments of the present application.

[0019] According to another embodiment of the present application, a storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the exception testing method described in the embodiments of the present application.

[0020] According to another embodiment of the present application, a computer program product or a computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the exception testing method provided in various alternative implementation manners described in the embodiments of the present application.

[0021] In an embodiment of the present application, first, content description information corresponding to the content published by a content publishing entity is obtained, and the content description information is marked by the content publisher; then, the obtained content description information is segmented to obtain the words that make up the content description information, and the words are converted into feature vectors of a predetermined dimension; then, among the feature vectors corresponding to the words that make up the content description information, the vectors of the target dimension are aggregated to obtain a topic vector of the content description information; then, among the topic vectors corresponding to the obtained content description information, the vectors of the target dimension are normalized to obtain a subject vector of the content publishing entity; finally, a preset subject vector sample that matches the subject vector is determined, and a subject category label corresponding to the preset subject vector sample is obtained as the category label of the content publishing entity to classify the content publishing entity.

[0022] Furthermore, by obtaining the limited content description information marked by the user, segmenting the content description information and converting it into feature vectors of a predetermined dimension, a topic vector representing the connotation of the content description information can be obtained through aggregation processing, and then a subject vector effectively representing the connotation of the content publishing entity can be obtained through normalization processing. Furthermore, by matching the preset subject vector sample, the content publishing entity can be classified accurately, efficiently and at low cost. It effectively improves the classification accuracy and efficiency of the content publishing entity, and at the same time reduces the classification cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0024] Figure 1 A schematic diagram of a system to which the embodiments of the present application can be applied is shown.

[0025] Figure 2 A flowchart of a method for classifying a content publishing entity according to an embodiment of the present application is shown.

[0026] Figure 3 A flowchart of a method for obtaining content description information according to an embodiment of the present application is shown.

[0027] Figure 4 A flowchart of a method for determining an importance score according to an embodiment of the present application is shown.

[0028] Figure 5 A flowchart of a method for calculating a second frequency according to an embodiment of the present application is shown.

[0029] Figure 6 The flowchart of a method for calculating a first frequency according to an embodiment of the present application is shown.

[0030] Figure 7 The flowchart of a method for transforming a feature vector according to an embodiment of the present application is shown.

[0031] Figure 8 The flowchart of a method for obtaining a theme vector according to an embodiment of the present application is shown.

[0032] Figure 9 The flowchart of a method for obtaining a training sample according to an embodiment of the present application is shown.

[0033] Figure 10 The terminal interface diagram for viewing a video number on the WeChat Video Number platform in a scenario is shown.

[0034] Figure 11 Shown is Figure 10 The terminal interface diagram for opening the target video content in a video number in a scenario.

[0035] Figure 12 The block diagram of a classification device for a content publishing entity according to an embodiment of the present application is shown.

[0036] Figure 13 The block diagram of a training sample acquisition device according to an embodiment of the present application is shown.

[0037] Figure 14 The block diagram of an electronic device according to an embodiment of the present application is shown. Detailed implementation manners

[0038] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0039] In the following description, specific embodiments of the present application will be described with reference to steps and symbols executed by one or more computers, unless otherwise stated. Therefore, these steps and operations will be referred to several times as being executed by a computer, and the computer execution referred to herein includes the operation of a computer processing unit that represents data in a structured form. This operation transforms the data or maintains it at a location in the computer's memory system, which can be reconfigured or otherwise changed in a manner well known to those skilled in the art to alter the operation of the computer. The data structure in which the data is maintained is a physical location in the memory that has specific characteristics defined by the data format. However, the principles of the present application are described in the above text, which does not represent a limitation, and those skilled in the art will understand that the various steps and operations described below can also be implemented in hardware.

[0040] Figure 1 FIG. shows a schematic diagram of a system 100 to which embodiments of the present application can be applied. As Figure 1 shown, the system 100 may include a server 101 and a terminal cluster 102. The server 101 may receive information published by a content publishing entity on a terminal in the terminal cluster 102, such as content and content description information corresponding to the content; a user may publish content through the content publishing entity based on a terminal in the terminal cluster 102.

[0041] The server 101 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0042] In one embodiment, the server 101 may provide artificial intelligence cloud services, such as providing artificial intelligence cloud services for a massively multiplayer online role-playing game (MMORPG). The so-called artificial intelligence cloud service is generally also referred to as AIaaS (AI as a Service, Chinese for "AI as a Service"). This is a current mainstream service mode of an artificial intelligence platform. Specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service mode is similar to opening an AI-themed mall: all developers can access and use one or more artificial intelligence services provided by the platform through the API interface, and some senior developers can also use the AI framework and AI infrastructure provided by the platform to deploy and operate their own exclusive cloud artificial intelligence services.

[0043] The terminals in the terminal cluster 102 may be edge devices, such as smart phones, computers, etc.

[0044] Among them, the terminal 102 and the server 101 can be directly or indirectly connected through a wireless communication method, and this application does not make special restrictions here.

[0045] In an implementation manner of this example, the server 101 can obtain content description information corresponding to the content published by the content publishing entity, and this content description information is marked by the content publisher; perform word segmentation on the obtained content description information to obtain the words that make up the content description information, and convert the words into feature vectors of a predetermined dimension; perform aggregation processing on the vectors of the target dimension among the feature vectors corresponding to the words that make up the content description information to obtain a topic vector of the content description information; then, perform normalization processing on the vectors of the target dimension among the topic vectors corresponding to the obtained content description information to obtain a main body vector of the content publishing entity; determine a preset main body vector sample that matches the main body vector, and obtain a main body category label corresponding to the preset main body vector sample as the category label of the content publishing entity, so as to classify the content publishing entity.

[0046] Figure 2 Schematically shows a flowchart of a method for classifying a content publishing entity according to an embodiment of the present application. The execution entity of this method for classifying a content publishing entity can be an electronic device with computing and processing capabilities, such as Figure 1 the server 101 or the terminal 102 shown in

[0047] As Figure 2 shown, this method for classifying a content publishing entity may include steps S210 to S250.

[0048] Step S210, obtain content description information corresponding to the content published by the content publishing entity, and this content description information is marked by the content publisher;

[0049] Step S220, perform word segmentation on the obtained content description information to obtain the words that make up the content description information, and convert the words into feature vectors of a predetermined dimension;

[0050] Step S230, perform aggregation processing on the vectors of the target dimension among the feature vectors corresponding to the words that make up the content description information to obtain a topic vector of the content description information;

[0051] Step S240, perform normalization processing on the vectors of the target dimension among the topic vectors corresponding to the obtained content description information to obtain a main body vector of the content publishing entity;

[0052] Step S250: Determine a preset main body vector sample that matches the main body vector, and obtain the main body category label corresponding to the preset main body vector sample as the category label of the content publishing main body, so as to classify the content publishing main body.

[0053] The following describes the specific processes of the respective steps performed when classifying the content publishing main body.

[0054] In step S210, obtain the content description information corresponding to the content published by the content publishing main body, where the content description information is calibrated by the content publishing user.

[0055] In the implementation manner of this example, the content publishing main body may be a carrier for users to publish content such as a video account, a WeChat Moments, and a Douyin account. The content published by the content publishing main body may include video content, picture content, and plain text content, etc. For example, the content published in a video account is usually video content. It can be understood that the number of content published by the content publishing main body may be 0 or at least 1. For example, video account A publishes 0 short videos, and video account B publishes 20 short videos; the content description information corresponding to each content may be 0 or at least one. For example, content A1 is calibrated with 2 content description information, and content A2 is not calibrated with content description information.

[0056] The content description information is calibrated by the content publishing user, such as the title of the content or the information label calibrated by the target symbol, etc. The applicant found that the description information calibrated by the user for the content usually includes the summary topic classification label, the extended reading label, and the content points that guide the user to pay attention to of the content they publish. Therefore, the content description information contained in the content publishing main body can overall effectively reflect the category characteristics of the content publishing main body.

[0057] When determining the content description information corresponding to the content published by the content publishing main body, it is possible to first determine the type of the content publishing main body (this type includes, for example, a video account or a WeChat Moments, etc.) and / or the type of the content published by the content publishing main body (this type includes, for example, video content or picture content, etc.); then, according to the corresponding relationship between the preset type and the manifestation form of the content description information (such as the title or the information calibrated by the target symbol), determine the manifestation form of the content description information that conforms to the type of the content publishing main body and / or the type of the content; furthermore, the content description information in the corresponding manifestation form can be crawled from the content publishing main body according to the determined manifestation form of the content description information. For example, crawl the information calibrated by the target symbol.

[0058] In one embodiment, the content description information is an information label with a target symbol; refer to Figure 3 , step S210, obtaining the content description information corresponding to the content published by the content publishing main body includes:

[0059] Step S310: Obtain the information tags corresponding to the content published by the content publishing entity, and obtain all the information tags included in the content publishing entity;

[0060] Step S320: Calculate the first frequency of each information tag appearing in the information tags included in the content publishing entity;

[0061] Step S330: Determine the importance score of each information tag according to the first frequency corresponding to each information tag;

[0062] Step S340: Use multiple information tags with the highest importance scores in the content publishing entity as the obtained content description information.

[0063] For an information tag with a target symbol, such as an information tag with symbol @ or an information tag with symbol #, it can be located by retrieving the target symbol, and then the descriptive text information associated with the target symbol can be crawled.

[0064] For example, the number of all information tags that can be obtained in each content publishing entity is 100. When the number of A tags among the 100 information tags is 5, the first frequency of the A tag can be obtained by dividing the number of A tags by the number of all information tags, or the number of A tags being 5 (i.e., the word frequency) can be used as the first frequency.

[0065] The first frequency corresponding to each information tag can reflect the importance degree of the information tag in the content publishing entity, that is, the higher the frequency of an information tag appears, the more important the tag is. In one example, the first frequency can be directly used as the importance score. In another example, the first frequency of each information tag can be multiplied by a preset weight coefficient to obtain the importance score, which can be obtained from a preset weight coefficient table corresponding to different information tags. This weight coefficient can represent the importance degree coefficient of the information tag on the target platform.

[0066] Finally, use multiple information tags with the highest importance scores in the content publishing entity as the obtained content description information. Several of the most important tags in the content publishing entity can be screened out from all the information tags to further ensure the classification accuracy.

[0067] In one embodiment, the content publishing entity is from the target publishing platform, and the target publishing platform also includes other content publishing entities; Refer to Figure 4 , Step S330: Determine the importance score of each information tag according to the first frequency corresponding to each information tag, including:

[0068] Step S410: Calculate the second frequency of each information tag appearing in other content publishing entities in the target publishing platform;

[0069] Step S420: Calculate the importance score of each information tag according to the first frequency and the second frequency. Here, the importance score is directly proportional to the first frequency and inversely proportional to the second frequency.

[0070] The target publishing platform is, for example, the WeChat Video Account platform or the Douyin Account platform, etc.; the target publishing platform may include at least one content publishing entity. The other content publishing entities are all the publishing entities in the target publishing platform except the aforementioned content publishing entity.

[0071] In one example, an information tag set can also be obtained from the other content publishing entities. Furthermore, the ratio of the number of occurrences of each information tag in this information tag set to the total number of all tags in the information tag set can be calculated to obtain the second frequency of each information tag in the other content publishing entities in the target publishing platform.

[0072] Then, the importance score of each information tag can be calculated according to the first frequency and the second frequency. For example, the ratio obtained by dividing the first frequency by the second frequency is used as the importance score. Here, the importance score is directly proportional to the first frequency and inversely proportional to the second frequency, and a score representing the uniqueness and importance of each information tag in the content publishing entity can be obtained, that is, a score representing the core value. That is, the higher the importance score, the stronger the ability of the corresponding information tag to distinguish the theme of the content publishing entity, and the more suitable it is to be used for the theme analysis of the content publishing entity in the subsequent steps.

[0073] In one embodiment, the first frequency is the word frequency; refer to Figure 5 , step S410: Calculate the second frequency of each information tag in the other content publishing entities in the target publishing platform, including:

[0074] Step S510: Obtain the first number of all content publishing entities included in the target publishing platform;

[0075] Step S520: Sequentially determine the second number of the content publishing entities containing the information tag from all the content publishing entities included in the target publishing platform to obtain the second number corresponding to each information tag;

[0076] Step S530: Take the logarithm of the quotient obtained by dividing the first number by the second number corresponding to each information tag as the second frequency corresponding to each information tag;

[0077] Step S420: Calculate the importance score of each information tag according to the first frequency and the second frequency, including:

[0078] Step S540: Calculate the product of the first frequency and the second frequency corresponding to each information tag as the importance score of each information tag.

[0079] The first number of all content publishing entities included in the target publishing platform. For example, the number of all video account entities on the video account platform is 50 million (i.e., the first number).

[0080] From all content publishing entities, sequentially determine the second number of content publishing entities that contain the information tag. For example, for the ASD tag, determine that the number of content publishing entities that have appeared this ASD tag among all content publishing entities is 10 million (i.e., the second number). The second number corresponding to each information tag can be determined sequentially.

[0081] Take the logarithm of the quotient obtained by dividing the first number by the second number corresponding to each information tag as the second frequency corresponding to each information tag. For example, the second frequency P corresponding to the ASD tag = log(10 million / 50 million).

[0082] In this way, each content publishing entity can be represented as a document, each information tag can be represented as a word in the document, and the second frequency is the inverse document frequency of each word (the frequency of each word appearing in other documents); at the same time, the first frequency is the word frequency, which can represent the word frequency of each word (the frequency of each word appearing in this document); furthermore, calculate the product of the first frequency and the second frequency corresponding to each information tag, that is, calculate the product of the word frequency and the inverse document frequency of each word, and the obtained importance score can effectively characterize the importance degree of each word (information tag).

[0083] In one embodiment, refer to Figure 6 , step S320: Calculate the first frequency of each information tag appearing in the information tags included in the content publishing entity, including:

[0084] Step S610: Calculate the similarity between the information tags included in the content publishing entity;

[0085] Step S620: Determine the number of times each information tag appears in the information tags included in the content publishing entity by determining the information tags with a similarity greater than the predetermined threshold as the same information tags, and calculate the first frequency.

[0086] In an example, the calculation of the first frequency of each information tag appearing in the information tags included in the content publishing entity is: taking the ratio of the number of times each tag appears (i.e., the number of the same tags) to the total number of information tags as the first frequency. The number of appearances calculated by determining the information tags with a similarity greater than the predetermined threshold as the same information tags can ensure the accuracy of the first frequency calculation.

[0087] In one example, when the first frequency of each information tag appearing in the information tags included in the content publishing entity is calculated as: directly using the number of times each tag appears (i.e., the number of identical tags), that is, the word frequency, as the first frequency, it is possible to avoid the first frequency being biased towards the content publishing entity containing rich information tags. For example, when directly considering tags with a similarity of 100% as identical tags, it is possible that the word frequency of a certain tag appears more likely in the content publishing entity containing rich information tags, making the first frequency larger, but the importance of this word is not necessarily so. By determining information tags with a similarity greater than a predetermined threshold as identical information tags, the reliability of the word frequency of information tags in the content publishing entity containing fewer information tags can be improved.

[0088] In one embodiment, step S210 of obtaining the content description information corresponding to the content published by the content publishing entity (before step S310) includes:

[0089] Obtaining the content type corresponding to the content published by the content publishing entity; determining the proportion of the target type content in the obtained content type, where the target type content is other types of content except text type content; when the proportion of the target type content is higher than a predetermined threshold, obtaining the content description information corresponding to the content published by the content publishing entity.

[0090] The content types corresponding to the content in the content publishing entity may include types such as video content, text content, and picture content. Generally, other types of content except text type content are carried by non-text means such as video frames or pictures.

[0091] When obtaining the content description information, by first determining the proportion of the target type content, and when the proportion of the target type content is higher than a predetermined threshold, obtaining the content description information corresponding to the content published by the content publishing entity, and then the entity can be classified in subsequent steps based on the content description information. In this way, entities containing less text information can be screened out and classified based on the embodiments of the present application, effectively ensuring the classification efficiency. For content publishing entities containing more text information (where the proportion of the target type content is lower than the predetermined threshold), they can be classified based on other existing classification methods. Among them, the predetermined threshold can be set according to requirements, such as 40% or the like.

[0092] It can be understood that in one embodiment, when the proportion of the target type content is lower than the predetermined threshold, it is also possible to obtain the content description information corresponding to the content published by the content publishing entity and classify the entity in subsequent steps based on the content description information.

[0093] In step S220, the acquired content description information is segmented to obtain words constituting the content description information, and the words are converted into feature vectors of a predetermined dimension.

[0094] In the implementation of this example, the acquired content description information may be segmented using a word segmenter to obtain a plurality of words constituting the content description information.

[0095] The words may be converted into feature vectors of a predetermined dimension using a vector conversion model, and each word constituting the content description information may be converted into a feature vector of a predetermined dimension. For example, each word may be converted into a 200-dimensional feature vector.

[0096] After obtaining the words constituting the content description information, some useless words may be eliminated, for example, target redundant words such as "的" may be eliminated.

[0097] In one embodiment, see Figure 7 , step S220, segmenting the acquired content description information to obtain the words constituting the content description information, and converting the words into feature vectors of a predetermined dimension, including:

[0098] Step S710, using the target word segmenter to segment the acquired content description information to obtain words constituting the content description information;

[0099] Step S720, input the words into a pre-trained vector conversion model to obtain a feature vector of a predetermined dimension corresponding to the words output by the vector conversion model, wherein the word samples of the training vector conversion model are obtained by segmenting the content description information samples through the target word segmenter.

[0100] When training the vector conversion model, we collect content description information samples, use the target word segmenter to segment words, and obtain word samples as the input of the model. The feature vectors corresponding to the word samples are used as output, and a vector conversion model that meets the standards is trained.

[0101] The target word segmenter is used to segment the acquired content description information. After obtaining the words that constitute the content description information, the words are input into the vector conversion model to effectively ensure the accuracy of the vector conversion.

[0102] In one embodiment, the vector conversion model is the Word2vec model. The Word2vec model is a word vector calculation model. Under the assumption of the bag-of-words model in word2vec, the order of words is not important. The Word2vec model can be used to map each word to a vector, which can represent the relationship between words. This vector is the hidden layer of the neural network. Through the Word2vec model, words can be converted into computable and structured vectors, which can effectively ensure the reliability of the aggregation / normalization process in subsequent steps.

[0103] In step S230, among the feature vectors corresponding to the words that make up the content description information, the vectors of the target dimension are aggregated to obtain the theme vector of the content description information.

[0104] In the implementation manner of this example, when there are multiple (at least two) words that make up the content description information, the corresponding feature vectors of the words that make up the content description information are also multiple, and the feature vectors are composed of vectors of multiple dimensions. It can be understood that the words that make up the content description information can be 1.

[0105] In one implementation manner, the vectors of the target dimension in all feature vectors can be vectors of all or some of the same dimensions. For example, the words that make up the content description information include A and B, and the corresponding feature vectors include: A feature vector (predetermined dimension of 200 dimensions), B feature vector (predetermined dimension of 200 dimensions); at this time, the first dimension in the A feature vector and the first dimension in the B feature vector are of the same dimension, and so on until the 200th dimension are all vectors of the same dimension, and by analogy until the 150th dimension are vectors of some of the same dimensions.

[0106] In another implementation manner, the vectors of the target dimension in all feature vectors can be the vectors of the corresponding specific dimensions in each feature vector. For example, the words that make up the content description information include C and D, and the corresponding feature vectors include: C feature vector (predetermined dimension of 200 dimensions), D feature vector (predetermined dimension of 200 dimensions); the vectors of the target dimension of the C feature vector and the D feature vector can be: the vector of the ith dimension in the C feature vector, the vector of the (i + 1)th dimension in the D feature vector.

[0107] The aggregation process of the vectors of the target dimension can be an accumulation process, a multiplication process or a vector concatenation process. By aggregating the vectors, the applicant finds that the theme vector obtained through the aggregation process of the vectors can reflect the core connotation of the content description information. It can be understood that when there is 1 word that makes up the content description information, no aggregation process is performed.

[0108] In one embodiment, refer to Figure 8, step S230, perform an aggregation process on the vectors of the target dimension in the feature vectors corresponding to the words that make up the content description information to obtain the topic vector of the content description information, including:

[0109] Step S810, add up the feature vectors of the same dimension in the feature vectors corresponding to the words that make up the content description information to obtain the topic vector of the content description information.

[0110] Add up the feature vectors of the same dimension in the feature vectors corresponding to the words that make up the content description information to obtain the topic vector of the content description information. The dimension of the topic vector is the same as the dimension of the feature vector corresponding to the word. The applicant finds that the obtained topic vector can effectively reflect the core connotation of the content description information.

[0111] For example, the words that make up the content description information include from A to N, and the corresponding feature vectors include: A feature vector (predetermined dimension 200 dimensions) to N feature vectors (predetermined dimension 200 dimensions); add up the feature vectors of the same dimension, that is, add up the vectors of the first dimension in the A feature vector, the first dimension in the B feature vector until the first dimension in the N feature vector, and the obtained accumulated vector is used as the vector of the first dimension of the topic vector, and so on until the 200th dimension is accumulated respectively.

[0112] In step S240, perform a normalization process on the vectors of the target dimension in the topic vector corresponding to the obtained content description information to obtain the subject vector of the content publishing entity.

[0113] In the implementation manner of this example, at least one content description information is obtained. Each content description information obtains the corresponding topic vector based on the foregoing steps. Then, perform a normalization process on the vectors of the target dimension in each topic vector to obtain the subject vector of the content publishing entity.

[0114] In one implementation manner, the vectors of the target dimension in all feature vectors can be vectors of all or part of the same dimension. For example, the words that make up the content description information include A and B, and the corresponding feature vectors include: A feature vector (predetermined dimension 200 dimensions), B feature vector (predetermined dimension 200 dimensions); at this time, the first dimension in the A feature vector and the first dimension in the B feature vector are of the same dimension, and so on until the 200th dimension is the vector of all the same dimensions, and so on until the 150th dimension is the vector of part of the same dimensions.

[0115] In another embodiment, the vectors of the target dimension in all feature vectors can be the vectors of the corresponding specific dimensions in each feature vector. For example, the words constituting the content description information include C and D, and the corresponding feature vectors include: C feature vector (predetermined dimension of 200 dimensions), D feature vector (predetermined dimension of 200 dimensions); the vectors of the target dimension of the C feature vector and the D feature vector can be: the vector of the i-th dimension in the C feature vector, and the vector of the (i + 1)-th dimension in the D feature vector.

[0116] The normalization process of the vectors of the target dimension can be the process of taking the average value after accumulation, or the process of taking the n-th root after multiplication. By normalizing the vectors, the applicant finds that the main body vector obtained by normalizing the theme vector obtained through the aggregation process can reflect the core connotation of the content publishing entity. For example, for two theme vectors corresponding to the content description information of "father" and "mother", the main body vector obtained after normalization will characterize the main body category as "child". It can be understood that when there is only one word constituting the content description information, no normalization process is performed.

[0117] In one embodiment, continue to refer to Figure 8 , step S240, in the theme vectors corresponding to the obtained content description information, normalize the vectors of the target dimension to obtain the main body vector of the content publishing entity, including:

[0118] Step S820, in the theme vectors corresponding to the obtained content description information, take the average value of the feature vectors of the same dimension to obtain the main body vector of the content publishing entity.

[0119] Take the average value of the feature vectors of the same dimension in the theme vectors corresponding to the obtained content description information to obtain the main body vector of the content publishing entity. The dimension of the main body vector is the same as that of the theme vector. The applicant finds that the obtained main body vector can effectively reflect the core connotation of the content publishing entity.

[0120] For example, the obtained content description information includes M and N, and the corresponding theme vectors include: M theme vector (predetermined dimension of 200 dimensions) to N theme vector (predetermined dimension of 200 dimensions); take the average value of the feature vectors of the same dimension, that is, add the vectors of the first dimension in the M theme vector and the first dimension in the N theme vector, and then take the average value as the vector of the first dimension of the main body vector, and so on until the 200th dimension is averaged respectively.

[0121] In step S250, determine the preset main body vector sample that matches the main body vector, and obtain the main body category label corresponding to the preset main body vector sample as the category label of the content publishing entity to classify the content publishing entity.

[0122] In the implementation of this example, the preset main body vector sample is the feature vector corresponding to the obtained main body category label. By calculating the similarity matching between the main body vector and all preset main body vector samples, the preset main body vector samples with similarity greater than a predetermined threshold can be used as the preset main body vector samples matched with the main body vector. Then, the main body category label corresponding to the preset main body vector sample can be used as the category label of the content publishing main body, and the content publishing main body can be classified according to this category label. The content publishing main body can be classified accurately and efficiently.

[0123] Among them, the method of calculating the similarity matching can be calculating the cosine vector or the Euclidean distance, etc.

[0124] In one embodiment, after segmenting the main body category label and converting the obtained words into feature vectors of a predetermined dimension, the vectors of the target dimension in the feature vectors corresponding to the words constituting the main body category label are aggregated to obtain the preset main body vector sample, which can effectively ensure the classification accuracy.

[0125] Refer to Figure 9 , this application also provides a method for obtaining training samples of a classification model for content publishing main bodies. Among them, the training samples are used to train the classification model. The training of the classification model takes the target information in the content publishing main body as the input and the category label corresponding to the content publishing main body as the output; the method includes:

[0126] Step S910, determine multiple content publishing main bodies included in the target publishing platform;

[0127] Step S920, based on the foregoing classification method of content publishing main bodies, respectively obtain the category labels corresponding to each content publishing main body;

[0128] Step S930, obtain the target information and the corresponding category label in each content publishing main body as the positive training samples of the classification model.

[0129] The training of the classification model takes the target information in the content publishing main body (such as text information such as the video account name, profile, and the titles / hashtags (topics) of all video account feeds under it) as the input and the category label corresponding to the content publishing main body as the output. Furthermore, by obtaining the target information and the corresponding category label in each content publishing main body as the input and output respectively, supervised training of the classification model can be carried out. Among them, the positive training samples are the samples belonging to the category corresponding to the category label. At the same time, the negative training samples can be randomly obtained from the content publishing main bodies that do not belong to the category corresponding to the category label.

[0130] In one embodiment, after determining multiple content publishing entities included in the target publishing platform, the importance of each content publishing entity is obtained, such as the number of fans and playback volume of the video account, and then the content publishing entity whose importance is higher than a predetermined threshold is taken as the acquired entity. Based on the aforementioned classification method of content publishing entities, the category label corresponding to each acquired entity is obtained respectively, and then the target information and the corresponding category label in the acquired entity are used as positive training samples of the classification model to further ensure the reliability of the training samples.

[0131] In one embodiment, the content description information in any of the aforementioned implementation modes is an information tag with the symbol #, namely, a topic (hashtag). A topic (hashtag) is usually a subject classification label and an extended reading label marked by a user for content. It can accurately and reliably reflect the content characteristics and the user's inner expression at the same time. Based on the topic (hashtag), the classification accuracy of the content classification subject can be effectively guaranteed.

[0132] See also Figure 10 and Figure 11 Taking WeChat video accounts as an example, each video account’s short video feeds are allowed to carry multiple topics (hashtags). At the same time, a video account often accumulates a large number of short videos (feeds) containing these different tags in history. Figure 10 The terminal interface shown is a video account named "Collection***", under which multiple video contents are published. Each video content has content description information under it, in particular, an information tag marked with the symbol #. Figure 11 When you open short video A in the terminal interface shown, you can see that there are financial topics "Finance" and technology topics "5G Era" under short video A. Although the name of the video account is centered on finance, by obtaining topics (hashtags) as content description information, that is, using information tags with the symbol # as content description information, the classification accuracy can be effectively guaranteed.

[0133] Figure 12 A block diagram of a classification device of a content publishing entity according to an embodiment of the present application is shown.

[0134] like Figure 12 As shown, the classification device 1200 of the content publishing entity may include a first acquisition module 1210 , a conversion module 1220 , an aggregation module 1230 , a normalization module 1240 and a classification module 1250 .

[0135] The first acquisition module 1210 can be used to acquire content description information corresponding to the content published by the content publishing entity, where the content description information is marked by the publishing user of the content; the conversion module 1220 can be used to perform word segmentation on the acquired content description information to obtain the words that make up the content description information, and convert the words into feature vectors of a predetermined dimension; the aggregation module 1230 can be used to perform an aggregation process on the vectors of the target dimension in the feature vectors corresponding to the words that make up the content description information to obtain the topic vector of the content description information; the normalization module 1240 can be used to perform a normalization process on the vectors of the target dimension in the topic vector corresponding to the acquired content description information to obtain the entity vector of the content publishing entity; the classification module 1250 can be used to determine a preset entity vector sample that matches the entity vector, and obtain an entity category label corresponding to the preset entity vector sample as the category label of the content publishing entity, so as to classify the content publishing entity.

[0136] In some embodiments of the present application, the content description information is an information label with a target symbol; the first acquisition module includes: a label acquisition unit, configured to acquire the information label corresponding to the content published by the content publishing entity, and obtain all the information labels included in the content publishing entity; a first frequency calculation unit, configured to calculate the first frequency of each information label appearing in the information labels included in the content publishing entity; a score determination unit, configured to determine the importance score of each information label according to the first frequency corresponding to each information label; a screening unit, configured to use multiple information labels with the highest importance scores in the content publishing entity as the acquired content description information.

[0137] In some embodiments of the present application, the content publishing entity is from a target publishing platform, and the target publishing platform further includes other content publishing entities; the score determination unit includes: a second frequency calculation subunit, configured to calculate the second frequency of each information label appearing in the other content publishing entities in the target publishing platform; a score determination subunit, configured to calculate the importance score of each information label according to the first frequency and the second frequency, where the importance score is proportional to the first frequency and inversely proportional to the second frequency.

[0138] In some embodiments of the present application, the second frequency calculation subunit is configured to: obtain a first number of all content publishing entities included in the target publishing platform; for each of the information tags, obtain a second number of content publishing entities that include the information tag among all content publishing entities included in the target publishing platform, to obtain a second number corresponding to each of the information tags; take the logarithm of the quotient obtained by dividing the first number by the second number corresponding to each of the information tags as the second frequency corresponding to each of the information tags; the score determination subunit is configured to: calculate the product of the first frequency and the second frequency corresponding to each of the information tags as the importance score of each of the information tags.

[0139] In some embodiments of the present application, the first frequency calculation unit includes: a similarity calculation subunit, configured to calculate the similarity between information tags included in the content publishing entity; a same tag determination subunit, configured to determine information tags with a similarity greater than a predetermined threshold as the same information tags; a tag number acquisition subunit, configured to obtain a first number of the same information tags and obtain a second number of information tags included in the content publishing entity; a first frequency calculation subunit, configured to obtain the first frequency of each information tag appearing in the information tags included in the content publishing entity by calculating the ratio of the first number to the second number.

[0140] In some embodiments of the present application, the aggregation module includes: an accumulation unit, configured to accumulate feature vectors of the same dimension in the feature vectors corresponding to the words constituting the content description information to obtain a topic vector of the content description information; the normalization module includes: an averaging unit, configured to take the average of feature vectors of the same dimension in the topic vector corresponding to the obtained content description information to obtain a subject vector of the content publishing entity.

[0141] In some embodiments of the present application, the conversion module includes: a word segmentation unit, configured to segment the obtained content description information using a target word segmenter to obtain words constituting the content description information; a vector conversion unit, configured to input the words into a pre-trained vector conversion model to obtain a feature vector of a predetermined dimension corresponding to the words output by the vector conversion model, wherein the word samples for training the vector conversion model are obtained by segmenting content description information samples using the target word segmenter.

[0142] In some embodiments of the present application, the first acquisition module includes: a content type acquisition unit configured to acquire the content type corresponding to the content published by the content publishing entity; a ratio determination unit configured to determine the proportion of content of a target type in the content type, where the content of the target type is other types of content except text type content; and a determination and acquisition unit configured to, when the proportion of the content of the target type is higher than a predetermined threshold, acquire the content description information corresponding to the content published by the content publishing entity.

[0143] Figure 13 FIG. shows a block diagram of a training sample acquisition device according to an embodiment of the present application.

[0144] As Figure 13 shown, a training sample acquisition device 1300 for a classification model of a content publishing entity, where the training samples are used to train the classification model, and the training of the classification model takes the target information in the content publishing entity as input and the category label corresponding to the content publishing entity as output; the training sample acquisition device 1300 may include a main body determination module 1310, a classification device 1200 of the content publishing entity, and a training sample acquisition module 1320.

[0145] The main body determination module 1310 may be configured to determine a plurality of content publishing entities included in the target publishing platform;

[0146] The classification device 1200 of the content publishing entity may be configured to respectively acquire the category label corresponding to each content publishing entity;

[0147] The training sample acquisition module 1320 may be configured to acquire the target information and the corresponding category label in each content publishing entity as the positive training sample of the classification model.

[0148] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.

[0149] In addition, an embodiment of the present application further provides an electronic device, which may be a terminal or a server. As Figure 14 shown, it shows a schematic structural diagram of the electronic device involved in the embodiment of the present application. Specifically:

[0150] The electronic device may include components such as a processor 1401 with one or more processing cores, a memory 1402 of one or more computer-readable storage media, a power supply 1403, and an input unit 1404. Those skilled in the art can understand that Figure 14 the structure of the electronic device shown in

[0151]

[0152]

[0153]

[0154] does not limit the electronic device, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Among them:

[0154] The processor 1401 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 1402, and by calling the data stored in the memory 1402, it executes various functions of the computer device and processes data, thereby monitoring the computer device as a whole. Optionally, the processor 1401 may include one or more processing cores; preferably, the processor 1401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1401.

[0152] The memory 1402 can be used to store software programs and modules. The processor 1401 executes various functional applications and data processing by running the software programs and modules stored in the memory 1402. The memory 1402 may mainly include a program storage area and a data storage area. Among them, the program storage area may store the operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the computer device. In addition, the memory 1402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 1402 may also include a memory controller to provide the processor 1401 with access to the memory 1402.

[0153] The computer device also includes a power supply 1403 that supplies power to each component. Preferably, the power supply 1403 can be logically connected to the processor 1401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 1403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0154] The computer device may further include an input unit 1404, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0155] Although not shown, the computer device may further include a display unit and the like, which will not be elaborated herein. Specifically, in this embodiment, the processor 1401 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 1402 according to the following instructions, and the processor 1401 will run the application programs stored in the memory 1402 to implement various functions.

[0156] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated herein.

[0157] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners in the above embodiments.

[0158] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a computer program or by controlling related hardware through a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0159] Therefore, the embodiments of the present application further provide a storage medium, in which a computer program is stored, and the computer program can be loaded by a processor to execute the steps in any one of the methods provided by the embodiments of the present application.

[0160] Wherein, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0161] Since the computer program stored in the storage medium can execute the steps in any one of the live message interaction methods provided by the embodiments of the present application, the beneficial effects achievable by the methods provided by the embodiments of the present application can be realized. For details, please refer to the previous embodiments, which will not be elaborated herein.

[0162] The above detailed introduction to the embodiments provided by the present application uses specific examples to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A classification method for content publishing entities, characterized in that, Including: Obtain content description information corresponding to the content published by a content publishing entity, where the content description information is marked by the user who publishes the content; Segment the obtained content description information to obtain words that make up the content description information, and convert the words into feature vectors of a predetermined dimension; Perform aggregation processing on vectors of a target dimension among the feature vectors corresponding to the words that make up the content description information to obtain a topic vector of the content description information; Perform normalization processing on vectors of a target dimension among the topic vectors corresponding to the obtained content description information to obtain a main body vector of the content publishing entity; Determine a preset main body vector sample that matches the main body vector, and obtain a main body category label corresponding to the preset main body vector sample as the category label of the content publishing entity, so as to classify the content publishing entity.

2. The method according to claim 1, characterized in that, The content description information is an information label with a target symbol; The obtaining content description information corresponding to the content published by a content publishing entity includes: Obtain the information labels corresponding to the content published by the content publishing entity to obtain all the information labels included in the content publishing entity; Calculate a first frequency of each information label appearing in the information labels included in the content publishing entity; Determine an importance score for each information label according to the first frequency corresponding to each information label; Use multiple information labels with the highest importance scores in the content publishing entity as the obtained content description information.

3. The method according to claim 2, wherein The content publishing entity is from a target publishing platform, and the target publishing platform also includes other content publishing entities; The determining an importance score for each information label according to the first frequency corresponding to each information label includes: Calculate a second frequency of each information label appearing in other content publishing entities in the target publishing platform; Calculate an importance score for each information label according to the first frequency and the second frequency, where the importance score is directly proportional to the first frequency and inversely proportional to the second frequency.

4. The method according to claim 3, characterized in that, The first frequency is a word frequency, and the calculating a second frequency of each information label appearing in other content publishing entities in the target publishing platform includes: Obtain a first number of all content publishing entities included in the target publishing platform; Sequentially determine a second number of content publishing entities that include the information label from all content publishing entities included in the target publishing platform to obtain a second number corresponding to each information label; Take the logarithm of the quotient obtained by dividing the first number by the second number corresponding to each information label as the second frequency corresponding to each information label; The calculating an importance score for each information label according to the first frequency and the second frequency includes: Calculate the product of the first frequency and the second frequency corresponding to each information label as the importance score of each information label.

5. The method according to claim 2, wherein Calculating the first frequency of each of the information tags appearing in the information tags included in the content publishing entity includes: Calculating the similarity between the information tags included in the content publishing entity; Determining the number of times each information tag appears in the information tags included in the content publishing entity by determining information tags with a similarity greater than a predetermined threshold as the same information tag, and calculating the first frequency.

6. The method according to claim 1, wherein Aggregating the vectors of the target dimension in the feature vectors corresponding to the words constituting the content description information to obtain the topic vector of the content description information includes: Adding up the feature vectors of the same dimension in the feature vectors corresponding to the words constituting the content description information to obtain the topic vector of the content description information; Normalizing the vectors of the target dimension in the topic vector corresponding to the obtained content description information to obtain the entity vector of the content publishing entity includes: Taking the average value of the feature vectors of the same dimension in the topic vector corresponding to the obtained content description information to obtain the entity vector of the content publishing entity.

7. The method according to claim 1, characterized in that, Performing word segmentation on the obtained content description information to obtain the words constituting the content description information, and converting the words into feature vectors of a predetermined dimension includes: Performing word segmentation on the obtained content description information using a target word segmenter to obtain the words constituting the content description information; Inputting the words into a pre-trained vector conversion model to obtain the feature vectors of a predetermined dimension corresponding to the words output by the vector conversion model, where the word samples for training the vector conversion model are obtained by segmenting content description information samples using the target word segmenter.

8. The method according to claim 1, wherein Obtaining the content description information corresponding to the content published by the content publishing entity includes: Obtaining the content type corresponding to the content published by the content publishing entity; Determining the proportion of the target type content in the content type, where the target type content is other types of content except text type content; When the proportion of the target type content is higher than a predetermined threshold, obtaining the content description information corresponding to the content published by the content publishing entity.

9. A method for obtaining training samples of a classification model for content publishing entities, characterized in that, The training samples are used to train the classification model. The training of the classification model uses the target information in the content publishing entity as the input and the category label corresponding to the content publishing entity as the output; the method includes: Determining a plurality of content publishing entities included in the target publishing platform; Based on the method according to any one of claims 1-8, respectively obtaining the category label corresponding to each content publishing entity; Obtaining the target information and the corresponding category label in each content publishing entity as the positive training sample of the classification model.

10. A classification device for content publishing entities, characterized in that, Includes: A first acquisition module, configured to acquire the content description information corresponding to the content published by the content publishing entity, where the content description information is calibrated by the publishing user of the content; A conversion module, configured to perform word segmentation on the acquired content description information to obtain the words constituting the content description information, and convert the words into feature vectors of a predetermined dimension; An aggregation module for aggregating vectors of a target dimension in the feature vectors corresponding to the words constituting the content description information to obtain a topic vector of the content description information; A normalization module for normalizing vectors of a target dimension in the topic vector corresponding to the obtained content description information to obtain a subject vector of the content publishing subject; A classification module for determining a preset subject vector sample that matches the subject vector, and obtaining a subject category label corresponding to the preset subject vector sample as the category label of the content publishing subject to classify the content publishing subject.

Citation Information

Patent Citations

  • Character information classification method and device and electronic device

    CN110309308A

  • Hotspot event classification method and apparatus, and storage medium

    WO2019184217A1