A method, device and computer device for label determination
By calculating the broadness and importance parameters of the tags in the content interaction platform, and integrating the tag recommendation reference information, the problem of low label accuracy in the prior art is solved, and the accuracy and exposure of tags are improved.
Patent Information
- Application Number
- CN202011410643.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-04
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-12-04
AI Technical Summary
In the prior art, the labels determined in the content interaction platform have low accuracy, mainly based on user semantics and operational experience, resulting in inaccurate labels.
By determining the label set corresponding to multiple contents in the content interaction platform, the broadness parameters and importance parameters of the target label are calculated, and these parameters are fused to obtain label recommendation reference information, and then the label of content to be published is determined.
Improve the accuracy and exposure of labels, ensuring the rationality and effectiveness of label recommendations.
Smart Images

Figure CN113392316B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a method, apparatus, and computer device for determining tags. Background Art
[0002] With the rapid development of information technology, users can publish various types of content on a content interaction platform and determine corresponding tags when publishing the content. Other users can browse the content corresponding to the tags on the content interaction platform by searching for the tags.
[0003] In the process of researching and practicing related technologies, the inventors of this application found that most of the tags currently determined on the content interaction platform are manually input by users based on the semantics of the content to be published or their own operation experience, and the accuracy of the tags is relatively low. Summary of the Invention
[0004] Embodiments of this application provide a method, apparatus, and computer device for determining tags, which can improve the accuracy of tags.
[0005] Embodiments of this application provide a method for determining tags, including:
[0006] Determine a tag set corresponding to multiple pieces of content in a content interaction platform, where the tag set includes content tags corresponding to the content, and the content tags characterize the content type of the content;
[0007] Determine, from the tag set, a target tag set having a word inclusion relationship with a target tag, where the target tag is used to determine the tag of the content to be published in the content interaction platform;
[0008] Determine the number of tags in the tag set and the tags in the target tag set used in the content interaction platform;
[0009] Based on the number of tags, obtain probability information of the target tag set in the tag set;
[0010] According to the probability information, obtain the information entropy of the target tag, and determine a breadth parameter of the target tag, where the breadth parameter characterizes the amount of information of the target tag;
[0011] Obtain the search frequency of the target tag in the tag set;
[0012] Based on the search frequency and the number of contents corresponding to the target tag, determine an importance parameter of the target tag, where the importance parameter characterizes the importance of the target tag in the tag set;
[0013] Fuse the breadth parameter and the importance parameter to determine the label recommendation reference information of the target label;
[0014] Determine the labels of the content to be published in the content interaction platform according to the label recommendation reference information of the target label.
[0015] Correspondingly, an embodiment of the present application provides a label determination device, including:
[0016] A first label set determination unit, configured to determine a label set corresponding to multiple contents in a content interaction platform, where the label set includes content labels corresponding to the contents, and the content labels represent the content types of the contents;
[0017] A second label set determination unit, configured to determine, from the label set, a target label set having a word inclusion relationship with a target label, where the target label is used to determine the labels of the content to be published in the content interaction platform;
[0018] A first parameter determination unit, configured to determine the number of labels in the label set and the target label set in the content interaction platform; based on the number of labels, obtain the probability information of the target label set in the label set; according to the probability information, obtain the information entropy of the target label, and determine the breadth parameter of the target label, where the breadth parameter represents the amount of information of the target label;
[0019] A first acquisition unit, configured to acquire the search frequency of the target label in the label set;
[0020] A second parameter determination unit, configured to determine the importance parameter of the target label based on the search frequency and the number of contents corresponding to the target label, where the importance parameter represents the importance of the target label in the label set;
[0021] A fusion unit, configured to fuse the breadth parameter and the importance parameter to determine the label recommendation reference information of the target label;
[0022] A first label determination unit, configured to determine the labels of the content to be published in the content interaction platform according to the label recommendation reference information of the target label.
[0023] In an embodiment, the first label determination unit includes:
[0024] A score determination subunit, configured to determine the label recommendation score of the target label according to the label recommendation reference information of the target label;
[0025] A label determination subunit, configured to determine, when the label recommendation score reaches a preset recommendation score, the label of the content to be published in the content publishing platform as the target label.
[0026] In one embodiment, the first parameter determination unit includes:
[0027] An information determination subunit, configured to determine the search information of the labels in the label set and the target label set, where the search information includes the search times of the labels in the label set and the target search times of the labels in the target label set;
[0028] A second calculation subunit, configured to calculate the breadth parameter of the target label based on the search times and the target search times.
[0029] In one embodiment, the first acquisition unit includes:
[0030] A first fusion subunit, configured to fuse the search times of the target label and the search times of the labels in the label set to obtain the search frequency of the target label.
[0031] In one embodiment, the second parameter determination unit includes:
[0032] An acquisition subunit, configured to acquire the content quantity of the multiple contents;
[0033] A second fusion subunit, configured to fuse the content quantity of the multiple contents and the content quantity corresponding to the target label to determine the importance parameter of the target label;
[0034] A third fusion subunit, configured to fuse the search frequency and the importance parameter to obtain the importance degree parameter of the target label.
[0035] In one embodiment, the label determination device further includes:
[0036] A second acquisition unit, configured to acquire the editing content of the label editing operation when detecting a label editing operation of a user for the current content to be published, where the current content to be published is the content to be published in the content publishing platform;
[0037] A first recognition unit, configured to recognize the editing content to obtain the label recommendation reference information of the editing content;
[0038] A second label determination unit, configured to determine the label of the current content to be published based on the label recommendation reference information of the editing content.
[0039] In one embodiment, the label determination device further includes:
[0040] A second recognition unit, configured to perform content recognition on the current content to be published when detecting a tag acquisition operation of the user for the current content to be published, so as to obtain a content recognition result;
[0041] A third acquisition unit, configured to acquire a content matching tag that matches the current content to be published from a tag library based on the content recognition result of the current content to be published;
[0042] A third tag determination unit, configured to determine a tag of the current content to be published from the content matching tags based on tag recommendation reference information of the content matching tags.
[0043] Correspondingly, an embodiment of the present application further provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the steps in any of the tag determination methods provided by the embodiments of the present application.
[0044] Correspondingly, an embodiment of the present application further provides a storage medium storing multiple instructions, which are suitable for being loaded by a processor to execute the steps in any of the tag determination methods provided by the embodiments of the present application.
[0045] Embodiments of the present application can determine a tag set corresponding to multiple contents in a content interaction platform. The tag set includes content tags corresponding to the contents, where the content tags represent the content types of the contents; determine a target tag set having a word inclusion relationship with a target tag from the tag set, where the target tag is used to determine a tag of the content to be published in the content interaction platform; determine a breadth parameter of the target tag based on the tag set and the target tag set, where the breadth parameter represents the amount of information of the target tag; obtain a search frequency of the target tag in the tag set; determine an importance parameter of the target tag based on the search frequency and the number of contents corresponding to the target tag, where the importance parameter represents the importance of the target tag in the tag set; fuse the breadth parameter and the importance parameter to determine tag recommendation reference information of the target tag; and determine a tag of the content to be published in the content interaction platform according to the tag recommendation reference information of the target tag. This solution can obtain the breadth parameter of the target tag to determine the breadth of the target tag, then determine the importance parameter of the target tag, that is, determine the posterior heat feature of the tag of the target tag, based on the search frequency of the target tag in the tag set and the number of contents of the target tag. Then, based on the breadth and the posterior heat of the tag, determine the tag recommendation reference information of the target tag. Finally, the tag of the content to be published in the content interaction platform can be determined according to the tag recommendation reference information, which can improve the tag accuracy. Description of the Drawings
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0047] Figure 1 It is a schematic diagram of the scenario of the tag determination method provided by the embodiment of the present application;
[0048] Figure 2a It is a flowchart of the tag determination method provided by the embodiment of the present application;
[0049] Figure 2b It is a multi-round conversation flowchart of the tag determination method provided by the embodiment of the present application;
[0050] Figure 3 It is another flowchart of the tag determination method provided by the embodiment of the present application;
[0051] Figure 4a It is a device diagram of the tag determination method provided by the embodiment of the present application;
[0052] Figure 4b It is another device diagram of the tag determination method provided by the embodiment of the present application;
[0053] Figure 4c It is another device diagram of the tag determination method provided by the embodiment of the present application;
[0054] Figure 4d It is another device diagram of the tag determination method provided by the embodiment of the present application;
[0055] Figure 4e It is another device diagram of the tag determination method provided by the embodiment of the present application
[0056] Figure 5 It is a schematic diagram of the structure of the computer device provided by the embodiment of the present application. Detailed implementation manners
[0057] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0058] The embodiments of the present application provide a method, an apparatus, a computer device, and a storage medium for determining tags. Specifically, the embodiments of the present application provide a tag determination apparatus applicable to a computer device. Among them, the computer device may be a device such as a terminal or a server. The server may be an independent physical server or a server cluster or a distributed system composed of multiple physical servers. The terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.
[0059] Referring Figure 1 , taking the computer device as a server as an example, the server may determine a tag set corresponding to multiple contents in the content interaction platform. The tag set includes content tags corresponding to the contents, where the content tags represent the content types of the contents; from the tag set, determine a target tag set that has a word inclusion relationship with the target tag, and the target tag is used to determine the tags of the content to be published in the content interaction platform; based on the tag set and the target tag set, determine the breadth parameter of the target tag, and the breadth parameter represents the amount of information of the target tag; obtain the search frequency of the target tag in the tag set; based on the search frequency and the number of contents corresponding to the target tag, determine the importance parameter of the target tag, and the importance parameter represents the importance of the target tag in the tag set; fuse the breadth parameter and the importance parameter to determine the tag recommendation reference information of the target tag; according to the tag recommendation reference information of the target tag, determine the tags of the content to be published in the content interaction platform.
[0060] Among them, determining the target tag set that has a word inclusion relationship with the target tag from the tag set can be implemented based on natural language processing technology in the field of artificial intelligence. For example, the target tag can be determined from the tag set, and then based on natural language processing technology, the target tag set that has a word inclusion relationship with the target tag can be determined from the tag set.
[0061] Among them, artificial intelligence (AI) is to utilize a digital computer or a machine model controlled by a digital computer to extend and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, method, technology, and application systems. Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence software technologies mainly include natural language processing, machine learning / deep learning, etc.
[0062] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs, and other technologies.
[0063] As can be seen from the above, the embodiments of the present application can obtain the breadth parameter of the target tag to determine the breadth of the target tag, and then determine the importance parameter of the target tag based on the search frequency of the target tag in the tag set and the content quantity of the target tag, that is, determine the posterior heat feature of the target tag. After that, based on the breadth and the posterior heat of the tag, determine the tag recommendation reference information of the target tag. Finally, the tags of the content to be published in the content interaction platform can be determined according to the tag recommendation reference information, which can improve the tag accuracy.
[0064] The present embodiment will be described in detail separately as follows. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0065] The embodiments of the present application provide a tag determination method. This method can be executed by a terminal or a server, or jointly executed by a terminal and a server. The embodiments of the present application take the example that the tag determination method is executed by the server for illustration. Specifically, it is executed by a tag determination device integrated in the server. As Figure 2a shown, the specific process of this tag determination method can be as follows:
[0066] 201. Determine the tag set corresponding to multiple contents in the content interaction platform. The tag set includes the content tags corresponding to the contents, where the content tags represent the content types of the contents.
[0067] Among them, the content tag is the tag of the corresponding content in the content interaction platform. The content tag can be a keyword, key phrase, or an image, audio, etc. that is highly relevant to the corresponding content. The content tag can help users easily describe and classify the contents in the content interaction platform for retrieval and classification.
[0068] Among them, the content type refers to the type corresponding to each content in the content interaction platform. For example, each content can introduce types such as pet raising, language training, daily necessities purchase, and work and study.
[0069] Among them, the tag set includes multiple content tags, and each content tag can correspond to one or more contents in one or more content interaction platforms. For example, a content tag in the tag set can be the content tag of one content in the content interaction platform, or the content tags of multiple contents in the content interaction platform, and the content tag of one content in the content interaction platform can be one or more.
[0070] In one example, in the content interaction platform, users can accurately locate the content under the content tag by searching for the content tag. For example, in the content interaction platform, users are supported to independently define and recommend some hashtags (tweet topics, such as words preceded by the # sign on Twitter to represent tweet topics) that can be used for classification, sorting, and searching, so as to facilitate users to, after browsing a content, through the content tag of this content, such as hashtag, further extend to read the content in the content interaction platform with the same content tag.
[0071] For example, in the content interaction platform, the content tag corresponding to the content read by the user is "English training". If the user is interested in "English training", then in the content interaction platform, the user can further extend to read the content that also has "English training" as the content tag.
[0072] 202. Determine, from the tag set, a target tag set that has a word inclusion relationship with the target tag, where the target tag is used to determine the tag of the content to be published in the content interaction platform.
[0073] Among them, the word inclusion relationship refers to the relationship between two words with a connection in content. For example, the inclusion sequence relationship of "fresh flowers" can be found according to the dictionary inclusion relationship, that is, words such as "fresh flower delivery", "fresh flower distribution", and "fresh flower packaging" have a word inclusion relationship with "fresh flowers".
[0074] Among them, the target tag can be used to determine the tag of the content to be published in the content interaction platform, and the target tag set refers to the set of content tags that have an inclusion relationship with the target tag. For example, if the target tag is "fresh flowers", then one tag in the target tag set can be "fresh flower delivery", and "fresh flower delivery" has an inclusion relationship with "fresh flowers", and the tag set includes the target tag set.
[0075] In one example, as Figure 2b shown, taking the target tag as "car" as an example, in the content interaction platform, by searching for the hashtag "car", users can display the content that also has the target tag "car" on the content list page.
[0076] 203. Based on the tag set and the target tag set, determine the breadth parameter of the target tag, where the breadth parameter characterizes the amount of information of the target tag.
[0077] Among them, the generality of the target label can be measured based on the information entropy between the target label set and the label set. The more content labels include it as a substring, the greater the entropy, that is, the higher the generality value and the greater the generality parameter.
[0078] Among them, information entropy is a measure of the complexity of a system. If the system is more complex and there are more types of different situations, then its information entropy is relatively large. If a system is simpler and there are few types of situations (the extreme case is 1 situation, then the corresponding probability is 1, and the corresponding information entropy is 0), the information entropy at this time is smaller. The formula is as follows:
[0079]
[0080] Among them, p(x i ) represents the probability that the random event x is x i . The amount of information H(X)=-logp(x i ). Information measures the information brought by the occurrence of a specific event, while entropy is the expectation of the amount of information that may be generated before the result comes out, considering all possible values of the random variable, that is, the expectation of the amount of information brought by all possible events that may occur, that is:
[0081] H(X)=-sum(p(x)log2p(x i ))
[0082] After transformation, the form of the aforementioned information entropy formula after derivation can be obtained:
[0083]
[0084] Among them, p(x) refers to the ratio of the number of times the label in the target label set is used to the number of times the label in the label set is used in the content interaction platform, that is, the probability of the target label set in the label set. For example, when the number of times the label in the target label set is used is 3 times and the number of times the label in the label set is used is 100 times, p(x)=3 / 100 = 0.03.
[0085] In one embodiment, the step of "determining the generality parameter of the target label based on the label set and the target label set" may include:
[0086] Determine the number of labels used in the label set and the target label set in the content interaction platform;
[0087] Based on the number of labels, calculate the information quantity generality of the target label to obtain the generality parameter of the target label.
[0088] In one embodiment, the detailed process of the step "calculating the information breadth of the target tag based on the number of tags to obtain the breadth parameter of the target tag" may include:
[0089] Based on the number of tags, obtain the probability information of the target tag set in the tag set;
[0090] According to the probability information, obtain the information entropy of the target tag and determine the breadth parameter of the target tag.
[0091] In one example, taking the target tag as "flowers" as an example to illustrate the calculation of the breadth parameter. In the content interaction platform, the target tag set and the number of occurrences of the content tags in the target tag set are as follows:
[0092] Tag No Tag Name Tag Times Tag1 Fresh Flowers 5 Tag2 Fresh Flower Delivery 10 Tag3 Fresh Flower Distribution 10 Tag4 Fresh Flower Packaging 70
[0093] Among them, if the number of occurrences of the content tag in the tag set is 100, the calculation process of the breadth parameter of the target tag "flowers" can be as follows:
[0094] P(flowers) = (5 + 10 + 10 + 70) / 100 = 0.95
[0095] H(flowers) = -(0.95 * log(0.95)) = 0.02
[0096] Among them, the hashtag information entropy of P(flowers) can be calculated through the above information entropy formula, and H(flowers) = 0.02 indicates that the breadth parameter of "flowers" is 0.02. The larger the breadth parameter, the more extensive it is.
[0097] In one embodiment, the step "determining the breadth parameter of the target tag based on the tag set and the target tag set" may include:
[0098] Determine the search information of the tags in the tag set and the target tag set. The search information includes the search times of the tags in the tag set and the target search times of the tags in the target tag set;
[0099] Based on the search times and the target search times, calculate the breadth parameter of the target tag.
[0100] Among them, the search information of the labels in the label set and the target label set can be directly obtained and determined by the server. Specifically, in one example, a mapping relationship can be established in the label database of the server between the labels in the label set and the number of times the labels in the label set are searched, and a mapping relationship between the labels in the target label set and the number of times the target labels in the target label set are searched. Whenever the search times of the labels in the label set and the labels in the target label set in the content interaction platform change, the mapping relationships in the label database between the labels in the label set and the number of times the labels in the label set are searched, and between the labels in the target label set and the number of times the target labels in the target label set are searched can be changed simultaneously. When it is necessary to determine the search information of the labels in the label set and the target label set, it can be directly obtained from the label database by the server.
[0101] Among them, the breadth parameter of the target label can also be obtained by the ratio of the target search times to the search times to get the search frequency of the target label, and then the information entropy of the target label can be obtained according to the search frequency, and further the breadth parameter of the target label can be determined.
[0102] In one example, the breadth parameter of the target label can be determined according to information such as the search times of each label in the label set and the target label set in the content interaction platform. Specifically, the search times of the labels in the label set and the search times of the labels in the target label set can be obtained first, and then the search times of the labels in the target label set are divided by the search times of the labels in the label set, and the ratio is the search frequency of the target label. The probability information of the above-mentioned target label in the label set can be replaced by the search frequency, and then the search probability information p(x) of the target label in the label set can be replaced by p(search frequency). Finally, the breadth parameter of the target label can be calculated through the above information entropy formula, that is, the breadth parameter of the target label can be calculated by the following formula:
[0103]
[0104] 204. Obtain the search frequency of the target label in the label set, and determine the importance parameter of the target label based on the search frequency and the number of contents corresponding to the target label. The importance parameter characterizes the importance of the target label in the label set.
[0105] Among them, the search frequency can be understood as the frequency of the target label appearing in the label set. This number is a normalization of the number of words to prevent it from biasing towards a large label set and can be used to determine the importance of the target label in the label set.
[0106] Among them, the importance parameter can be obtained by fusing the search frequency and the number of contents of the target label and can be used to judge the importance of the target label in the label set.
[0107] In one embodiment, the step of "obtaining the search frequency of the target tag in the tag set" may include:
[0108] Fusing the number of searches for the target tag and the number of searches for the tags in the tag set to obtain the search frequency of the target tag;
[0109] "Determining the importance parameter of the target tag based on the search frequency and the number of contents corresponding to the target tag" may include:
[0110] Obtaining the number of contents of multiple contents;
[0111] Fusing the number of contents of multiple contents and the number of contents corresponding to the target tag to determine the importance parameter of the target tag;
[0112] Fusing the search frequency and the importance parameter to obtain the importance parameter of the target tag.
[0113] Among them, to fuse the number of searches for the target tag and the number of searches for the tags in the tag set, specifically, the number of searches for the target tag can be divided by the number of searches for the tags in the tag set, and the obtained ratio is the search frequency of the target tag.
[0114] In one example, the tag set can be regarded as a relatively large hashtag article, and then the importance of each hashtag in this hashtag article can be calculated as idf (Inverse Document Frequency), and the number of searches for this hashtag in the tag set can be used as tf (term frequency). That is, the main idea of TFIDF (term frequency–inverse document frequency, a commonly used weighting technique for information retrieval and data mining) can be used to calculate the importance parameter of the target tag (or it can be understood as the prior heat feature).
[0115] Among them, TF-IDF is a classical statistical method used to evaluate the importance of a word for a document set or a single document in a corpus. The importance of a word increases in direct proportion to the number of times it appears in the document, but at the same time decreases in inverse proportion to the frequency of its appearance in the corpus.
[0116] The introduction of the traditional tf*idf idea is as follows: If a word or phrase appears frequently (TF) in an article and rarely appears in other articles, it is considered that this word or phrase has good category discrimination ability and is suitable for classification. FIDF is actually: TF*IDF, where TF represents the frequency of the term in document d. The main idea of IDF is: If the number of documents containing the term t is smaller, that is, n is smaller, and the IDF is larger, it indicates that the term t has good category discrimination ability. If the number of documents containing the term t in a certain category of documents C is m, and the total number of documents containing t in other categories is k, obviously the total number of documents containing t, n = m + k. When m is large, n is also large, and the IDF value obtained according to the IDF formula will be small, indicating that the category discrimination ability of the term t is not strong. However, in fact, if a term appears frequently in the documents of a class, it means that this term can well represent the characteristics of the text of this class. Such terms should be given higher weights and selected as the characteristic words of this class of text to distinguish them from other classes of documents. This is the shortcoming of IDF. In a given document, TF refers to the frequency of a given word in that document. This number is a normalization of the term count to prevent it from biasing towards long documents. (The same word may have a higher term count in a long document than in a short document, regardless of the importance of the word.) For a word in a specific document, its importance can be expressed as:
[0117]
[0118] Among them, in the above formula, the numerator is the hashtag. For example, it is the number of searches for the target tag in the content interaction platform, and the denominator is the sum of the number of searches for all tags in the tag set in the content interaction platform. The inverse document frequency IDF is a measure of the general importance of a hashtag. The IDF of a specific hashtag can be obtained by dividing the number of contents of multiple contents in the content interaction platform by the number of contents containing the content corresponding to the hashtag, and then taking the logarithm of the obtained quotient. The formula can be as follows:
[0119]
[0120] Among them, |D| refers to the number of contents of multiple contents in the content interaction platform. The IDF of the target tag can be obtained by dividing the number of contents of multiple contents in the content interaction platform by the number of contents containing the target tag, and then taking the logarithm of the obtained quotient.
[0121] Among them, |{j:t i ∈d j}| refers to the number of contents containing the target tag. For example, if the number of contents containing the target tag is 10, then |{j:ti ∈d j} = 10.
[0122] 205. Integrate the breadth parameter and the importance parameter to determine the label recommendation reference information for the target label.
[0123] Among them, the label recommendation reference information indicates the recommendation degree for the target label. For example, the label recommendation parameter information can include the label recommendation score. Based on the size of the label recommendation score, determine the recommendation degree of the target label corresponding to the label recommendation score. For a higher label recommendation score, the corresponding recommendation degree of the target label can be higher, and for a lower label recommendation score, the corresponding recommendation degree of the target label can be lower.
[0124] In one example, the label recommendation score of the target label can be obtained by integrating the breadth parameter and the importance parameter of the target label. The integration of the breadth parameter and the importance of the target label can be through the following formula:
[0125] Compete_Score(Hashtag) = H(hashtag) * log(Hotness(hashtag))
[0126] Among them, Compete_Score(Hashtag) represents the label recommendation score of this hashtag. Taking this hashtag as an example of the target label, Compete_Score(Hashtag) represents the label recommendation score of the target label.
[0127] Among them, H(hashtag) represents the breadth parameter of this hashtag. Taking this hashtag as an example of the target label, H(hashtag) represents the breadth parameter of the target label.
[0128] Among them, log(Hotness()hashtag) represents the importance parameter of this hashtag. Taking the log for smoothing processing, that is, it can be understood that after performing TF-IDF on this hashtag, taking the log for smoothing processing, and finally multiplying and integrating with the breadth parameter of this hashtag to obtain the label recommendation score of this hashtag.
[0129] 206. Determine the labels of the content to be published in the content interaction platform according to the label recommendation reference information of the target label.
[0130] Among them, for the tags of each content to be released in the content interaction platform, the content tags used by each content in the history of the content interaction platform can be obtained, or the content tags that have never been used by each content in the content interaction platform can be used. For example, new content tags can be obtained according to the tag editing operation of the user as the tags of the content to be released.
[0131] Among them, it is possible to determine whether the content tag can be used as the tag of the content to be released by determining the tag recommendation reference information corresponding to the content tags used by each content in the history, and it is also possible to determine whether the new content tag edited by the user can be used as the tag of the content to be released by calculating the tag recommendation reference information of the new content tag.
[0132] Among them, obtain the breadth parameter of the target tag in the content interaction platform to determine the breadth of the target tag, and then determine the importance parameter of the target tag according to the search frequency of the target tag in the tag set and the number of contents of the target tag. Then, based on the breadth parameter and the importance parameter of the target tag, determine the tag recommendation reference information of the target tag, that is, determine the tag recommendation reference information of the target tag through the two aspects of prior breadth and posterior heat to determine the tag of the content to be released in the content interaction platform, which can not only improve the tag accuracy but also improve the exposure of the content to be released.
[0133] In one embodiment, the step of "determining the tag of the content to be released in the content interaction platform according to the tag recommendation reference information of the target tag" may include:
[0134] Determine the tag recommendation score of the target tag according to the tag recommendation reference information of the target tag;
[0135] When the tag recommendation score reaches the preset recommendation score, determine that the tag of the content to be released in the content publishing platform is the target tag.
[0136] In one example, the tag recommendation score of the target tag can be determined based on the breadth parameter of the target tag and the importance parameter of the target tag. Specifically, it can be calculated by the following formula:
[0137] Compete_Score(Hashtag)=H(hashtag)*log(Hotness(hashtag))
[0138] Among them, Compete_Score(Hashtag) refers to the hashtag recommendation score of the target hashtag, H(hashtag) refers to the popularity parameter of the target hashtag, and Hotness(hashtag) refers to the importance parameter of the target hashtag, which can be calculated by tf×idf. Finally, the popularity parameter of the target hashtag is multiplied and fused with the importance parameter Hotness(hashtag) of the target hashtag to obtain the hashtag recommendation score Compete_Score(Hashtag) of the target hashtag. Considering that the value range of the importance parameter Hotness(hashtag) of the target hashtag may be relatively large, log can be taken for smoothing processing.
[0139] In one example, based on the hashtag recommendation score of the target hashtag, it can be determined whether the target hashtag can be used as the hashtag of the content to be published. Selecting the content hashtag with a larger hashtag recommendation score as the content hashtag of the content to be published can increase the exposure rate of the content to be published on the content interaction platform.
[0140] For example, in the content interaction platform, when a user actively enters a certain candidate content hashtag for the content to be published, the hashtag recommendation score of the content hashtag on the content interaction platform can be prompted. A larger hashtag recommendation score may mean that if this content hashtag is selected as the content hashtag of the content to be published, the relationship between the popularity and competitiveness of the content hashtag can be balanced.
[0141] Among them, if the hashtag recommendation score is small, it may mean either competing with many more competitive published contents, that is, it is more difficult to obtain additional traffic by clicking on the hashtag search list by users, or it may mean that although there will be less competition from similar contents, it may mean that the hashtag is too unpopular. If this hashtag is selected as the content hashtag of the user's content to be published, the exposure rate of the content to be published on the content interaction platform is low.
[0142] In one embodiment, the hashtag determination method may further include:
[0143] When detecting a hashtag editing operation by the user for the current content to be published, obtain the editing content of the hashtag editing operation, where the current content to be published is the content to be published on the content publishing platform;
[0144] Identify the editing content to obtain the hashtag recommendation reference information of the editing content;
[0145] Based on the hashtag recommendation reference information of the editing content, determine the hashtag of the current content to be published.
[0146] In one example, the tags of the current content to be published can also be edited to obtain candidate tags for the current content to be published. Then, the popularity parameter and importance parameter of the candidate tags can be calculated. After that, the popularity parameter and importance parameter of the candidate tags are multiplied and fused to obtain the tag recommendation reference information for the candidate tags. The result of multiplying and fusing the popularity parameter and importance parameter is the tag recommendation score for the candidate tags. Finally, the tags of the current content to be published can be determined according to the tag recommendation score of the candidate tags. For example, candidate tags with higher tag recommendation scores can be determined from the candidate tags as the tags of the current content to be published, which can increase the exposure of the current content to be published.
[0147] In one embodiment, the tag determination method may further include:
[0148] When it is detected that the user performs a tag acquisition operation on the current content to be published, the current content to be published is identified to obtain a content identification result;
[0149] Based on the content identification result of the current content to be published, content matching tags that match the current content to be published are obtained from the tag library;
[0150] Based on the tag recommendation reference information of the content matching tags, the tags of the current content to be published are determined from the content matching tags.
[0151] In one example, the content of the current content to be published can be identified, and then based on the content identification result of the current content to be published, content matching tags whose tag content matches the current content to be published are found from the numerous tags in the tag library. For example, semantic matching, etc. Then, the popularity parameter and importance parameter of the content matching tags are calculated. After that, the popularity parameter and importance parameter of the content matching tags are multiplied and fused to obtain the tag recommendation reference information for the content matching tags. Furthermore, the tag recommendation score of the content matching tags can be determined. Finally, based on the tag recommendation score of the content matching tags, the tags of the current content to be published can be determined from the content matching tags. For example, tags with higher tag recommendation scores can be determined from the content matching tags as the tags of the current content to be published to increase the exposure of the current content to be published.
[0152] As can be seen from the above, the embodiments of the present application can obtain the popularity parameter of the target tag to determine the popularity of the target tag, and then based on the search frequency of the target tag in the tag set and the number of contents of the target tag, determine the importance parameter of the target tag, that is, determine the tag posterior heat feature of the target tag. After that, based on the popularity and tag posterior heat, determine the tag recommendation reference information of the target tag. Finally, the tags of the content to be published in the content interaction platform can be determined according to the tag recommendation reference information, which can improve the tag accuracy.
[0153] Based on the content introduced above, examples will be given below to further illustrate the method for determining tags in the present application. Refer to Figure 3 , a method for determining tags, and the specific process can be as follows:
[0154] 301. The server determines tag sets corresponding to multiple contents in the content interaction platform.
[0155] In an example, multiple video accounts can be registered in the content interaction platform, and each video account can publish content in the content interaction platform. The published content can be in the forms of videos, images, audios, texts, etc. Then, the tag set refers to the set of content tags corresponding to the content published by each video account.
[0156] Among them, each content can correspond to one content tag or multiple content tags. The content tag can characterize the type of the content and can play a good role in classifying, sorting, and searching for the content.
[0157] In an example, in the content interaction platform, users can search for the content corresponding to the content tag through the content tag. As Figure 2b shown, by searching for the content tag "car", the content with the content tag "car" can be displayed on the page for users to browse. By setting corresponding content tags for the published content in the content interaction platform, the exposure rate of each content in the content interaction platform can be improved.
[0158] 302. The server determines, from the tag set, a target tag set that has a word inclusion relationship with the target tag, where the target tag is used to determine the tag of the content to be published in the content interaction platform.
[0159] In an example, there is a word inclusion relationship between the target tag and the content tags in the target tag set. For example, the target tag is the substring "flowers". Under the substring "flowers", various different words such as "express delivery", "delivery", and "packaging" can be matched around it, making its entropy larger. According to the word inclusion relationship, for example, according to the dictionary inclusion relationship, the target tag set that has an inclusion relationship with the target tag "flowers" can be found from the tag set. At this time, the target tag set can include substrings such as "flowers express delivery", "flowers delivery", and "flowers packaging".
[0160] Among them, it is easy to understand that if "flowers express delivery" is used as the target tag, the target tag set that has a word inclusion relationship with "flowers express delivery" can also be found. For example, substrings such as "flowers express delivery to the door", "flowers express delivery express", and "flowers express delivery service" are used as the target tag set.
[0161] 303. The server determines the broadness parameter of the target tag based on the tag set and the target tag set.
[0162] Among them, the target breadth parameter refers to the amount of information brought by the target label, which can be determined based on the information entropy of the target label.
[0163] In one example, the information entropy of the hashtag of H(flowers) and its breadth calculation process can be obtained through the information entropy formula as follows:
[0164] P(flowers) = (5 + 10 + 10 + 70) / 100 = 0.95
[0165] H(flowers) = -(0.95 * log(0.95)) = 0.02
[0166] Among them, 5, 10, 10, and 70 respectively represent the number of occurrences of each content label in the target label set, and 100 represents the number of occurrences of each content label in the label set.
[0167] In one example, if you want to calculate the breadth parameter of a new content label, you can obtain a new label set from the label set that has a word inclusion relationship with the new content label. Then, similarly, based on the label set and the new label set, determine the breadth parameter of the new label set. For example, it can be calculated through the information entropy formula of the new content label.
[0168] 304. The server obtains the search frequency of the target label in the label set and the importance parameter of the target label.
[0169] In one example, all the content labels in the label set can be regarded as an extremely large label article. Then, the search frequency of the target label in the label set and the importance parameter of the target label can be determined through the main idea of TFIDF. For example, the number of searches for the target label in the content interaction platform and the number of searches for the content labels in the label set in the content interaction platform can be obtained, and then the search times of the target label and the search times of the label set are fused. For example, the search times of the target label are divided by the search times of the label set to obtain the term frequency TF, that is, the search frequency of the target label in the label set can be obtained.
[0170] In one example, the number of contents using the target label in the content interaction platform can be obtained, and then the number of contents of multiple contents in the content interaction platform can be obtained. After that, the two numbers of contents can be fused to obtain the importance parameter IDF of the target label.
[0171] 305. The server determines the importance degree parameter of the target label based on the search frequency and the importance parameter.
[0172] Among them, the search frequency is the search frequency of the target tag in the tag set, and the importance parameter is a measure of the general importance of the target tag.
[0173] In one example, the search frequency and the importance parameter can be fused to obtain the importance degree parameter of the target tag. For example, the search frequency can be multiplied by the importance parameter to obtain the importance degree parameter of the target tag.
[0174] In one example, the higher the importance degree parameter of the target tag, the higher the popularity of the target tag can be indicated. It can be understood that in the content interaction platform, when users publish content through the video number, they relatively like to use the target tag as the content tag of the content to be published, and the popularity of the target tag is relatively high.
[0175] 306. Fuse the breadth parameter and the importance degree parameter of the target tag to determine the tag recommendation reference information of the target tag.
[0176] In one example, the importance degree parameter of the target tag can be logarithmically smoothed first, and then multiplicatively fused with the breadth parameter of the target tag to obtain the tag recommendation reference information of the target tag.
[0177] 307. The server determines that the tag of the content to be published in the content interaction platform is the target tag according to the tag recommendation reference information of the target tag.
[0178] In one example, each content tag in the tag set can correspond to a determined tag recommendation reference information. The tag recommendation reference information includes a tag recommendation score. The tag of the content to be published in the content interaction platform can be determined according to the tag recommendation score. For example, when the tag recommendation score reaches the preset recommendation score, the content tag corresponding to the tag recommendation score can be used as the content tag of the current content to be published.
[0179] In one example, the user can edit the corresponding content tag for the current content to be published, and then obtain the tag recommendation score corresponding to the edited content tag. Whether the edited content tag can be used as the content tag of the current content to be published can be determined according to the tag recommendation score.
[0180] In one example, the user can obtain multiple candidate content tags generated for the current content to be published, and then determine whether to determine the content tag of the current content to be published from the multiple candidate content tags based on the tag recommendation scores corresponding to the candidate content tags.
[0181] As can be seen from the above, the embodiments of the present application can obtain the breadth parameter of the target tag to determine the breadth of the target tag, and then determine the importance parameter of the target tag, that is, determine the posterior heat feature of the target tag, based on the search frequency of the target tag in the tag set and the number of contents of the target tag. After that, the tag recommendation reference information of the target tag is determined based on the breadth and the posterior heat of the tag, and finally, the tags of the content to be published in the content interaction platform can be determined according to the tag recommendation reference information, which can improve the tag accuracy.
[0182] To better implement the above method, correspondingly, the embodiments of the present application further provide a tag determination device, where the tag determination device can be specifically integrated in a server. Refer to Figure 4a , the tag determination device may include a first tag set determination unit 401, a second tag set determination unit 402, a first parameter determination unit 403, a first acquisition unit 404, a second parameter determination unit 405, a fusion unit 406, and a first tag determination unit 407, as follows:
[0183] (1) The first tag set determination unit 401;
[0184] The first tag set determination unit 401 is configured to determine a tag set corresponding to multiple contents in the content interaction platform, and the tag set includes content tags corresponding to the contents, where the content tags represent the content types of the contents.
[0185] In one embodiment, as Figure 4b shown, the first tag determination unit 401 includes:
[0186] A score determination subunit 4011, configured to determine a tag recommendation score of the target tag according to the tag recommendation reference information of the target tag;
[0187] A tag determination subunit 4012, configured to determine the tag of the content to be published in the content publishing platform as the target tag when the tag recommendation score reaches a preset recommendation score.
[0188] (2) The second tag set determination unit 402;
[0189] The second tag set determination unit 402 is configured to determine a target tag set having a word inclusion relationship with the target tag from the tag set, and the target tag is used to determine the tag of the content to be published in the content interaction platform.
[0190] (3) The first parameter determination unit 403;
[0191] The first parameter determination unit 403 is configured to determine a breadth parameter of the target tag based on the tag set and the target tag set, and the breadth parameter represents the amount of information of the target tag.
[0192] In one embodiment, as Figure 4c shown, the first parameter determination unit 403 includes:
[0193] A quantity determination subunit 4031, configured to determine the quantity of labels in the label set and the target label set in the content interaction platform;
[0194] A first calculation subunit 4032, configured to calculate the information breadth of the target label based on the quantity of labels, and obtain the breadth parameter of the target label.
[0195] In one embodiment, the first calculation subunit 4032 is further configured to obtain the probability information of the target label set in the label set based on the quantity of labels; according to the probability information, obtain the information entropy of the target label, and determine the breadth parameter of the target label.
[0196] In one embodiment, as Figure 4c shown, the first parameter determination unit 403 includes:
[0197] An information determination subunit 4033, configured to determine the search information of the labels in the label set and the target label set, where the search information includes the search times of the labels in the label set and the target search times of the labels in the target label set;
[0198] A second calculation subunit 4034, configured to calculate the breadth parameter of the target label based on the search times and the target search times.
[0199] (4) The first acquisition unit 404;
[0200] The first acquisition unit 404 is configured to acquire the search frequency of the target label in the label set.
[0201] In one embodiment, as Figure 4d shown, the first acquisition unit 404 includes:
[0202] A first fusion subunit 4041, configured to fuse the search times of the target label and the search times of the labels in the label set to obtain the search frequency of the target label.
[0203] (5) The second parameter determination unit 405;
[0204] The second parameter determination unit 405 is configured to determine the importance degree parameter of the target label based on the search frequency and the quantity of content corresponding to the target label, where the importance degree parameter represents the importance degree of the target label in the label set.
[0205] In one embodiment, as Figure 4e shown, the second parameter determination unit 405 includes:
[0206] An acquisition subunit 4051, configured to acquire the quantity of multiple contents;
[0207] A second fusion subunit 4052, configured to fuse the number of contents of multiple contents with the number of contents corresponding to the target label to determine an importance parameter of the target label;
[0208] A third fusion subunit 4053, configured to fuse the search frequency and the importance parameter to obtain an importance degree parameter of the target label.
[0209] (6) A fusion unit 406;
[0210] The fusion unit 406 is configured to fuse the breadth parameter and the importance degree parameter to determine label recommendation reference information of the target label.
[0211] (7) A first label determination unit 407;
[0212] The first label determination unit 407 is configured to determine a label of the content to be published in the content interaction platform according to the label recommendation reference information of the target label.
[0213] In an embodiment, the label determination device further includes:
[0214] A second acquisition unit 408, configured to acquire the edited content of the label editing operation when detecting a label editing operation of a user for the current content to be published, where the current content to be published is the content to be published in the content publishing platform;
[0215] A first recognition unit 409, configured to recognize the edited content to obtain label recommendation reference information of the edited content;
[0216] A second label determination unit 410, configured to determine a label of the current content to be published based on the label recommendation reference information of the edited content.
[0217] In an embodiment, the label determination device further includes:
[0218] A second recognition unit 411, configured to perform content recognition on the current content to be published when detecting a label acquisition operation of a user for the current content to be published, to obtain a content recognition result;
[0219] A third acquisition unit 412, configured to acquire a content matching label matching the current content to be published from a label library based on the content recognition result of the current content to be published;
[0220] A third label determination unit 413, configured to determine a label of the current content to be published from the content matching labels based on the label recommendation reference information of the content matching labels.
[0221] As can be seen from the above, the first tag set determination unit 401 of the tag determination device according to the embodiments of the present application determines the tag sets corresponding to multiple contents in the content interaction platform. The tag set includes the content tags corresponding to the contents, where the content tags represent the content types of the contents. Then, the second tag set determination unit 402 determines, from the tag sets, a target tag set that has a word inclusion relationship with the target tag, where the target tag is used to determine the tags of the content to be published in the content interaction platform. The first parameter determination unit 403 determines the breadth parameter of the target tag based on the tag set and the target tag set, where the breadth parameter represents the amount of information of the target tag. The first acquisition unit 404 acquires the search frequency of the target tag in the tag set. The second parameter determination unit 405 determines the importance parameter of the target tag based on the search frequency and the number of contents corresponding to the target tag, where the importance parameter represents the importance of the target tag in the tag set. The fusion unit 406 fuses the breadth parameter and the importance parameter to determine the tag recommendation reference information of the target tag. The first tag determination unit 407 determines the tags of the content to be published in the content interaction platform according to the tag recommendation reference information of the target tag. This solution can obtain the breadth parameter of the target tag to determine the breadth of the target tag, and then determine the importance parameter of the target tag based on the search frequency of the target tag in the tag set and the number of contents of the target tag, that is, determine the posterior heat feature of the tag of the target tag. Then, based on the breadth and the posterior heat of the tag, the tag recommendation reference information of the target tag is determined. Finally, the tags of the content to be published in the content interaction platform can be determined according to the tag recommendation reference information, which can improve the tag accuracy.
[0222] In addition, the embodiments of the present application further provide a computer device, which can be a device such as a terminal or a server, as Figure 5 shown, which shows a schematic structural diagram of the computer device involved in the embodiments of the present application. Specifically:
[0223] The computer device may include a processor 501 with one or more processing cores, a memory 502 with one or more storage media, a power supply 503, an input unit 504, and other components. Those skilled in the art can understand that Figure 5 the structural diagram of the computer device shown in
[0224] does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or different component arrangements.
[0225] The processor 501 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 502, and by invoking the data stored in the memory 502, it executes various functions of the computer device and processes data, thereby performing an overall detection of the computer device. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 501 either.
[0226] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store the data created according to the use of the computer device. In addition, the memory 502 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 502 can also include a memory controller to provide the processor 501 with access to the memory 502.
[0227] The computer device also includes a power supply 503 that powers each component. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, thereby realizing functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 503 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0228] The computer device may also include an input unit 504, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0229] Although not shown, the computer device may also include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 501 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 502 according to the following instructions, and the processor 501 will run the application programs stored in the memory 502 to achieve various functions as follows:
[0230] Determine the tag sets corresponding to multiple contents in the content interaction platform. The tag sets include content tags corresponding to the contents, where the content tags characterize the content types of the contents; from the tag sets, determine the target tag set that has a word inclusion relationship with the target tag, and the target tag is used to determine the tags of the content to be published in the content interaction platform; based on the tag set and the target tag set, determine the breadth parameter of the target tag, and the breadth parameter characterizes the amount of information of the target tag; obtain the search frequency of the target tag in the tag set; based on the search frequency and the number of contents corresponding to the target tag, determine the importance parameter of the target tag, and the importance parameter characterizes the importance of the target tag in the tag set; fuse the breadth parameter and the importance parameter to determine the tag recommendation reference information of the target tag; according to the tag recommendation reference information of the target tag, determine the tags of the content to be published in the content interaction platform.
[0231] As can be seen from the above, the embodiments of the present application can obtain the breadth parameter of the target tag to determine the breadth of the target tag, and then based on the search frequency of the target tag in the tag set and the number of contents of the target tag, determine the importance parameter of the target tag, that is, determine the posterior heat feature of the tag of the target tag. Then, based on the breadth and the posterior heat of the tag, determine the tag recommendation reference information of the target tag. Finally, the tags of the content to be published in the content interaction platform can be determined according to the tag recommendation reference information, which can improve the tag accuracy.
[0232] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a storage medium and loaded and executed by a processor.
[0233] Therefore, the embodiments of the present application provide a storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps in any one of the tag determination methods provided by the embodiments of the present application. For example, the instructions can execute the following steps:
[0234] Determine the tag set corresponding to multiple contents in the content interaction platform. The tag set includes the content tags corresponding to the contents, where the content tags represent the content types of the contents; determine, from the tag set, the target tag set that has a word inclusion relationship with the target tag, where the target tag is used to determine the tags of the content to be published in the content interaction platform; based on the tag set and the target tag set, determine the breadth parameter of the target tag, where the breadth parameter represents the amount of information of the target tag; obtain the search frequency of the target tag in the tag set; based on the search frequency and the number of contents corresponding to the target tag, determine the importance parameter of the target tag, where the importance parameter represents the importance of the target tag in the tag set; fuse the breadth parameter and the importance parameter to determine the tag recommendation reference information of the target tag; according to the tag recommendation reference information of the target tag, determine the tags of the content to be published in the content interaction platform.
[0235] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0236] Since the instructions stored in the storage medium can execute the steps in any of the tag determination methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any of the tag determination methods provided in the embodiments of the present application can be realized. For details, see the previous embodiments and will not be elaborated here.
[0237] Among them, according to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the tag determination method provided in the above-mentioned invention content and embodiments.
[0238] The above has introduced in detail a tag determination method, device, computer device and storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for determining a label, characterized in that, Including: Determine a tag set corresponding to multiple contents in the content interaction platform, where the tag set includes content tags corresponding to the contents, and the content tags characterize the content types of the contents; Determine, from the tag set, a target tag set having a word inclusion relationship with a target tag, where the target tag is used to determine the tags of the content to be published in the content interaction platform; Determine the quantities of the tags in the tag set and the tags in the target tag set used in the content interaction platform; Based on the quantities of the tags, obtain the probability information of the target tag set in the tag set; According to the probability information, obtain the information entropy of the target tag, and determine the breadth parameter of the target tag, where the breadth parameter characterizes the amount of information of the target tag; Obtain the search frequency of the target tag in the tag set; Based on the search frequency and the quantity of the content corresponding to the target tag, determine the importance parameter of the target tag, where the importance parameter characterizes the importance degree of the target tag in the tag set; Fuse the breadth parameter and the importance parameter to determine the tag recommendation reference information of the target tag; According to the tag recommendation reference information of the target tag, determine the tags of the content to be published in the content interaction platform.
2. The method according to claim 1, wherein The method further includes: Determine the search information of the tags in the tag set and the target tag set, where the search information includes the search times of the tags in the tag set and the target search times of the tags in the target tag set; Based on the search times and the target search times, calculate the breadth parameter of the target tag.
3. The method according to claim 2, wherein The search information further includes the search times of the target tag; The obtaining the search frequency of the target tag in the tag set includes: Fuse the search times of the target tag and the search times of the tags in the tag set to obtain the search frequency of the target tag; The determining the importance parameter of the target tag based on the search frequency and the quantity of the content corresponding to the target tag includes: Obtain the quantities of the multiple contents; Fuse the quantities of the multiple contents and the quantity of the content corresponding to the target tag to determine the importance parameter of the target tag; Fuse the search frequency and the importance parameter to obtain the importance parameter of the target tag.
4. The method according to claim 1, characterized in that, The determining the tags of the content to be published in the content interaction platform according to the tag recommendation reference information of the target tag includes: According to the tag recommendation reference information of the target tag, determine the tag recommendation score of the target tag; When the tag recommendation score reaches a preset recommendation score, determine that the tag of the content to be published in the content publishing platform is the target tag.
5. The method according to claim 1, characterized in that, The method further includes: When detecting a tag editing operation by a user for the current content to be published, obtain the editing content of the tag editing operation, where the current content to be published is the content to be published in the content publishing platform; Identify the editing content to obtain the tag recommendation reference information of the editing content; Determine the tags of the current content to be published based on the tag recommendation reference information of the edited content.
6. The method according to claim 5, wherein The method further includes: When it is detected that the user performs a tag acquisition operation on the current content to be published, perform content recognition on the current content to be published to obtain a content recognition result; Based on the content recognition result of the current content to be published, obtain content matching tags that match the current content to be published from the tag library; Based on the tag recommendation reference information of the content matching tags, determine the tags of the current content to be published from the content matching tags.
7. A label determination device, characterized in that, Includes: A first tag set determination unit for determining a tag set corresponding to multiple contents in the content interaction platform, where the tag set includes content tags corresponding to the content, and the content tags represent the content types of the content; A second tag set determination unit for determining, from the tag set, a target tag set having a word inclusion relationship with the target tag, where the target tag is used to determine the tags of the content to be published in the content interaction platform; A first parameter determination unit for determining the number of tags in the tag set and the target tag set used in the content interaction platform; based on the number of tags, obtain the probability information of the target tag set in the tag set; according to the probability information, obtain the information entropy of the target tag, and determine the broadness parameter of the target tag, where the broadness parameter represents the amount of information of the target tag; A first acquisition unit for acquiring the search frequency of the target tag in the tag set; A second parameter determination unit for determining the importance parameter of the target tag based on the search frequency and the number of contents corresponding to the target tag, where the importance parameter represents the importance of the target tag in the tag set; A fusion unit for fusing the broadness parameter and the importance parameter to determine the tag recommendation reference information of the target tag; A first tag determination unit for determining the tags of the content to be published in the content interaction platform according to the tag recommendation reference information of the target tag.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the steps in the tag determination method according to any one of claims 1 to 6.
9. A storage medium storing multiple instructions applicable to be loaded by a processor to execute the steps in the tag determination method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Relational data label cleaning method and device, equipment and storage medium
CN111177132A