Recommended content determination methods, media, apparatuses, and computing devices
Patent Information
- Application Number
- CN202310730529.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-19
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2043-06-19
AI Technical Summary
[0005]本公开提供一种推荐内容确定方法、介质、装置和计算设备,以解决相关技术中对用户兴趣偏好确定的全面性较差的问题
[0019]根据本公开实施方式的推荐内容确定方法、介质、装置和计算设备,通过基于预先确定的参考内容和至少一种内容拓展方式,确定每种内容拓展方式对应的拓展内容,再确定拓展内容对应的关联对象,并向关联对象投放拓展内容,以得到每种内容拓展方式对应的拓展内容的投放效率,最后基于投放效率和拓展内容,确定用于推送的推荐内容。由此,通过不同的内容拓展方式,可以更全面的挖掘用户潜在需求的内容,进而更好的满足用户的需求;同时通过根据投放效率优化拓展内容,能有效保证最终推送给用户的推荐内容具有较高的转化效率,显著缩短了在向用户推荐内容时的兴趣探索过程,显著减弱了探索过程用户体验下降的问题,从而同时保证了推荐内容的转化效率和用户体验。
Smart Images

Figure CN116821488B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure relate to the field of Internet technology, and more specifically, the embodiments of this disclosure relate to a method, medium, apparatus, and computing device for determining recommended content. Background Technology
[0002] This section is intended to provide background or context for the embodiments of this disclosure as set forth in the claims. The description herein is not intended to be related art simply because it is included in this section.
[0003] In related technologies, content platforms need to determine the content to push to users based on their interests and preferences in order to better meet user experience. Therefore, it is necessary to accurately and comprehensively determine users' interests and preferences to ensure that recommended content meets user needs.
[0004] Existing methods for determining recommended content typically rely on user interaction records with various content on content platforms. By converting the tags of each content in the interaction records into word vectors, and using word vectors to explore similarity, user interests and preferences are determined. Recommended content is then tentatively determined based on these interests and preferences. However, this approach lacks a systematic approach, fails to fully reflect user interests, and requires constantly trying different recommended content to obtain more interaction records to ensure the accuracy of the explored interests and preferences. This can easily lead to low conversion rates for explored traffic and a decline in user experience. Summary of the Invention
[0005] This disclosure provides a method, medium, apparatus, and computing device for determining recommended content, in order to solve the problem of poor comprehensiveness in determining user interests and preferences in related technologies.
[0006] In a first aspect of this disclosure, a method for determining recommended content is provided, comprising:
[0007] Based on predetermined reference content and at least one content expansion method, determine the expansion content corresponding to each content expansion method;
[0008] Identify the associated objects corresponding to the extended content, and deliver the extended content to the associated objects to obtain the delivery efficiency of the extended content for each content extension method;
[0009] Based on delivery efficiency and expanded content, determine the recommended content to be pushed.
[0010] In a second aspect of this disclosure, a computer-readable storage medium is provided, comprising:
[0011] A computer-readable storage medium stores computer-executable instructions that, when executed by a processor, are used to implement the recommended content determination method as described in the first aspect of this disclosure.
[0012] In a third aspect of this disclosure, a recommended content determination apparatus is provided, comprising:
[0013] The preparation module is used to determine the extended content corresponding to each content extension method based on the pre-determined reference content and at least one content extension method.
[0014] The delivery module is used to determine the associated objects corresponding to the extended content and deliver the extended content to the associated objects in order to obtain the delivery efficiency of the extended content for each content extension method.
[0015] The determination module is used to determine the recommended content to be pushed based on delivery efficiency and expanded content.
[0016] In a fourth aspect of this disclosure, a computing device is provided, comprising: at least one processor;
[0017] and memory that is communicatively connected to at least one processor;
[0018] The memory stores instructions executable by at least one processor, which, when executed by at least one processor, cause the computing device to perform the recommended content determination method as described in the first aspect of this disclosure.
[0019] The recommended content determination method, medium, apparatus, and computing device according to embodiments of this disclosure determine the extended content corresponding to each content expansion method based on pre-determined reference content and at least one content expansion method. Then, the associated objects corresponding to the extended content are determined, and the extended content is delivered to the associated objects to obtain the delivery efficiency of the extended content for each content expansion method. Finally, based on the delivery efficiency and the extended content, recommended content for push notifications is determined. Therefore, by using different content expansion methods, the potential needs of users can be more comprehensively explored, thereby better meeting user needs. Simultaneously, by optimizing the extended content according to the delivery efficiency, it is possible to effectively ensure that the recommended content finally pushed to users has a high conversion efficiency, significantly shortening the interest exploration process when recommending content to users, and significantly mitigating the problem of decreased user experience during the exploration process, thus simultaneously ensuring the conversion efficiency and user experience of the recommended content. Attached Figure Description
[0020] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which:
[0021] Figure 1 An application scenario diagram illustrating an embodiment of the present disclosure is shown schematically;
[0022] Figure 2 A flowchart illustrating a method for determining recommended content according to another embodiment of this disclosure is shown schematically;
[0023] Figure 3a A flowchart illustrating a method for determining recommended content according to another embodiment of this disclosure is shown schematically;
[0024] Figure 3b schematically shown Figure 3a The flowchart of the method for determining the extension content corresponding to the first type of extension provided in the embodiment shown;
[0025] Figure 3c schematically shown Figure 3a The flowchart of the method for determining extended words using word vectors provided in the illustrated embodiment is shown.
[0026] Figure 3d schematically shown Figure 3a The flowchart of the method for determining extended words using a knowledge graph provided in the embodiment shown is as follows;
[0027] Figure 3e schematically shown Figure 3a The flowchart of the method for determining extended words by combining knowledge graph and word vectors provided in the embodiment shown is shown.
[0028] Figure 3f schematically shown Figure 3a The flowchart of the method for determining the extension content corresponding to the second type of extension provided in the embodiment shown;
[0029] Figure 3g schematically shown Figure 3a The flowchart of the method for determining the location of topic words provided in the embodiment shown;
[0030] Figure 3h schematically shown Figure 3a The flowchart of the method for determining extended content by the location and corresponding content of topic words provided in the embodiment shown;
[0031] Figure 4a A flowchart illustrating a method for determining recommended content according to another embodiment of this disclosure is shown schematically;
[0032] Figure 4b schematically shown Figure 4a The flowchart of the method for calculating delivery efficiency provided in the illustrated embodiment is shown.
[0033] Figure 4c schematically shown Figure 4a The flowchart of the method for determining the maximum delivery efficiency provided in the illustrated embodiment;
[0034] Figure 5 A schematic diagram of the structure of a storage medium according to another embodiment of the present disclosure is shown;
[0035] Figure 6 A schematic diagram of the structure of a recommendation content determination device according to another embodiment of the present disclosure is shown;
[0036] Figure 7 A schematic diagram of the structure of a computing device according to another embodiment of the present disclosure is shown.
[0037] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation
[0038] The principles and spirit of this disclosure will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are given merely to enable those skilled in the art to better understand and implement this disclosure, and are not intended to limit the scope of this disclosure in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.
[0039] Those skilled in the art will recognize that embodiments of this disclosure can be implemented as a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0040] According to embodiments of this disclosure, a method, medium, apparatus, and computing device for determining recommended content are proposed.
[0041] In this document, it should be understood that the terminology used is for convenience of understanding only and does not imply any limitation on its meaning. Furthermore, any number of elements in the accompanying drawings is for illustrative purposes only and not for limitation, and any naming is for distinction only and has no limiting meaning.
[0042] The following is a description of the terminology used in this disclosure:
[0043] Word vectors: a collective term for a set of language modeling and feature learning techniques in Natural Language Processing (NLP), and in this disclosure, it refers to a vector consisting of a set of real numbers that are mapped to words, phrases, or tags.
[0044] Knowledge graph: A graphical database that uses visualization technology to describe the relationships between content resources corresponding to different tags. It can display the relationships between different tags in the form of a graph.
[0045] Recommended content: A pre-determined set of content based on user interests and preferences. In actual push notifications, a portion of the recommended content that meets the user's selected function (such as "Daily Recommendations") will be selected and pushed to the user.
[0046] In this document, it should be understood that the terminology used is for convenience of understanding only and does not imply any limitation on its meaning. Furthermore, any number of elements in the accompanying drawings is for illustrative purposes only and not for limitation, and any naming is for distinction only and has no limiting meaning.
[0047] In addition, the data involved in this disclosure may be data authorized by the user or fully authorized by all parties. The collection, dissemination and use of the data shall comply with the requirements of relevant national laws and regulations. The implementation methods / executives of this disclosure may be combined with each other. Invention Overview
[0049] The inventors have discovered that in related technologies, content platforms need to determine the content to push to users based on their interests and preferences in order to better meet user experience. Therefore, it is necessary to accurately and comprehensively determine users' interests and preferences to ensure that recommended content meets user needs.
[0050] Existing recommendation systems first need to determine the content to recommend before recommending specific content to users. For content that users have expressed interest in, related content can be directly recommended to users. However, for content that users have not expressed interest in, different methods need to be used to explore it, that is, to determine the content to recommend.
[0051] Existing methods for determining recommended content typically rely on user interaction records with various content on content platforms. By converting the tags of each content in the interaction records into word vectors, and using word vectors to explore similarity, user interests and preferences are determined. Recommended content is then tentatively determined based on these interests and preferences. However, this approach lacks a systematic approach, depends solely on the similarity exploration of content tags, involves a limited range of content types, and cannot fully reflect user interests. Furthermore, it cannot guarantee that all explored content will satisfy user interests and preferences. Therefore, it is necessary to continuously try recommending different content to obtain more interaction records to ensure the accuracy of the interest preferences corresponding to the explored content. This can easily lead to low efficiency in converting exploration traffic into user interaction feedback and result in a poor user experience.
[0052] The recommended content determination method disclosed herein determines the corresponding extended content by pre-determining reference content and different content expansion methods. Then, by delivering these extended contents to related objects and adjusting and optimizing the delivery efficiency, the recommended content to be pushed to users can be determined from these extended contents. Thus, the extended content obtained by different methods can be comprehensively determined, and the delivery efficiency can be effectively guaranteed, thereby ensuring the user experience when pushing to users.
[0053] After introducing the basic principles of this disclosure, various non-limiting embodiments of this disclosure will be described in detail below.
[0054] Application Scenarios Overview
[0055] First refer to Figure 1 The diagram illustrates an application scenario of the recommended content determination method provided in this disclosure. When determining recommended content, server 110 identifies the target user 120 corresponding to the recommended content and determines the associated extended content 140 based on the reference content 130 corresponding to the target user 120, thereby determining the extended content 140 as recommended content and thus realizing the recommended content determination process.
[0056] It should be noted that, Figure 1 In the scenario shown, only one or two of the server, target user, reference content, and extended content are used as examples for illustration, but this disclosure is not limited to this. That is to say, the number of servers, target users, reference content, and extended content can be arbitrary.
[0057] The following is combined with Figure 1 Application scenarios, refer to Figures 2 to 4a This document describes a method for determining recommended content according to exemplary embodiments of this disclosure. It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of this disclosure, and the embodiments of this disclosure are not limited in any way. Rather, the embodiments of this disclosure can be applied to any applicable scenario.
[0058] Figure 2 A flowchart illustrating a method for determining recommended content according to an embodiment of this disclosure. Figure 2 As shown, the method for determining recommended content provided in this embodiment includes the following steps:
[0059] Step S201: Based on the predetermined reference content and at least one content expansion method, determine the expansion content corresponding to each content expansion method.
[0060] Specifically, the recommended content for different users usually varies to some extent. However, if the content that different users are interested in (or the content they have interacted with, such as content they have liked or shared) has the same interest tags (such as the "fishing" tag) or content tags (such as the "piano performance video" tag), then the recommended content for these users is usually quite similar.
[0061] To improve the efficiency of content recommendation, users with a set number of common interest tags or content tags (often hundreds or thousands of such users) are identified as users for whom corresponding recommended content can be directly determined. For ease of description later, these users are collectively referred to as target users (e.g., users with tags such as "Mandarin", "rap", "music", and "underground" can be considered as the same type of target users and recommended the same or similar content).
[0062] For each type of target user, similar or identical recommended content can be determined, while for different types of target users, the recommended content usually differs significantly.
[0063] The determination of target users is usually based on the popularity of interest tags (or content tags, hereinafter referred to as tags for convenience). The number of tags to be set is determined by the popularity of the tags. If the popularity of the tags is high (such as multiple tags for singers with high interaction, such as "Hong Kong and Taiwan singers", "male singers", "folk songs", etc.), then a larger number of tags are set. If there are seven or more users with the same tags, they can be identified as a group of target users. If the popularity of the tags is low (such as video theme tags with very low interaction, such as "complex function" and "shock fluid"), then a smaller number of tags are set. If there are three users with the same tags, then they can be identified as a group of target users.
[0064] A user typically belongs to only one category of target users, but for users who have interacted with a large number of tags, they may belong to two or more categories of target users at the same time.
[0065] For each type of target user, the content that each user has interacted with can reflect the user's interests and preferences to some extent. Therefore, the content that the target user has interacted with can be identified as reference content (i.e., content that can help determine the user's interests and preferences). Alternatively, the content that the target user has interacted with can be identified as reference content if the number of interactions (or the proportion of interactions) reaches a set interaction threshold (e.g., content that at least 1 / 5 of the target users have interacted with). This ensures that the reference content can better match the target user's interests and preferences.
[0066] Therefore, before determining the reference content, it is first necessary to identify the target users of the content to be recommended, so as to determine the corresponding reference content based on the target users.
[0067] Content expansion methods mainly include two categories: content expansion based on the similarity or relevance of tags in the reference content (e.g., based on the "orchestral" tag in the reference content, content with the "classical music" tag that is similar to it is identified as expanded content), and content expansion based on the relevance of the tags to the corresponding topics (e.g. based on the related topics of the "orchestral" tag, the name of a rock singer in a news article involving "orchestral" is used as a tag, and the content corresponding to that tag is identified as expanded content; obviously, the singer and "orchestral" itself are not similar, but they are related in terms of topic). Each category contains different specific methods depending on the specific identification method, such as identification based on word vector similarity (e.g., "keyboard instruments" and "keyboard music"), identification based on the similarity of tags extracted from the context of the reference content (e.g., "classical novels" and "Dream of the Red Chamber"), identification based on highly relevant topics (e.g., "folk songs" and "music festivals"), or identification based on the inclusion relationship between topics (e.g., "rap" and "West Coast rap").
[0068] Based on the different content expansion methods, the tags are called expanded content tags. Content with these expanded content tags can be identified as expanded content.
[0069] Step S202: Determine the associated objects corresponding to the extended content, and deliver the extended content to the associated objects to obtain the delivery efficiency of the extended content corresponding to each content extension method.
[0070] Specifically, since the determined extended content comes from different extended methods and there are significant differences between them (for example, extended content may contain content tagged "Suzhou Pingtan" and content tagged "Mexican food"), it is necessary to determine the proportion of these different extended content to be pushed (for example, if the target user is more interested in content tagged "latest popular songs", then not all the content pushed to them should have that tag, but only a certain proportion of the content should have that tag) in order to meet the needs of users with different interests.
[0071] Among them, the user objects used for testing to determine the proportion of different extended content pushes are called associated objects. Associated objects are usually users who have interacted with the reference content (not necessarily meeting the requirement of a set number of common tags among target users), that is, user objects that are associated with the reference content.
[0072] The process of delivering extended content to related objects involves selecting a subset of associated objects and pushing them a certain proportion of different extended content.
[0073] After delivering extended content to related objects, the delivery efficiency of each type of extended content can be determined based on the interaction ratio of related objects within a set time period (such as 24 hours). This efficiency reflects the interest and preference of related objects for the extended content.
[0074] Step S203: Based on delivery efficiency and expanded content, determine the recommended content to be pushed.
[0075] Specifically, the distribution ratio of various extended content types is adjusted based on their distribution efficiency. Then, the content is distributed to other related users and the distribution efficiency is tested. By repeating this process, the overall distribution efficiency can be maximized. The distribution ratio of various extended content types at this point is determined as the distribution ratio when pushing extended content to target users. Then, the extended content and distribution ratio can be combined as the recommended content pushed to target users.
[0076] In other words, when it is necessary to push recommended content to target users, the extended content is delivered according to a certain delivery ratio, which can ensure the highest delivery efficiency, that is, the best satisfaction with the target users' interests and preferences (because the target users' interests and preferences have a high degree of similarity with related objects), thereby ensuring the target users' satisfaction with the pushed recommended content, while reducing the interest testing process when delivering recommended content, and thus improving the user experience of the target users.
[0077] Since the number of associated objects selected for testing is usually much smaller than the target users (for example, if there are 10,000 target users, 1,000 associated objects can usually be selected), and the content delivered to the associated objects is also content that is related to or similar to their interests and preferences, and each associated object is usually only delivered extended content once, compared to the process of multiple interest probing in related technologies, the probing process can be reduced, thereby ensuring the user experience of the associated objects.
[0078] According to the recommended content determination method of this disclosure, based on pre-determined reference content and at least one content expansion method, the expanded content corresponding to each content expansion method is determined. Then, the associated objects corresponding to the expanded content are determined, and the expanded content is delivered to the associated objects to obtain the delivery efficiency of the expanded content corresponding to each content expansion method. Finally, based on the delivery efficiency and the expanded content, the recommended content for push notification is determined. Therefore, by using different content expansion methods, the potential needs of users can be more comprehensively explored, thereby better meeting user needs. At the same time, by optimizing the expanded content according to the delivery efficiency, it is possible to effectively ensure that the recommended content finally pushed to users has a high conversion efficiency, significantly shortening the interest exploration process when recommending content to users, and significantly mitigating the problem of decreased user experience during the exploration process, thus simultaneously ensuring the conversion efficiency and user experience of the recommended content.
[0079] Figure 3a A flowchart illustrating a method for determining recommended content provided in one embodiment of this disclosure. Figure 3a As shown, the method for determining recommended content provided in this embodiment includes the following steps:
[0080] Step S301: Use the tags corresponding to the reference content as reference words for the reference content.
[0081] The content expansion methods include the first type of expansion based on reference words and the second type of expansion based on topic trees.
[0082] Specifically, the determination of extended content is mainly based on tag expansion. In this embodiment, subsequent expansion can be based on the tags of the reference content.
[0083] Reference content consists of user-interacted material, such as liked music, saved videos, and commented articles. This reference content is pre-tagged on the content platform, so these tags can be directly used as reference words for subsequent expansion. By identifying similar expansion tags to the reference words, the content corresponding to these tags can be further identified as expansion content.
[0084] Topic trees are systems that generate topics (or topic keywords) based on real-time content collection from content platforms. Through topic trees, the relationships between different topics (or topic keywords) can be determined, such as mutual inclusion (e.g., two topic keywords are the singer and the title of their released song), similarity (e.g., two topic keywords are the titles of two songs performed in a concert), relevance (e.g., two topic keywords are the titles of two works by two singers associated with the same composer), and no direct relationship (e.g., two composers and lyricists from different countries with no related news).
[0085] Based on the topic tree, we can identify extended tags that are related to the main tags in terms of topics, and then determine the corresponding extended content.
[0086] Step S302: Based on the extended words corresponding to the reference words, determine the extended content corresponding to the first type of extension.
[0087] Specifically, the first type of expansion is based on similar expansion tags of reference words, and content with these expansion tags is identified as expanded content.
[0088] These extended content items are usually similar to the reference content. For example, if the reference content is the lyrics and translation of a country song, and the reference term is "American country music," the similar extended tag might be "American folk music." The resulting extended content might also be articles related to American folk music, which will have a high degree of similarity to the reference content itself, and users will usually be more inclined to read them.
[0089] like Figure 3b The diagram shown is a flowchart of the method for determining the extended content corresponding to the first type of extension. It specifically includes the following steps:
[0090] Step S3021: Determine the extended words based on the word vectors corresponding to the reference words.
[0091] Specifically, there are different methods for the first type of expansion. One method is to expand through word vectors and vector similarity.
[0092] The conversion of word vectors can be achieved by any existing word vector conversion model, such as word2vec, glove, ELMo, BERT, etc., without any restrictions here.
[0093] In one embodiment of this disclosure, such as Figure 3c The diagram shows a flowchart of a method for determining extended words using word vectors. The specific steps include:
[0094] Step A1: Determine the position of the corresponding word in the reference content.
[0095] Specifically, the similarity tags used to identify words similar to the reference words are usually words that have appeared in the reference content (the corresponding tags are obtained based on the words), rather than searching for similar words on the content platform, in order to narrow the search scope and improve search efficiency.
[0096] Because the reference content may vary in length, such as a long article, searching for similar tags in the entire text would result in an excessively broad search scope. Furthermore, when the reference content is long, the meaning of similar words may differ significantly from the reference word (e.g., if the reference content is an article introducing a band, the reference word "End Band" is the band name, but the corresponding article may contain the similar word "End." If content tagged with "End" is used as extended content, it will differ too much from the reference content, and recommending it to users would significantly reduce the user experience). Therefore, similar words in nearby words are usually identified as similar tags (i.e., extended words) based on the position of the reference word in the reference content.
[0097] Step A2: Select words whose distance to the word position is within the set range as candidate words.
[0098] Specifically, the range can be defined as a range where the distance to the specified word position is less than a set number of characters, or it can be defined based on the length of the reference content (e.g., the distance to the word position is less than 1 / 20 of the total number of characters in the article).
[0099] There are usually many candidate words. For example, if there are 50 words whose distance from the position of the reference word is within a set range, then all 50 words are candidate words.
[0100] Step A3: Calculate the similarity between the word vector of the reference word and the word vector of the candidate word.
[0101] Specifically, for the selected candidate words, the word vector of each candidate word will be calculated separately, and the similarity with the word vector of the reference word will be calculated.
[0102] The specific similarity calculation method can be any vector similarity calculation method, and there are no restrictions here.
[0103] Step A4: Select candidate words whose similarity falls within the set vector similarity range and determine them as expanded words.
[0104] Specifically, the set vector similarity range is usually predetermined. If there are multiple candidate words with similarity within the set vector similarity range, these candidate words can all be used as expanded words (that is, the aforementioned similarity tags).
[0105] If there are no candidate words whose similarity falls within the set vector similarity range, it can be directly determined that no corresponding extended content has been obtained through this expansion method, without needing to adjust the vector similarity range to ensure the similarity between the obtained extended content and the reference content. Alternatively, the vector similarity range and / or the distance to the word position can be adjusted, and candidate words and extended words can be re-determined to increase the number of extended content, improve the amount of recommended extended content, and meet users' needs for a large number of push notifications.
[0106] Step S3022: Based on the predetermined knowledge graph and reference words, determine the extended words.
[0107] Specifically, knowledge graph-based content expansion is another concrete implementation of reference word-based expansion. The difference between it and word vector-based expansion is that knowledge graph-based expansion is based on the similarity between the reference word and the tags in the knowledge graph (which can be understood as a database), while word vector-based expansion is based on the similarity between the reference word and the corresponding tags of adjacent words in the reference content.
[0108] This step is an optional step parallel to step S3021, and those skilled in the art can choose to perform any step according to the actual situation.
[0109] In one embodiment of this disclosure, such as Figure 3d The diagram shown is a flowchart of a method for determining extended words using a knowledge graph. The specific steps include:
[0110] Step B1: Determine the triplet features of the reference word and determine the word vectors corresponding to the triplet features.
[0111] Among them, the triplet feature is used to represent the vector relationship between the subject, relative pronoun, and object of the reference word.
[0112] Specifically, triplet features are used to describe different entities and their relationships in word vectors (subject and object are entities, and relation words are the relationships between these two entities). For example, if the reference word is "jazz album," then "jazz" is the subject entity, "album" is the object entity, and its relation word is "genre" or "style." Thus, the entities in triplet features can usually be directly extracted from the reference word, while the relation words are usually determined based on pre-labeled or pre-set labeling rules (for words in specific domains, the rules for determining their relation words can be pre-entered, such as the relation word for a word composed of a music genre word and a work type word being "genre").
[0113] Step B2: Words whose word vectors corresponding to triple features in a pre-determined knowledge graph have a similarity within a set triple similarity range are identified as extended words.
[0114] Specifically, after converting each part of the triplet features into word vectors, the similarity of vectors is calculated for each pair of parts. Then, the similarities of the three parts are combined (e.g., by calculating the mean, weighted average, product, summation, etc.) to obtain the similarity between the word in the knowledge graph and the reference word, which is the triplet similarity.
[0115] Based on the set triple similarity range, we can obtain words within the triple similarity range in the knowledge graph, which are also known as extended words.
[0116] Step S3023: Determine extended words based on word vectors and knowledge graphs.
[0117] Specifically, the third implementation method based on reference words is to combine word vectors and knowledge graphs to jointly determine the extended words.
[0118] This step is an optional step parallel to steps S3021 and S3022. Those skilled in the art can choose to perform any step according to the actual situation.
[0119] In one embodiment of this disclosure, such as Figure 3e The diagram shows a flowchart of a method for determining extended words using both knowledge graphs and word vectors. The specific steps include:
[0120] Step C1: Determine the position of the corresponding word in the reference content.
[0121] Step C2: The word whose distance to the word position is within a set range and whose similarity to the word vector of the reference word is within a set vector similarity range is determined as the first entity word.
[0122] Specifically, the determination of the first entity word, that is, the content of steps C1 to C2, is the same as the corresponding content in the aforementioned step S3021, and will not be repeated here.
[0123] Step C3: Determine the triplet features of the reference word and determine the word vector corresponding to the triplet features.
[0124] Step C4: Select words whose similarity to the word vectors corresponding to the triple features in the pre-determined knowledge graph is within the set triple similarity range as the second entity words.
[0125] Specifically, the determination of the second entity word, that is, the content of steps C3 to C4, is the same as the corresponding content in the aforementioned step S3022, and will not be repeated here.
[0126] Step C5: Rank the first entity word and the second entity word by comparing their text similarity with the reference content.
[0127] Specifically, the first and second entity words are sequentially compared with the words in the reference content. The maximum value of the text similarity for each entity word (including the first and second entity words) (usually the text similarity with the reference word, but it could also be the text similarity with other words in the reference content) is taken as the text similarity between the entity word and the reference content. Then, the words are sorted based on their text similarity.
[0128] Text similarity can be calculated using existing semantic similarity calculation methods, such as Jacard similarity calculation and Hamming distance similarity calculation.
[0129] Step C6: Sort the results to obtain at least one entity word, and identify it as an expanded word.
[0130] Specifically, after ranking the text similarity of the first and second entity words, the entity words with the highest similarity (a predetermined number) are identified as expanded words. These expanded words may originate from word vector similarity or from triplet feature similarity in the knowledge graph, thus potentially yielding results with richer sources compared to the previous two expansion methods.
[0131] Step S3024: Determine the inventory content tagged with extended words as the extended content corresponding to the first type of extended content.
[0132] Specifically, inventory content refers to the content in the content library of the content platform. There are usually multiple inventory contents corresponding to the same extended keyword. For example, there are multiple videos with a certain anchor's name as the tag.
[0133] Extended content can be of the same type as the reference content (e.g., both are articles) or it can be of a different type (e.g., articles and videos). There are no restrictions on the type of extended content, which can expand the diversity of extended content and meet users' needs for different types of content.
[0134] Step S303: Use the tags and / or entity features corresponding to the reference content as the topic words corresponding to the reference content.
[0135] Specifically, determining the expanded content through a topic tree is the second type of expansion method. In this second type of expansion method, the first step is to determine the topic keywords corresponding to the reference content.
[0136] Topic keywords are usually determined based on the tags and entity features corresponding to the reference content (topic keywords can be tags or a combination of tags and entity features). For example, if the tag is a song title and the entity feature of the reference content (that is, the entity properties of the reference content itself) is an introductory video of the song, then its topic keywords can be an introduction to the song corresponding to the tag or the song itself.
[0137] Step S304: Based on the correspondence between topic words and a pre-determined topic tree, determine the expansion content corresponding to the second type of expansion.
[0138] The topic tree includes existing topics and the existing content corresponding to those topics.
[0139] Specifically, the topic tree usually contains several topic words (i.e., existing topics), and each existing topic usually corresponds to multiple existing contents (e.g., multiple news articles on a certain topic). Therefore, the existing topics corresponding to the second type of expansion can be determined based on the topic words corresponding to the reference content, and the content of the corresponding existing topics can be determined as the expansion content.
[0140] like Figure 3f The diagram shown is a flowchart of the method for determining the extended content corresponding to the second type of extension. The specific steps include the following:
[0141] Step S3041: Determine the position of the topic word in the topic tree based on the granularity of the topic word.
[0142] Specifically, the existing topics in the topic tree are pre-divided into different granularities. By determining the granularity corresponding to the topic words of the reference content, and the existing topic that is closest to the topic words in the same granularity, the position of the obtained existing topic can be determined as the position of the topic words, and the extended content can be determined based on the position.
[0143] In one embodiment of this disclosure, such as Figure 3g The diagram shown is a flowchart of the method for determining the location of topic words. The specific steps include:
[0144] Step D1: Determine the granularity of topic keywords.
[0145] Specifically, the granularity of topic terms can be determined based on pre-configured rules, such as the number of dimensions involved in the topic term (generally, the more dimensions, the finer the granularity; for example, "the first single from a singer's new album this year" involves three dimensions: singer, album, and single, so its granularity is relatively fine), the granularity corresponding to each dimension (the finer the granularity of each dimension, the finer the combined granularity; for example, "a high-definition recording of a certain performance version of a classic piece of music"—the three dimensions of piece of music, performance version, and recording are all very fine granular, so the topic term obtained by combining the granularity of the three dimensions will have very fine granularity), or it can be determined by combining the granularity of each dimension (if categorized, then summing the number of categories, weighted summation, product, or taking the maximum or minimum value, etc.).
[0146] For example, if a topic word has two dimensions, and each dimension has a granularity of one or two (the larger the number of granularity, the smaller the granularity level, and the finer the granularity, such as granularity one being the largest granularity and granularity five being the smallest granularity), then the granularity of the topic word can be two (i.e., taking the maximum or minimum value or summing the number of dimensions) or three (i.e., summing the granularity of different dimensions).
[0147] Step D2: Identify existing topics in the topic tree that have the same granularity as the topic words.
[0148] Specifically, once the granularity of the topic words is determined, existing topics with the same granularity as the topic words can be obtained from the total number of topics. Generally, the finer the granularity, the more existing topics there are. The largest granularity usually has a large number of existing topics (e.g., more than fifty). Therefore, there are usually many existing topics with the same granularity as the topic words.
[0149] Step D3: Determine the layer in which the existing topic is located as the target layer to which the topic word belongs.
[0150] Specifically, the topic tree achieves a tree-like structure through different layers. For example, in the same topic tree structure, the main topic or root topic is located in a higher layer, and its corresponding subtopics are located in a lower layer (a main topic and all its subtopics constitute a topic tree structure).
[0151] Existing topics of different granularities belong to different layers in the topic tree. A single layer may contain multiple existing topics of different granularities. For example, the smallest and smallest granularities of existing topics may belong to the same layer's sub-topics (e.g., "a singer's concert on a certain day" and "a singer's concert on a certain day in a certain location" belong to the same layer / position, but the latter has finer granularity). Therefore, based on the determined granularity of the existing topics, it is necessary to determine the layer where the reference content's corresponding topic keywords belong, i.e., the target layer.
[0152] Step D4: Based on the cluster analysis of existing topics and target hierarchies, determine the cluster to which the topic words belong.
[0153] Specifically, by stratifying topic terms and target terms and performing cluster analysis on existing topics, the cluster to which the topic terms belong can be determined. Any existing text clustering algorithm can be used for this clustering analysis; no restrictions are imposed here.
[0154] After cluster analysis, a topic word may belong to one cluster or multiple clusters (such as two or three clusters). The cluster to which each topic word belongs can be used to determine the location of the topic word in the subsequent analysis.
[0155] Step D5: Determine the position of the existing topic corresponding to the cluster in the topic tree as the position of the topic word.
[0156] Specifically, the existing topics in each cluster obtained by cluster analysis may correspond to multiple topic tree structures. The positions of the corresponding existing topics in all these topic tree structures can be considered as the positions of the topic words corresponding to the reference content. That is, the position of the same topic word may be in multiple topic tree structures, rather than corresponding one-to-one with the position in a certain topic tree structure.
[0157] The location of a topic keyword usually has one or more existing topics. After extracting the extended content, the topic keywords corresponding to the reference content can also be added to that location, and the reference content corresponding to the topic keywords can be added to the topic tree as existing content for that location.
[0158] Step S3042: Based on the location and the corresponding existing content, determine the extended content corresponding to the second type of extension.
[0159] Specifically, existing content that corresponds to the topic words in the reference content (the same position in the same topic tree structure) or is similar to the topic words (the same layer in the same topic tree structure and belongs to the same cluster as the topic words) can usually be used as the second type of extended content.
[0160] In one embodiment of this disclosure, such as Figure 3h The diagram shows a flowchart illustrating the method for determining expanded content based on the location and corresponding content of topic keywords. The specific steps include:
[0161] Step E1: If the existing topic at the current location contains subtopics, determine the existing content corresponding to the subtopics as extended content.
[0162] Specifically, if there are multiple existing topics where the topic words of the reference content are located, then all existing content corresponding to the subtopics (topic words) of each existing topic can be determined as extended content.
[0163] Step E2: Determine the adjacent locations as those whose distance to the current location is less than the set distance value.
[0164] Specifically, the distance value here can be the distance between different topics in the same stratum during cluster analysis, or it can be the distance based on a defined subordinate relationship (e.g., the distance is zero when they are in the same position, the distance between topic words in the same direction within the same sub-topic is one unit, the distance between topic words in different directions within the same sub-topic is two units, the distance between topic words in different sub-topics within the same main topic is three units, the distance between topic words in different main topics is four units, and so on, dividing the distance).
[0165] For example, adjacent positions can be selected from existing topics that are located in the same sub-topic as the reference content's topic word. In this case, these existing topics are related to the reference content's topic word content, and when their content is pushed to users, users are more willing to interact.
[0166] Step E3: Determine the existing content corresponding to the existing topics in adjacent positions as extended content.
[0167] Specifically, all existing content corresponding to existing topics in adjacent positions can be identified as extended content; when the reference content topic word corresponds to multiple positions, and there are also multiple adjacent positions, then the existing content of all existing topics corresponding to these adjacent positions can be used as extended content.
[0168] Steps E2 to E3 are optional steps parallel to step E1, and those skilled in the art can choose to perform any step as needed.
[0169] Step S305: Determine the associated objects corresponding to the extended content, and deliver the extended content to the associated objects to obtain the delivery efficiency of the extended content corresponding to each content extension method.
[0170] Step S306: Based on delivery efficiency and expanded content, determine the recommended content to be pushed.
[0171] Specifically, steps S305 to S306 and Figure 2 The corresponding steps in the illustrated embodiments are the same and will not be repeated here.
[0172] According to the recommended content determination method of this disclosure, after determining reference content, based on the tags corresponding to the reference content, a first type of expansion is determined using those tags as reference words to determine content corresponding to similar tags, and a second type of expansion is determined using those tags as topic words to determine existing content corresponding to existing topics. The expanded content obtained from both types of expansion is then delivered, and the recommended content is determined by combining the delivery efficiency. Therefore, by using different types of expansion and multiple specific expansion methods, it is possible to fully acquire expanded content that may satisfy user interests and preferences. By introducing topic-related expansion methods and determining expanded content based on hot topics, the user experience of the target users can be improved. Compared to the single expansion method in related methods, this method can more comprehensively and systematically reflect user interests, fully explore users' potential needs, and improve user satisfaction with recommended content.
[0173] Figure 4a A flowchart illustrating a method for determining recommended content provided in one embodiment of this disclosure. Figure 4a As shown, the method for determining recommended content provided in this embodiment includes the following steps:
[0174] Step S401: Based on the predetermined reference content and at least one content expansion method, determine the expansion content corresponding to each content expansion method.
[0175] Specifically, this step is related to Figure 2 The corresponding steps shown are the same, so they will not be repeated here.
[0176] Step S402: Based on the content expansion method corresponding to the expanded content, add the corresponding method tag for the expanded content.
[0177] Specifically, the content expansion methods here are mainly aimed at Figure 3a The embodiments shown focus on the first and second types of expansion, rather than on more specific expansion implementation methods (such as expansion based on word vectors or knowledge graphs). This is because the first type of expansion is mainly based on the similarity between the reference content and its tags, while the second type of expansion is more focused on topics related to time-sensitive reference content. Depending on whether users care about the timeliness of the pushed content, the delivery ratio of the two types of expanded content can be adjusted. The specific implementation methods within each type of expansion are relatively similar and have little impact on user experience, so they do not need to be distinguished in detail.
[0178] Step S403: Determine the proportion of extended content corresponding to each type of tag.
[0179] Specifically, after adding a corresponding expansion method tag to each piece of extended content, the number of extended content items corresponding to each expansion method tag can be determined. Then, based on the proportion of extended content items for each tag to the total number of all extended content items, the proportion of extended content items corresponding to each expansion method tag can be determined.
[0180] Step S404: Identify objects that interact with the reference content as associated objects.
[0181] Specifically, after determining the expanded content, it is necessary to determine the target audience for the expanded content to be delivered, i.e., the relevant targets. To ensure that the relevant targets are interested in the expanded content, they need to have interacted with the reference content (at least one reference content).
[0182] The requirements for associated objects are lower than those for the target users of the reference content. Therefore, associated objects are usually not within the scope of the target users, but they are still interested in the reference content. Therefore, sending extended content to associated objects will not affect the subsequent push of recommended content to the target users (because the target users are different from the associated objects), but it can reflect whether the extended content can meet the interests and preferences of the target users (because both associated objects and target users will be interested in the reference content).
[0183] Step S405: Divide the associated objects into a set number of flow units.
[0184] Specifically, in order to adjust the distribution ratio of extended content and optimize distribution efficiency, it is necessary to divide the associated objects into multiple groups, and distribute a certain amount of extended content to each group. Therefore, the combination of the associated objects and the corresponding traffic in each group constitutes a traffic unit.
[0185] Traffic units can usually be randomly divided, but the number of associated objects and traffic in each traffic unit are usually kept the same (e.g., each traffic unit has a fixed number of one thousand associated objects, and the traffic for each traffic unit is 50,000 extended content items, which is to say, the traffic is set).
[0186] Step S406: Based on the set traffic volume and the set extension content delivery ratio for each type of tag, deliver extension content to the associated objects in the traffic unit.
[0187] Specifically, during the campaign, extended content can be delivered to a single traffic unit at a time. The delivery method involves sending extended content to a related target when a request for content is received (e.g., a request generated when clicking on the "Today's Recommended Music" feature), with the corresponding proportion of extended content pushed to that target.
[0188] The initial distribution ratio of extended content to traffic units can be based on the proportion of extended content corresponding to each tag. Then, the distribution ratio can be improved based on the distribution efficiency collected after the distribution.
[0189] Step S407: Take the delivery efficiency corresponding to each method tag fed back by the traffic unit as the delivery efficiency.
[0190] Specifically, after delivering the extended content to the traffic unit with the set traffic volume, the delivery efficiency reported by the traffic unit can be statistically analyzed after a set time period (such as 24 hours).
[0191] like Figure 4b The diagram shown illustrates the flowchart for calculating delivery efficiency. Delivery efficiency is calculated as follows:
[0192] Step S4071: Obtain the exposure of the extended content corresponding to each type of tag from the traffic unit feedback, and the number of interactions between the associated object and the extended content.
[0193] Specifically, when traffic units deliver extended content to associated objects, there may be situations where the associated objects do not choose to receive or view it (e.g., when an associated object opens the application client provided by the content platform, the client pushes extended content by generating a "recommended playlist" for the associated object, but the associated object does not click on the "recommended playlist"). In this case, it is considered that the extended content has not been exposed. In such cases, it is meaningless to count whether the associated objects are willing to interact with the extended content, because the associated objects have not seen the extended content itself.
[0194] Therefore, when determining the efficiency of ad placement, it is necessary to collect the exposure of the extended content (i.e., the number of extended content views or traffic by the associated objects) and the number of interactions between the associated objects and the extended content (including interaction methods such as likes, plays, forwards, shares, and comments) to determine the efficiency of ad placement.
[0195] Step S4072: Determine the campaign efficiency based on the ratio of interaction count to exposure.
[0196] Specifically, the higher the ratio of interaction frequency to exposure, the higher the willingness of the associated target audience to interact with the extended content, and the more the extended content matches the interests and preferences of the associated target audience. Therefore, the delivery efficiency is also higher.
[0197] Step S408: The maximum value of the delivery efficiency is used to determine the delivery ratio of the expanded content.
[0198] Specifically, after receiving the delivery efficiency, the delivery ratio needs to be adjusted, and the delivery should be directed to another traffic unit. The delivery efficiency should be collected again. By continuously adjusting the delivery ratio, the optimal delivery efficiency (maximum delivery efficiency) can be obtained, and the corresponding delivery ratio can be used as the target delivery ratio.
[0199] like Figure 4c The flowchart shown illustrates the method for determining the maximum delivery efficiency. The maximum delivery efficiency is obtained as follows:
[0200] Step S4081: Based on the delivery efficiency of each method tag in the traffic unit feedback, adjust the delivery ratio of the extended content corresponding to the method tag.
[0201] Specifically, adjusting the ad delivery ratio involves increasing the proportion of tags with high ad delivery efficiency and decreasing the proportion of tags with low ad delivery efficiency. Each adjustment value can be determined based on pre-configured rules. For example, it can be determined by the difference in ad delivery efficiency between the two tags and the corresponding difference in ad delivery ratio. If the difference in ad delivery efficiency is large, the adjustment value should be increased; if the difference in ad delivery ratio is large, the adjustment value for the ad delivery ratio should also be increased, and vice versa. Additionally, the number of traffic units can be considered; the more traffic units, the smaller the adjustment value usually should be, allowing for multiple small adjustments to accurately approximate the optimal ad delivery efficiency.
[0202] For example, if the first type of expansion method has a deployment ratio of 60% and a deployment efficiency of 30%, and the second type of expansion method has a deployment ratio of 40% and a deployment efficiency of 60%, then the difference in deployment efficiency is -30% (assuming it is a large difference range) and the difference in deployment ratio is 20% (assuming it is also a large range). Then the adjustment value can be set to 10%, that is, every time the deployment ratio of the second type of expansion method increases by 10%, the deployment ratio of the first type of expansion method decreases by 10%.
[0203] Step S4082: Deliver the adjusted delivery ratio of extended content to traffic units that have not yet received extended content, and obtain the corresponding delivery efficiency.
[0204] Specifically, after obtaining the adjusted delivery ratio based on the adjustment value, the expanded content can be delivered to the next traffic unit based on the new delivery ratio.
[0205] The extended content to be delivered at this time can be exactly the same as the extended content of the previous traffic delivery, or it can be different extended content (because there is a large amount of extended content obtained through different extension methods, when delivering to the traffic unit, it may only extract a part of the extended content for delivery, and there is still an undelivered part. These undelivered parts can be pushed to the associated objects in the next delivery, or they can not be pushed).
[0206] Step S4083: Determine that the delivery efficiency of traffic units that have not delivered extended content is lower than the delivery efficiency before adjusting the delivery ratio.
[0207] Specifically, after adjusting the delivery ratio, the delivery efficiency usually increases compared to the previous delivery (because the adjustment increases the proportion of high-efficiency extended content). When the delivery efficiency starts to decline, it means that the related object does not need so much high-efficiency extended content (such as not needing so much extended content corresponding to the second type of extended content). In this case, continuing to increase the proportion of high-efficiency extended content will usually only further reduce the overall delivery efficiency. Therefore, it is only necessary to take the highest delivery efficiency corresponding to the traffic unit that has been delivered so far as the maximum delivery efficiency (at this time, the traffic unit with the highest delivery efficiency is usually a traffic unit that was delivered before this adjustment).
[0208] Step S4084: Determine the maximum value of the delivery efficiency before adjusting the delivery ratio.
[0209] Specifically, the delivery efficiency before adjusting the delivery ratio, that is, the delivery efficiency of the previous traffic unit.
[0210] Step S409: Determine the expanded content of the target delivery ratio as recommended content.
[0211] Specifically, the recommended content pushed to the target users is the extended content delivered according to the target delivery ratio.
[0212] After determining the target delivery ratio, the target delivery ratio and the determined extended content can be stored. This allows for the delivery of extended content corresponding to different extension methods according to the target delivery ratio when a content recommendation request is received from a target user, thus completing the content recommendation process. Since the extended content with the target delivery ratio receives the highest probability of interaction from associated users (i.e., the highest delivery efficiency), it usually best meets the user's interests and preferences.
[0213] According to the recommended content determination method of this disclosure, after obtaining extended content, associated objects are identified and divided into multiple traffic units. Extended content with a set traffic volume is then sequentially delivered. The proportion of different extended content is optimized based on feedback delivery efficiency, thereby obtaining the target delivery ratio corresponding to the extended content with the best delivery efficiency. Extended content is then delivered based on the target delivery ratio as recommended content pushed to target users. Thus, by delivering content in stages and continuously optimizing the composition of extended content, a combination of extended content with a high probability of interaction with target users is output more efficiently, thereby ensuring high conversion efficiency of the pushed recommended content and avoiding the problem of decreased traffic conversion efficiency caused by excessive interest exploration.
[0214] Exemplary media
[0215] After introducing the methods of exemplary embodiments of this disclosure, the following references are made. Figure 5The storage medium of the exemplary embodiments of this disclosure will be described.
[0216] refer to Figure 5 As shown, a program product 50 for implementing the above-described method according to an embodiment of the present disclosure is described. This product may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto.
[0217] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0218] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium.
[0219] Program code for performing the operations disclosed herein can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN).
[0220] Exemplary device
[0221] Having introduced the medium of exemplary embodiments of this disclosure, the following references are made to... Figure 6 The recommended content determination apparatus according to exemplary embodiments of this disclosure will be described. Its implementation principle and technical effects are similar to those of the corresponding methods described above, and will not be repeated here. Figure 6The recommended content determination device shown is used to achieve the above. Figures 2 to 4a The recommended content determination method in the illustrated method embodiment.
[0222] The recommended content determining device 600 provided in this disclosure includes:
[0223] Preparation module 610 is used to determine the extended content corresponding to each content extension method based on pre-determined reference content and at least one content extension method;
[0224] The delivery module 620 is used to determine the associated objects corresponding to the extended content and deliver the extended content to the associated objects in order to obtain the delivery efficiency of the extended content corresponding to each content extension method.
[0225] Module 630 is used to determine the recommended content to be pushed based on delivery efficiency and expanded content.
[0226] In an exemplary embodiment of this disclosure, the preparation module 610 is specifically configured to: if the content expansion method includes a first type of expansion based on reference words and a second type of expansion based on a topic tree, use the tags corresponding to the reference content as reference words corresponding to the reference content; determine the expansion content corresponding to the first type of expansion based on the expansion words corresponding to the reference words; use the tags and / or entity features corresponding to the reference content as topic words corresponding to the reference content; and determine the expansion content corresponding to the second type of expansion based on the correspondence between the topic words and a pre-determined topic tree.
[0227] In one exemplary embodiment of this disclosure, the preparation module 610 is specifically used to: determine extended words based on the word vectors corresponding to the reference words; or, determine extended words based on a pre-determined knowledge graph and reference words; or, determine extended words based on word vectors and knowledge graphs; and determine the inventory content tagged as extended words as the extended content corresponding to the first type of extended content.
[0228] In an exemplary embodiment of this disclosure, the preparation module 610 is specifically used to: determine the position of the corresponding word in the reference content; determine words whose distance to the word position is within a set range as candidate words; calculate the similarity between the word vector of the reference word and the word vector of the candidate words; and determine the candidate words whose similarity is within the set vector similarity range as extended words.
[0229] In an exemplary embodiment of this disclosure, the preparation module 610 is specifically used to: determine the triplet features of the word vector of the reference word, and determine the word vector corresponding to the triplet features, wherein the triplet features are used to represent the vector relationship between the subject, relation word and object of the reference word; and determine words in the pre-determined knowledge graph whose similarity to the word vectors corresponding to the triplet features is within a set triplet similarity range as extended words.
[0230] In an exemplary embodiment of this disclosure, the preparation module 610 is specifically configured to: determine the position of the corresponding word in the reference content; determine the words whose distance to the word position is within a set range and whose similarity to the word vector of the corresponding word is within a set vector similarity range as first entity words; determine the triplet features of the reference word and determine the word vector corresponding to the triplet features; determine the words whose similarity to the word vectors corresponding to the triplet features in the pre-determined knowledge graph is within a set triplet similarity range as second entity words; sort the first entity words and the second entity words by comparing their text similarity with the reference content; and determine at least one entity word obtained from the sorting as an extended word.
[0231] In one exemplary embodiment of this disclosure, the preparation module 610 is specifically used to: if the topic tree includes existing topics and existing content corresponding to existing topics, determine the position of the topic word in the topic tree based on the granularity of the topic word; and determine the extended content corresponding to the second type of extension based on the position and the corresponding existing content.
[0232] In an exemplary embodiment of this disclosure, the preparation module 610 is specifically used to: determine the granularity of the topic word; determine existing topics in the topic tree that have the same granularity as the topic word; determine the layer in which the existing topic is located as the target layer to which the topic word belongs; determine the cluster to which the topic word belongs based on the cluster analysis of the topic word and the existing topics in the target layer; and determine the position of the existing topic corresponding to the cluster in the topic tree as the position of the topic word.
[0233] In one exemplary embodiment of this disclosure, the preparation module 610 is specifically configured to: if the existing topic at the current location contains a subtopic, determine the existing content corresponding to the subtopic as extended content; determine the location whose distance to the current location is less than a set distance value as an adjacent location; and determine the existing content corresponding to the existing topic at the adjacent location as extended content.
[0234] In one exemplary embodiment of this disclosure, the preparation module 610 is further configured to: determine the extended content corresponding to each content extension method based on predetermined reference content and at least one content extension method; add a method tag corresponding to the extended content based on the content extension method corresponding to the extended content; and determine the proportion of the number of extended content corresponding to each method tag.
[0235] In one exemplary embodiment of this disclosure, the delivery module 620 is specifically configured to: identify objects that interact with the reference content as associated objects; divide the associated objects into a set number of traffic units; deliver extended content to the associated objects in the traffic units based on the set delivery traffic and the set delivery ratio of extended content corresponding to each method tag; and use the delivery efficiency corresponding to each method tag fed back by the traffic units as the delivery efficiency.
[0236] In one exemplary embodiment of this disclosure, the delivery module 620 is specifically used to: calculate the delivery efficiency by: obtaining the exposure volume of the extended content corresponding to each method tag and the number of interactions between the associated object and the extended content from the traffic unit feedback; and determining the delivery efficiency based on the ratio of the number of interactions to the exposure volume.
[0237] In one exemplary embodiment of this disclosure, the determining module 630 is specifically used to: determine the target delivery ratio of the extended content by the delivery ratio corresponding to the maximum value of the delivery efficiency; and determine the extended content of the target delivery ratio as recommended content.
[0238] In one exemplary embodiment of this disclosure, the determining module 630 is specifically configured to: obtain the maximum value of the delivery efficiency by: adjusting the delivery ratio of the extended content corresponding to each method tag based on the delivery efficiency of each method tag fed back by the traffic unit; delivering the extended content after adjusting the delivery ratio to the traffic units that have not delivered extended content, and obtaining the corresponding delivery efficiency; determining that the delivery efficiency corresponding to the traffic units that have not delivered extended content is lower than the delivery efficiency before adjusting the delivery ratio; and determining the delivery efficiency before adjusting the delivery ratio as the maximum value of the delivery efficiency.
[0239] Exemplary computing device
[0240] Having described the methods, media, and apparatus of exemplary embodiments of this disclosure, the following references... Figure 7 A computing device according to an exemplary embodiment of the present disclosure will be described.
[0241] Figure 7 The computing device 700 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0242] like Figure 7 As shown, the computing device 700 is presented in the form of a general-purpose computing device. The components of the computing device 700 may include, but are not limited to: at least one processing unit 701, at least one storage unit 702, and a bus 703 connecting different system components (including the processing unit 701 and the storage unit 702).
[0243] The 703 bus includes a data bus, a control bus, and an address bus.
[0244] Storage unit 702 may include readable media in the form of volatile memory, such as random access memory (RAM) 7021 and / or cache memory 7022, and may further include readable media in the form of non-volatile memory, such as read-only memory (ROM) 7023.
[0245] Storage unit 702 may also include a program / utility 7025 having a set (at least one) program module 7024, such program module 7024 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0246] The computing device 700 can also communicate with one or more external devices 704 (e.g., keyboard, pointing device, etc.). This communication can be performed via the input / output (I / O) interface 705. Furthermore, the computing device 700 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via a network adapter 707. Figure 7 As shown, network adapter 707 communicates with other modules of computing device 700 via bus 703. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with computing device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0247] It should be noted that although several units / modules or sub-units / modules of the supply chain strategy determination device and the object scoring model training device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0248] Furthermore, although the operations of the methods disclosed herein are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0249] While the spirit and principles of this disclosure have been described with reference to several specific embodiments, it should be understood that this disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for convenience of expression. This disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A method for determining recommended content, characterized in that, The method includes: Based on predetermined reference content and at least one content expansion method, determine the expansion content corresponding to each content expansion method; Based on the content expansion method corresponding to the aforementioned expanded content, add the corresponding method tag for the expanded content; The objects that interact with the reference content are identified as the associated objects corresponding to the extended content; The associated objects are divided into a set number of traffic units; Based on the set traffic volume and the set extension content delivery ratio corresponding to each type of tag, the extension content is delivered to the associated objects in the traffic unit; The delivery efficiency corresponding to each method tag fed back by the traffic unit is used as the delivery efficiency of the extended content corresponding to each content extension method; Based on the delivery efficiency and the expanded content, the recommended content for push notifications is determined.
2. The method for determining recommended content according to claim 1, characterized in that, The content expansion methods include the first type of expansion based on reference words and the second type of expansion based on topic trees. The step of determining the extended content corresponding to each content extension method based on predetermined reference content and at least one content extension method includes: The tags corresponding to the reference content are used as reference words corresponding to the reference content; Based on the extended words corresponding to the reference words, determine the extended content corresponding to the first type of extension; The tags and / or entity features corresponding to the reference content are used as the topic words corresponding to the reference content; Based on the correspondence between the topic words and the pre-determined topic tree, the extended content corresponding to the second type of extension is determined.
3. The method for determining recommended content according to claim 2, characterized in that, The step of determining the extended content corresponding to the first type of extended content based on the extended words corresponding to the reference words includes: Based on the word vectors corresponding to the reference words, determine the expanded words; Alternatively, based on a pre-determined knowledge graph and the reference terms, extended terms can be determined; Alternatively, expandable words can be determined based on the word vectors and the knowledge graph; The inventory content tagged with the extended term is determined as the extended content corresponding to the first type of extension.
4. The method for determining recommended content according to claim 3, characterized in that, The process of determining expanded words based on the word vectors corresponding to the reference words includes: Determine the position of the corresponding word in the reference content; Words whose distance to the word's location is within a set range are identified as candidate words; Calculate the similarity between the word vectors of the reference word and the word vectors of the candidate words; Candidate words whose similarity falls within the set vector similarity range are selected as expanded words.
5. The method for determining recommended content according to claim 3, characterized in that, The process of determining extended terms based on a pre-determined knowledge graph and the reference terms includes: The triplet features of the reference word are determined, and the word vectors corresponding to the triplet features are determined. The triplet features are used to represent the vector relationships between the subject, relational words, and object of the reference word. Words whose word vectors corresponding to the triplet features in a pre-determined knowledge graph have a similarity within a set triplet similarity range are identified as extended words.
6. The method for determining recommended content according to claim 3, characterized in that, The process of determining extended words based on the word vectors and the knowledge graph includes: Determine the position of the corresponding word in the reference content; The word whose distance to the word position is within a set range and whose corresponding word vector is similar to the word vector of the reference word within a set vector similarity range is determined as the first entity word; Determine the triplet features of the reference word, and determine the word vector corresponding to the triplet features; Words whose similarity to word vectors corresponding to the triplet features in a pre-determined knowledge graph falls within a set triplet similarity range are designated as second entity words. The first entity word and the second entity word are ranked according to their text similarity with the reference content. The sorting process yields at least one entity word, which is then identified as an expanded word.
7. The method for determining recommended content according to claim 2, characterized in that, The topic tree includes existing topics and the existing content corresponding to those topics. The step of determining the extended content corresponding to the second type of extended content based on the correspondence between the topic words and a pre-determined topic tree includes: Based on the granularity of the topic words, determine the position of the topic words in the topic tree; Based on the location and the corresponding existing content, determine the extended content corresponding to the second type of extension.
8. The method for determining recommended content according to claim 7, characterized in that, Determining the position of a topic word in the topic tree based on its granularity includes: Determine the granularity of the topic terms; Identify existing topics in the topic tree that have the same granularity as the topic word; The layer in which the existing topic is located is determined as the target layer to which the topic term belongs; Based on the cluster analysis of topic words and existing topics in the target hierarchy, the cluster to which the topic words belong is determined; The position of the existing topic corresponding to the cluster in the topic tree is determined as the position of the topic word.
9. The method for determining recommended content according to claim 7, characterized in that, The step of determining the extended content corresponding to the second type of extension based on the location of the topic word and its corresponding existing content includes: If the existing topic at the location contains subtopics, the existing content corresponding to the subtopics is determined as the extended content; Positions whose distance to the given location is less than a set distance value are identified as adjacent positions; The existing content corresponding to the existing topics in adjacent positions is determined as the extended content.
10. The method for determining recommended content according to any one of claims 1 to 9, characterized in that, The step involves adding a method tag corresponding to the extended content based on the content extension method corresponding to the extended content; subsequently, it also includes: Determine the proportion of extended content corresponding to each type of tag.
11. The method for determining recommended content according to claim 1, characterized in that, The delivery efficiency is calculated as follows: Obtain the exposure volume of the extended content corresponding to each method tag fed back by the traffic unit and the number of interactions between the associated object and the extended content; The delivery efficiency is determined based on the ratio of the number of interactions to the number of exposures.
12. The method for determining recommended content according to claim 1, characterized in that, The step of determining the recommended content for push notifications based on the delivery efficiency and the expanded content includes: The maximum value of the delivery efficiency is used as the delivery ratio to determine the target delivery ratio for the expanded content. The expanded content of the target delivery ratio is determined as the recommended content.
13. The method for determining recommended content according to claim 12, characterized in that, The maximum delivery efficiency is obtained in the following way: Based on the delivery efficiency of each method tag as reported by the traffic unit, adjust the delivery ratio of the extended content corresponding to the method tag; Deliver the adjusted delivery ratio of the extended content to traffic units that have not previously received the extended content, and obtain the corresponding delivery efficiency. It was determined that the delivery efficiency of the traffic unit that had not delivered the extended content was lower than the delivery efficiency before the delivery ratio was adjusted. The delivery efficiency before adjusting the delivery ratio is determined as the maximum value of the delivery efficiency.
14. A computer-readable storage medium comprising: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the recommended content determination method as described in any one of claims 1 to 13.
15. A device for determining recommended content, characterized in that, The device includes: The preparation module is used to determine the extended content corresponding to each content extension method based on the pre-determined reference content and at least one content extension method. The delivery module is used to determine the associated object corresponding to the extended content and deliver the extended content to the associated object to obtain the delivery efficiency of the extended content corresponding to each content extension method. The determination module is used to determine the recommended content to be pushed based on the delivery efficiency and the expanded content; The preparation module is further configured to: determine the extended content corresponding to each content extension method based on the predetermined reference content and at least one content extension method, and then add the method tag corresponding to the extended content based on the content extension method corresponding to the extended content. The delivery module is specifically used for: The objects that interact with the reference content are identified as the associated objects; The associated objects are divided into a set number of traffic units; Based on the set traffic volume and the set extension content delivery ratio corresponding to each type of tag, the extension content is delivered to the associated objects in the traffic unit; The delivery efficiency corresponding to each method tag fed back by the traffic unit is taken as the delivery efficiency.
16. The recommended content determination device according to claim 15, characterized in that, The preparation module is specifically used for: If the content expansion method includes the first type of expansion based on reference words and the second type of expansion based on topic trees... The tags corresponding to the reference content are used as reference words corresponding to the reference content; Based on the extended words corresponding to the reference words, determine the extended content corresponding to the first type of extension; The tags and / or entity features corresponding to the reference content are used as the topic words corresponding to the reference content; Based on the correspondence between the topic words and the pre-determined topic tree, the extended content corresponding to the second type of extension is determined.
17. The recommended content determination device according to claim 16, characterized in that, The preparation module is specifically used for: Based on the word vectors corresponding to the reference words, determine the expanded words; Alternatively, based on a pre-determined knowledge graph and the reference terms, extended terms can be determined; Alternatively, expandable words can be determined based on the word vectors and the knowledge graph; The inventory content tagged with the extended term is determined as the extended content corresponding to the first type of extension.
18. The recommended content determination device according to claim 17, characterized in that, The preparation module is specifically used for: Determine the position of the corresponding word in the reference content; Words whose distance to the word's location is within a set range are identified as candidate words; Calculate the similarity between the word vectors of the reference word and the word vectors of the candidate words; Candidate words whose similarity falls within the set vector similarity range are selected as expanded words.
19. The recommended content determination device according to claim 17, characterized in that, The preparation module is specifically used for: The triplet features of the word vector of the reference word are determined, and the word vectors corresponding to the triplet features are determined. The triplet features are used to represent the vector relationship between the subject, relative words, and object of the reference word. Words whose word vectors corresponding to the triplet features in a pre-determined knowledge graph have a similarity within a set triplet similarity range are identified as extended words.
20. The recommended content determination device according to claim 17, characterized in that, The preparation module is specifically used for: Determine the position of the corresponding word in the reference content; The word whose distance to the word position is within a set range and whose corresponding word vector is similar to the word vector of the reference word within a set vector similarity range is determined as the first entity word; Determine the triplet features of the reference word, and determine the word vector corresponding to the triplet features; Words whose similarity to word vectors corresponding to the triplet features in a pre-determined knowledge graph falls within a set triplet similarity range are designated as second entity words. The first entity word and the second entity word are ranked according to their text similarity with the reference content. The sorting process yields at least one entity word, which is then identified as an expanded word.
21. The recommended content determination device according to claim 16, characterized in that, The preparation module is specifically used for: If the topic tree includes existing topics and the existing content corresponding to those existing topics. Based on the granularity of the topic words, determine the position of the topic words in the topic tree; Based on the location and the corresponding existing content, determine the extended content corresponding to the second type of extension.
22. The recommended content determination device according to claim 21, characterized in that, The preparation module is specifically used for: Determine the granularity of the topic terms; Identify existing topics in the topic tree that have the same granularity as the topic word; The layer in which the existing topic is located is determined as the target layer to which the topic term belongs; Based on the cluster analysis of topic words and existing topics in the target hierarchy, the cluster to which the topic words belong is determined; The position of the existing topic corresponding to the cluster in the topic tree is determined as the position of the topic word.
23. The recommended content determination device according to claim 21, characterized in that, The preparation module is specifically used for: If the existing topic at the location contains subtopics, the existing content corresponding to the subtopics is determined as the extended content; Positions whose distance to the given location is less than a set distance value are identified as adjacent positions; The existing content corresponding to the existing topics in adjacent positions is determined as the extended content.
24. The recommender content determination apparatus according to any one of claims 15 to 23, characterized in that, The preparation module is also used for: Based on the content expansion method corresponding to the aforementioned expanded content, add the corresponding method tag for the expanded content; Determine the proportion of extended content corresponding to each type of tag.
25. The recommended content determination device according to claim 15, characterized in that, The delivery module is specifically used for: The delivery efficiency is calculated as follows: Obtain the exposure volume of the extended content corresponding to each method tag fed back by the traffic unit and the number of interactions between the associated object and the extended content; The delivery efficiency is determined based on the ratio of the number of interactions to the number of exposures.
26. The recommended content determination device according to claim 15, characterized in that, The determining module is specifically used for: The maximum value of the delivery efficiency is used as the delivery ratio to determine the target delivery ratio for the expanded content. The expanded content of the target delivery ratio is determined as the recommended content.
27. The recommended content determination device according to claim 26, characterized in that, The determining module is specifically used for: The maximum delivery efficiency can be obtained in the following way: Based on the delivery efficiency of each method tag as reported by the traffic unit, adjust the delivery ratio of the extended content corresponding to the method tag; Deliver the adjusted delivery ratio of the extended content to traffic units that have not previously received the extended content, and obtain the corresponding delivery efficiency. It was determined that the delivery efficiency of the traffic unit that had not delivered the extended content was lower than the delivery efficiency before the delivery ratio was adjusted. The delivery efficiency before adjusting the delivery ratio is determined as the maximum value of the delivery efficiency.
28. A computing device, comprising: At least one processor; and memory that is communicatively connected to at least one processor; The memory stores instructions executable by at least one processor, which, when executed by at least one processor, cause the computing device to perform the recommended content determination method as described in any one of claims 1 to 13.
Citation Information
Patent Citations
Topic recommendation method and system based on user tags
CN111914079A
Content resource recommendation method and device, electronic equipment and storage medium
CN112528146A
Content recommendation method and device, equipment and storage medium
CN114428846A
Article recommendation method and device based on new energy cloud and user behaviors
CN116070024A