Label Processing Method, Device, Electronic Device and Computer Readable Storage Medium
By integrating user behavior information into the construction of the tag map and using the tag jump relationship to build the tag map, the problem of inaccurate tag correlation characterization in the existing technology is solved, and more accurate tag maps and richer content recommendations are achieved.
Patent Information
- Application Number
- CN202210674793.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-14
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-06-14
AI Technical Summary
The prior art is difficult to accurately characterize the correlation between labels, resulting in inaccurate construction of label maps, affecting the diversity of content recommendations and user experience.
Through the tag jump relationship of multiple contents processed continuously by the user, a tag map is built and user behavior information is integrated to improve the accuracy of the map information.
Fully explore the correlation between tags and improve the accuracy of tag maps, thereby improving the diversity and user experience of content recommendations.
Smart Images

Figure CN114970548B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing, and in particular, to a method, apparatus, electronic device, and computer-readable storage medium for label processing. Background Art
[0002] In the Internet era, a large amount of content is produced every day, such as news articles, advertising tweets, knowledge content, etc. In a piece of content describing a core thing, there are often tag words strongly related to the thing. Using artificial intelligence technology, tag words describing the core of the content can be accurately extracted from single-modal (only text or only image) or multi-modal (text and image) content. Such technologies are usually directly used for content understanding and can then be applied to downstream tasks such as content recommendation. However, the tag-related information contained in a single piece of content is limited, and the relevance of the recommended content recalled based on a certain piece of content is likely to be too high, resulting in insufficient diversity. Specifically, there are some tag words that have a certain relevance to the content but are not included in the tagging results of the content and thus cannot be used for relevant recommendations.
[0003] Therefore, on the basis that the single / multi-modal tagging technology for a single piece of content has been relatively mature, it is necessary to use the tagging results of a large amount of content to establish a cross-article dimension tag semantic relevance network, that is, a tag graph, to better serve downstream applications. The tag graph uses tags as nodes in the graph and represents the relevance between tags in the edges between nodes, thereby realizing the semantic relevance analysis of different tags. In related technologies, the relevance between tags is often determined by the co-occurrence frequency of tags. However, since the co-occurrence frequency of different tags has a great relationship with the occurrence frequency of each tag itself, the tag graph established in related technologies is difficult to accurately represent the relevance between tags. Summary of the Invention
[0004] Embodiments of the present application provide a method, apparatus, electronic device, and computer-readable storage medium for label processing to solve the problems existing in related technologies. The technical solutions are as follows:
[0005] In a first aspect, an embodiment of the present application provides a method for label processing, including:
[0006] Determine a first user behavior sequence; wherein the first user behavior sequence includes multiple contents continuously processed by a user;
[0007] Based on multiple tags corresponding to the multiple contents, determine multiple tag jump relationships;
[0008] Based on the multiple tag jump relationships, determine first graph information.
[0009] In a second aspect, an embodiment of the present application provides a label processing apparatus, including:
[0010] A sequence determination module for determining a first user behavior sequence, where the first user behavior sequence includes a plurality of contents consecutively processed by the user.
[0011] A jump determination module for determining a plurality of label jump relationships based on a plurality of labels corresponding to the plurality of contents.
[0012] A first graph determination module for determining first graph information based on the plurality of label jump relationships.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the method provided in any embodiment of the present application is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method provided in any embodiment of the present application is implemented.
[0015] The technical solution of the embodiment of the present application, for a plurality of contents consecutively processed by a user, mines the jump relationship between labels based on the plurality of labels corresponding to them, and obtains graph information based on the label jump relationship. In this way, user behavior information is incorporated into the construction process of the label graph, which can fully mine the correlation between labels, thereby improving the accuracy of the graph information.
[0016] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the above-described illustrative aspects, embodiments, and features, further aspects, embodiments, and features of the present application will be readily apparent by reference to the drawings and the following detailed description. Description of the Drawings
[0017] In the drawings, unless otherwise specified, the same reference numerals throughout the drawings denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments disclosed in the present application and should not be regarded as limiting the scope of the present application.
[0018] Figure 1 is a schematic diagram of an exemplary application scenario of an embodiment of the present application;
[0019] Figure 2 is a flowchart of a label processing method according to an embodiment of the present application;
[0020] Figure 3 is a flowchart of a label processing method according to another embodiment of the present application;
[0021] Figure 4 It is a flowchart of a tag processing method according to another embodiment of the present application;
[0022] Figure 5 It is a schematic diagram of an application example of the tag processing method provided by the embodiment of the present application;
[0023] Figure 6 It is a structural block diagram of a tag processing device according to an embodiment of the present application;
[0024] Figure 7 It is a structural block diagram of an electronic device for implementing the tag processing method of the present application. Detailed implementation manners
[0025] In the following, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present application. Therefore, the drawings and descriptions are considered to be exemplary in nature rather than restrictive.
[0026] To facilitate the understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.
[0027] The content tagging technology can be applied to content recommendation. For example, after a user reads an article, news products can recommend other articles sharing tags with this article to the user according to the tagging results of the article. Since the scope involved in the tag extraction results of a single piece of content is too small, there are some tag words that have a certain relevance to this content but are not included in the tagging results of this content, so they cannot be used for relevant recommendations, resulting in insufficient diversity of recalled content. This not only easily causes user reading fatigue but also wastes relevant content in the content library that may arouse reading interest.
[0028] For the tagging results of a large number of different contents, constructing a tag graph and representing the tag relevance in the edges between nodes in the graph can realize the semantic relevance analysis of different tags. On the one hand, it can establish an updatable semantic network for content understanding. Due to the natural characteristic of strong timeliness of media content, the construction of the above semantic network can obtain the latest associated tags of a tag word based on the latest corpus; on the other hand, mining the deep semantic relationships between cross-article tagging results can improve the richness of recommended content in downstream tasks such as relevant recommendations and arouse readers' reading interest.
[0029] In the related art, a tag graph construction scheme based on co-occurrence frequency is generally adopted. This scheme calculates the correlation between two tags according to the co-occurrence frequency of the two tags. For example, for tag a, the correlation between tag a and tag b can be characterized as:
[0030]
[0031] However, since some high-frequency tags generally appear in various types of content and there are many connected nodes in the tag graph, the correlation scores with other tags obtained using the above formula are all relatively low. On the contrary, some low-frequency tags only appear in a small number of manuscripts, and the number of connected nodes is limited. At this time, the correlation scores of such tags with other tags calculated using the above formula are all very high.
[0032] Therefore, it is difficult for the above scheme to accurately measure the association relationship between the above tags. That is, the tag graph construction method based on co-occurrence frequency fails to incorporate information such as the out-degree of the node itself into the correlation evaluation. When a content contains multiple tags, it is very difficult to perform a reasonable correlation ranking on the multiple edges connected by the multiple tags of the content in the graph.
[0033] In addition, on the basis of constructing the above tag graph, the information contained in the graph can be used to generate the representation information of each tag node in the graph. This representation information can exist in the form of a vector, which can not only quantitatively describe the semantic information of a tag node, but also be used as an input feature for relevant recommendations or various deep learning models. In the related art, the representation information of tags is generally generated based on the context of the text. Specifically, through unsupervised learning technology, the text corpus containing words is input into the language model, and the mathematical dense representation vector of the words is generated according to the context of the words. This scheme is limited by the input length and can only incorporate the context information within a certain length window into the learning scope of the model.
[0034] The technical solution of the embodiment of the present application mainly solves at least one of the above technical problems. Figure 1 An exemplary application scenario is shown. Such as Figure 1As shown, a tag processing device for implementing the tag processing method of the embodiments of the present application can receive content from a content production platform, perform tagging based on this content, and then process it according to the tags of each content to output the graph information of the tags. Here, the content production platform is, for example, a news and information platform, an information sharing platform, a social platform, etc. The graph information contains the correlation information between the tags, which can be used for semantic understanding of the tags themselves or the content containing the tags, and then can be applied to downstream application tasks such as content recommendation and deep learning prediction tasks. For example, based on the graph information, other content related to the content already read in the content production platform can be recommended to the user. Another example is that based on the graph information, other tags related to the tag can be extracted to perform relevant predictions on the entity corresponding to the tag, such as predicting the tourism popularity of a location entity and predicting the consumption ability of a person name entity.
[0035] Optionally, the tag processing device can also output a representation vector of the tag. This representation vector can quantitatively represent the semantic information of the tag, that is, quantitatively represent the semantic understanding result of the tag, so as to be applied to the above-mentioned downstream application tasks.
[0036] In order to more comprehensively understand the features and technical content of the embodiments of the present application, the implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are only for reference and illustration, and are not used to limit the embodiments of the present application.
[0037] Figure 2 The flowchart of a tag processing method according to an embodiment of the present application is shown. This method can be executed by the above-mentioned tag processing device, but is not limited thereto. The tag processing method can include:
[0038] S210. Determine a first user behavior sequence; wherein, the first user behavior sequence includes a plurality of contents continuously processed by the user;
[0039] S220. Determine a plurality of tag jump relationships based on the plurality of tags corresponding to the plurality of contents;
[0040] S230. Determine a first graph information based on the plurality of tag jump relationships.
[0041] In the embodiments of the present application, a user behavior sequence may refer to a content sequence composed of multiple contents continuously processed by the same user within a period of time. The first user behavior sequence may be determined based on the access records of the content database. In the embodiments of the present application, the user's processing of the content may include clicking (opening), that is, the multiple contents continuously processed by the user may include the multiple contents continuously clicked by the user. In some scenarios, the user's processing of the content may also include browsing (reading), liking, collecting, etc. For example, only the user's liking and collecting behaviors may be concerned. The multiple contents continuously processed by the user may include at least one content liked by the user, and some other contents collected after liking, etc.
[0042] In the embodiments of the present application, a tag may refer to a word that characterizes the core semantics of the whole text in a content. Exemplarily, the tag may include an entity tag and a core word tag. Among them, the entity tag may refer to an entity word of the main description object in a content. The core word tag may refer to a non-entity word that can characterize the core of the content, such as a conceptual word.
[0043] Optionally, after determining the first user behavior sequence, tags may be determined (i.e., tagged) for multiple contents therein respectively. It is also possible to pre-determine the tags of each content in the content database, and after determining the first user behavior sequence, extract the tags of each content in the first user behavior sequence.
[0044] The following provides an exemplary way of tagging.
[0045] The extraction of entity tags can be implemented by using a language model. For example, a language model is established using deep learning technology to achieve the extraction and classification of specific segments in the text. In the training stage, the input corpus with annotated entity recognition results is input into the language model. The input corpus includes the input text and the entity annotation. Among them, the input text is ordinary text information; the entity annotation is the entity recognition result of the same length as the input text, that is, the entity annotation includes the entity recognition result corresponding to each character in the input text, used to inform the language model which segments of the text content are entity categories such as person names and place names. In the inference stage, the content for which the entity tag is to be extracted is used as the input to the language model, and the entity segments and their entity categories included in the content can be output.
[0046] Since the content involved in the core word tags is more extensive than that of entity tags, the process of word segmentation recall, non-core word filtering, and relevance ranking can generally be used to extract core words. To ensure accurate tag extraction, a tag library with a wide range of contents and sustainable updates is prepared in advance before extracting core word tags. The tag library can contain core word tags in various fields. This tag library serves as a whitelist to filter out the recall results that are not suitable as tags. In specific applications, a word segmentation tool is first used to segment the content to obtain the word sequence corresponding to the content. Single words (terms) in the word sequence and multiple words (i.e., phrases) covered by a time window of a preset length can be used as recall results. Words or phrases that are not in the whitelist of the tag library in the recall results are filtered. And for the multiple words / phrases included in the filtered recall results, based on features such as word frequency and whether they are in the title, the relevance between the word / phrase and the content is judged, so as to rank the multiple words / phrases based on this relevance to obtain the core word tags most relevant to the content theme.
[0047] In the embodiments of the present application, the multiple tags corresponding to the multiple contents in the first user behavior sequence include the tags of each content in the first user behavior sequence. Based on the tags of each content, multiple tag jump relationships can be determined. Here, the tag jump relationship can refer to the jump relationship between the tags of adjacent contents processed by the user, that is, the correlation between the tags extracted due to content jumps. Exemplarily, the multiple tag jump relationships can include the correspondence between the tag of the third content in the first user behavior sequence and the tag of the fourth content in the first user behavior sequence. For example, if a certain content in the first user behavior sequence corresponds to tag 1 and tag 2, and other contents correspond to tag 3 and tag 4, then the multiple tag jump relationships can include the correspondence between tag 1 and tag 3, the correspondence between tag 1 and tag 4, the correspondence between tag 2 and tag 3, the correspondence between tag 2 and tag 4, etc. It can be seen that even if two tags do not co-occur in the same content, according to the method of the embodiments of the present application, the correlation between these two tags can be captured based on the user behavior sequence.
[0048] In the embodiments of the present application, the first graph information can be determined based on multiple tag jump relationships. Among them, the first graph information includes node and edge information. The nodes in the first graph information can include the tags of each content in the first user behavior sequence. The edge information includes the correlation information between these tags. Specifically, when there is the above correspondence between two tags, these two tags can be connected by an edge, and the edge information is the number of times of connecting these two tags, that is, the number of times the above correspondence occurs between these two tags. In the embodiments of the present application, this number can be referred to as the jump number between two tags.
[0049] In an example, the first graph information in the embodiments of the present application can be represented by the following formula:
[0050] consume = {x 1 : {x 2 : c 12 , x 3 : c 13 ,...}, x 2 : {x 1 : c 21 , x 3 : c 23 ...},..., x n : {x 1 : c n1 , x 2 : c n2 ...}} Formula (1)
[0051] where consume is the mathematical representation of the atlas information, x i is the i-th label involved in the atlas information, and c mn is the number of jumps between x m and x n , or rather, the number of jumps from x m to x n . Therefore, the atlas information reflects the number of jumps between each label and other labels, and this number of jumps reflects the correlation between each label and other labels. That is to say, the atlas information can reflect the correlation between each label and other labels.
[0052] The above formula can be used as an example of the first atlas information. It can be understood that in some embodiments of the present application, other information can also be integrated to construct a label atlas. Therefore, the first atlas information can also be characterized based on other information.
[0053] Exemplarily, the first atlas information can be determined only based on the multiple label jump relationships for the first user behavior sequence as described above, or other corresponding relationships between labels can be determined based on other user behavior sequences or based on the co-occurrence frequency / number of labels, and then the first atlas information can be determined by combining the above multiple label jump relationships and other corresponding relationships.
[0054] In summary, in the method of the embodiments of the present application, for multiple contents continuously processed by the user, the jump relationships between labels are mined based on the multiple corresponding labels, and the atlas information is obtained based on the label jump relationships. In this way, the user behavior information is incorporated into the construction process of the label atlas, and the correlation between labels can be fully mined, thereby improving the accuracy of the atlas information.
[0055] In the embodiments of the present application, an application method of the above label processing method can also be provided. Exemplarily, the above label processing method can further include:
[0056] Determine the second content recommended to the user based on the label of the first content processed by the user and the first graph information.
[0057] Specifically, the content recommendation service can monitor the content processed by the user in real time, for example, read the label of the content currently clicked by the user in real time. The content recommendation service can determine at least one related label of the label based on the label of the content currently clicked by the user and the first graph information, and select the matching content in the content database based on the related label as the content recommended to the user.
[0058] Exemplarily, in the above step S210, determining the first user behavior sequence may include:
[0059] Select a second user behavior sequence from among multiple user behavior sequences associated with the first content set;
[0060] Determine at least one subsequence in the second user behavior sequence using a time sliding window of a preset length; wherein, the at least one subsequence includes the first user behavior sequence.
[0061] That is to say, the first user behavior sequence is one of the subsequences of the second user behavior sequence selected from among multiple user behavior sequences associated with the first content set.
[0062] Specifically, the first content set may include multiple contents. Exemplarily, the first content set may include the contents produced on the content production platform within a certain time range, such as within one year or within half a year. Correspondingly, the multiple user behavior sequences may also be the user behavior sequences within the corresponding time range.
[0063] Exemplarily, the method of selecting the second user behavior sequence from among multiple user behavior sequences may be random selection, or may be based on the activity of the user, for example, only select the user behavior sequences of active users.
[0064] Exemplarily, the time sliding window in the embodiments of the present application, which may also be referred to as a sliding time window, is, for example, a window containing multiple contents. And this window is translated forward based on a unit length (for example, one content), and at least one data range, that is, at least one subsequence, can be determined based on a user behavior sequence of a certain length. Taking the second user behavior sequence including content A to content F and the length of the time sliding window (i.e., the above preset length) being 4 contents, then based on the second user behavior sequence {A, B, C, D, E, F}, three subsequences can be determined, including {A, B, C, D}, {B, C, D, E}, and {C, D, E, F}.
[0065] Exemplarily, each of the at least one subsequence above can be used as the first user sequence, and corresponding multiple tag jump relationships can be obtained according to the above steps S210 and S220 respectively. In practical applications, the above step of determining the first user behavior sequence can be executed multiple times. For example, in multiple user behavior sequences associated with the first content set, the second user behavior sequence is selected multiple times. Each time the second user behavior sequence is selected, a corresponding subsequence is obtained using a time sliding window, and corresponding multiple tag jump relationships are obtained based on each subsequence.
[0066] Correspondingly, in the above step S230, determining the first graph information based on multiple tag jump relationships includes:
[0067] Adding the multiple tag jump relationships to the tag jump relationship set corresponding to the first content set;
[0068] Determining the first graph information corresponding to the first content set based on the tag jump relationship set.
[0069] That is to say, in the user behavior sequence associated with the first content set, subsequences are determined in a time sliding window manner, and multiple tag jump relationships are obtained based on each subsequence. These tag jump relationships can be summarized into the tag jump relationship set corresponding to the first content set, and graph information is constructed based on this set.
[0070] Since the subsequences are determined in a time sliding window manner, based on the above embodiments, a large number of subsequences can be obtained, thus obtaining a set containing rich tag jump relationships, providing a sufficient data source for determining graph information, and being conducive to improving the accuracy of graph information.
[0071] Optionally, in an exemplary embodiment, the determining the first graph information corresponding to the first content set based on the tag jump relationship set may include:
[0072] Obtaining second graph information based on the tag jump relationship set; wherein, the second graph information includes Y groups of jump times corresponding to Y tags; wherein, the i-th group of jump times in the Y groups of jump times includes the jump times between the i-th tag among the Y tags and each other tag, Y is an integer greater than or equal to 1, and i is a positive integer less than or equal to Y;
[0073] Processing the second graph information based on a preset number threshold to obtain the first graph information.
[0074] Exemplarily, the second graph information may be the graph information shown in the foregoing formula (1), and the correlation between tags is mainly represented by the number of tag jumps. On this basis, the second graph information may be processed based on a preset number threshold to obtain optimized graph information, that is, the first graph information. For example, the number of tag jumps c less than the preset number threshold in formula (1) may be mn reset to zero to weaken the influence of the contingency of two unrelated tags in the same user behavior sequence on subsequent processing work.
[0075] The above embodiments give some examples of determining graph information based on user behavior sequences. In practical applications, a tag graph may also be constructed by fusing other information with the first graph information. Specifically, Figure 3 FIG. shows a flowchart of a tag processing method according to another embodiment of the present application. This method may be executed by the above-mentioned tag processing device, but is not limited thereto. The tag processing method may include:
[0076] S310. Determine a first user behavior sequence; wherein, the first user behavior sequence includes a plurality of contents continuously processed by the user;
[0077] S320. Determine a plurality of tag jump relationships based on a plurality of tags corresponding to the plurality of contents;
[0078] S330. Determine the first graph information based on the plurality of tag jump relationships;
[0079] S340. Determine the co-occurrence times between every two of the plurality of tags corresponding to the first content set;
[0080] S350. Determine the third graph information based on the co-occurrence times; wherein, the third graph information includes X groups of co-occurrence times corresponding to X tags; wherein, the jth group of co-occurrence times in the X groups of co-occurrence times includes the co-occurrence times of the jth tag among the X tags with each of the other tags, X is an integer greater than or equal to 1, and j is a positive integer less than or equal to X;
[0081] S360. Obtain the fourth graph information corresponding to the first content set based on the first graph information and the third graph information.
[0082] It should be noted that the sequence order of the above steps S310 to S330 (i.e., the steps of determining the first graph information) and steps S330 to S340 (i.e., the steps of determining the third graph information) is not limited. For example, the first graph information may be determined first and then the third graph information, or the third graph information may be determined first and then the first graph information, or the first graph information and the third graph information may be determined in parallel.
[0083] By way of example and not limitation, in the embodiments of the present application, the number of tags X corresponding to the third map information may be equal to the number of tags Y corresponding to the first map information and the second map information, so as to perform fusion between map information based on the same data range and improve information accuracy.
[0084] In the embodiments of the present application, co-occurrence means that two tags appear in the same content at the same time. For example, if the tags of content A include tag 1 and tag 2, and the tags of content B include tag 2 and tag 3, then the co-occurrence count of tag 2 can be incremented by one.
[0085] In practical applications, the co-occurrence count between multiple tags pairwise can be statistically obtained by traversing the tags of each content in the above-mentioned first content set. For example, assume that the tags of a content can be represented as {x 1 ,x 2 ,...,x n}. Traverse this list. When traversing to the i-th element x i ,use x i as the starting node, and use the tags x j in the remaining n - 1 elements as the ending nodes (i ≠ j) to establish connections. In this way, after traversing n elements, the construction of connections between tag nodes in a single content is completed. Perform the above processing for each content in the first content set, then the construction of connections between the corresponding tags in the first content set can be completed. The count of the connections between the tags is the co-occurrence count between the tags.
[0086] In one example, the third map information can be represented by the following formula:
[0087] map = {x 1 : {x 2 : t 12 ,x 3 : t 13 ,...}, x 2 : {x 1 : t 21 ,x 3 : t 23 ...},..., x n : {x 1 : t n1 ,x 2 : t n2 ,...}} Formula (2)
[0088] Among them, map is the mathematical representation of the third map information, x i is the i-th tag involved in this map information, and t ij is the co-occurrence count between x i and x jThe co-occurrence times. Therefore, the graph spectrum information reflects the co-occurrence times between each pair of tags, and this co-occurrence times can also reflect the correlation between each tag and other tags. That is to say, the third graph spectrum information can also reflect the correlation between each tag and other tags. Exemplarily, the third graph spectrum information can also be processed based on a preset number threshold to reduce the impact of the contingency that two unrelated words co-occur in the same content on subsequent processing work.
[0089] In the embodiments of the present application, the fourth graph spectrum information can be obtained based on the first graph spectrum information and the third graph spectrum information. Since the third graph spectrum information obtained based on the co-occurrence times is fused, the fourth graph spectrum information can more accurately reflect the correlation between tags, providing a more accurate data basis for downstream application tasks.
[0090] Exemplarily, in the above method, there are various ways to fuse the first graph spectrum information and the third graph spectrum information. For example, the jump times in the first graph spectrum information and the co-occurrence times in the third graph spectrum information are weighted and summed based on a preset weight, or the third graph spectrum information is corrected based on the first graph spectrum information, etc.
[0091] The following provides an exemplary way to determine the fourth graph spectrum information based on the first graph spectrum information and the third graph spectrum information. Specifically, obtaining the fourth graph spectrum information corresponding to the first content set based on the first graph spectrum information and the third graph spectrum information includes:
[0092] Determining the association probability between each pair of multiple tags corresponding to the first content set based on the first graph spectrum information, the third graph spectrum information, and a preset random search probability;
[0093] Determining the weight information of each tag among the multiple tags based on the association probability between each pair of the multiple tags;
[0094] Obtaining the fourth graph spectrum information based on the weight information of each tag.
[0095] Exemplarily, the association probability between two tags can represent the probability of connecting edges with one tag as the starting node and the other tag as the ending node in the tag graph. This probability can be determined based on the probability of connecting edges between the two tags in the first graph spectrum information and the probability of connecting edges between the two tags in the third graph spectrum information.
[0096] The above random search probability can be a preset probability value. The random search probability can be used to represent the probability of connecting edges based on the tag jump times. Correspondingly, based on the random search probability, the probability of connecting edges based on the co-occurrence times can also be determined. For example, if the random search probability is rsp, then there is a probability of rsp to connect edges according to the tag jump times, and a probability of (1 - rsp) to connect edges according to the co-occurrence times between tags.
[0097] Exemplarily, if there are a total of size(comsume(x i )) other tags having a label jump relationship with the label x i ), then without considering the co-occurrence times, for the label x i , the probability that the label connected to the label x i is any one of the size(comsume(x i )) tags. Exemplarily, if the number of edges with the label x i as the starting node of the edge and the label x j as the ending node of the edge is t ij , then without considering the user behavior sequence, the probability of using the label x i as the starting node of the edge and the label x j as the ending node of the edge can be expressed as That is, the proportion of the co-occurrence times of x i and x j in the total co-occurrence times of x i . Based on the above example, when the label x j is one of the size(comsume(x i )) tags, the association probability between x i and x j is When the label x j is not one of the size(comsume(x i )) tags, the association probability between x i and x j is
[0098] Exemplarily, the association probabilities between the above multiple tags pairwise can be represented by a matrix T'. The element in the i-th row and j-th column of this matrix T' is the above t' ij , that is, the association probability between x i and x j , and can also be understood as the probability of transitioning from x i to x j .
[0099] Exemplarily, the weight information of each tag among the multiple tags can be determined based on the above matrix. This weight information can represent the importance degree of this tag in the first content set, for example, it can represent the popularity of this tag. In practical applications, this weight information can also be used as a feature information of the tag, and downstream application tasks such as content recommendation and / or deep learning prediction can be performed according to this feature.
[0100] An exemplary method for determining the weight information of each tag based on the association probability between pairs of multiple tags is introduced below. Specifically, determining the weight information of each tag among multiple tags based on the association probability between pairs of multiple tags may include:
[0101] Performing multiple iterations based on the initial weight vector and the association probability between pairs of multiple tags to obtain the weight information of each tag;
[0102] Among them, the t-th iteration in multiple iterations includes:
[0103] Based on the t-th weight vector and the probability matrix, obtaining the (t + 1)-th weight vector; where the probability matrix is used to represent the association probability between pairs of multiple tags;
[0104] When the (t + 1)-th weight vector meets the preset conditions, obtaining the weight information of each tag based on the (t + 1)-th weight vector, where t is an integer greater than or equal to 1.
[0105] Among them, the initial weight vector is the first weight vector. In the embodiments of the present application, the number of elements in the weight vector may be the number of tags corresponding to the above first content set. For example, if the above multiple tags are n tags, then the weight vector has n elements, and each element in the weight vector corresponds to each tag one by one, representing the weight information of the corresponding tag. The initial weight vector can be preset. For example, the initial weight vector can be That is, the initial weight information of each tag is
[0106] According to the above method, based on P 0 Perform multiple iterations. In the first iteration, based on the initial weight vector P 0 And the probability matrix, such as the above matrix T′, obtain the second weight vector P 1 , and determine whether P 1 Meets the preset conditions. If not, perform the second iteration; in the second iteration, based on the second weight vector P 1 And the probability matrix, obtain the third weight vector P 2 , and determine whether P 2 Meets the preset conditions. If not, perform the third iteration. And so on. When the h-th iteration is executed, obtain the (h + 1)-th weight vector P h . If P h Meets the preset conditions, then obtain the weight information of each tag based on P h . That is to say, the elements in P h Are the weight information of the corresponding tags.
[0107] Exemplarily, the above preset conditions may be related to the weight vector obtained in the previous iteration. For example, the preset condition corresponding to the (t + 1)-th weight vector is: the t-th weight vector P t-1 and the (t + 1)-th weight vector P t The distance between them is less than or equal to a preset distance threshold. Specifically, this distance can be characterized based on the Euclidean norm, and the preset condition can be expressed as ‖P t - P t-1 ‖ ≤ v, where ‖·‖ represents the Euclidean norm and v is the preset distance threshold.
[0108] After obtaining the above weight information, in the embodiments of the present application, based on the weight information of each label, the fourth graph information is obtained, which may be to correct the first graph information and / or the third graph information based on the weight information of each label to obtain the fourth graph information. Exemplarily, obtaining the fourth graph information based on the weight information of each label includes:
[0109] Processing the third graph information based on the weight information of each label to obtain the fourth graph information.
[0110] That is to say, in the embodiments of the present application, taking the third graph information obtained based on the co-occurrence times as the main information, the user behavior data and the co-occurrence data are combined using the random search probability to obtain the weight information of each label, so as to correct the third graph information. Since the co-occurrence relationship is the most direct manifestation of the correlation between labels, the fourth graph information obtained based on the above method can more accurately reflect the correlation between labels.
[0111] Exemplarily, the above way of processing the third graph information may be to multiply the weight information of the j-th label by each co-occurrence time in the j-th group of co-occurrence times in the third graph information.
[0112] The above embodiments give some examples of determining graph information based on user behavior sequences and co-occurrence relationships. Using these examples can improve the accuracy of graph information. Optionally, the embodiments of the present application also provide some examples that can improve the accuracy of the representation information of each label using graph information. Figure 4 The flowchart of a label processing method according to another embodiment is shown. This method can be executed by the above label processing device, but is not limited thereto. The label processing method may include:
[0113] S410. Determine the fourth graph information corresponding to the first content set;
[0114] S420. Determine a plurality of label sequences based on the weight information of each label and the fourth graph information;
[0115] S430. Process the initial representation vectors of each label among multiple labels based on multiple label sequences and a word vector generation model to obtain the target representation vectors of each label.
[0116] Among them, the step S410 of determining the fourth graph information corresponding to the first content set can be implemented with reference to the above embodiments and will not be elaborated here.
[0117] Exemplarily, each label sequence among the multiple label sequences may include a certain number of labels, and this number can be preset. In practical applications, labels in each label sequence can be selected based on the weight information of each label and the fourth graph information.
[0118] In one implementation manner, determining multiple label sequences based on the second weight information of each label and the fourth graph information includes:
[0119] Select initial labels from multiple labels based on the weight information of each label;
[0120] Select multiple wandering labels from multiple labels based on the initial labels and the fourth graph information;
[0121] Obtain the first label sequence among the multiple label sequences based on the initial labels and the multiple wandering labels.
[0122] That is to say, the weight information can be used to select the initial labels in a label sequence, and the association probability between labels in the fourth graph information can be used to select the wandering labels in this label sequence, that is, the kth label, where k is an integer greater than or equal to 2. Among them, the number of wandering labels in the label sequence can be preset.
[0123] Exemplarily, the selection probability of each label can be set based on the weight information of each label, and a random function can be set based on this probability, and the random function can be used to select the initial labels.
[0124] One exemplary probability setting method is to normalize the weight information of each label and use the normalized weight information as the selection probability of the label.
[0125] Another exemplary probability setting method is to normalize the weight information of each label and determine the selection probability of the label based on a preset probability ratio and the normalized weight information. For example, assume there are n labels in total, and the preset probability ratio is That is, set The probability of randomly selecting from all n labels, and The probability of selecting according to the normalized weight vector For selection. Among them, P gThe i-th element in represents the weight information of the i-th label x i , denoted as That is, the selection probability of the i-th label x i is
[0126] Exemplarily, the way to select multiple wandering labels from multiple labels can be to perform multiple iterations. Among them, the process of the q-th iteration is: based on the information corresponding to the (q - 1)-th wandering label in the fourth graph information, select the q-th wandering label from multiple labels, where q is an integer greater than or equal to 1. Among them, the initial label can be regarded as the 0-th wandering label. In the case of q = l, the selection of the wandering label ends. Among them, l is the number of wandering labels in the label sequence. In practical applications, in order to obtain more accurate label representation information, l can be set as an integer greater than or equal to 4.
[0127] The above steps of selecting the label sequence based on the weight information of each label and the fourth graph information can be iteratively executed m times. Each time a label sequence is obtained, then m label sequences are obtained. Among them, m is a preset value.
[0128] After obtaining multiple label sequences, based on the multiple label sequences and the word vector generation model, the initial representation vectors of each label among the multiple labels can be processed to obtain the target representation vectors of each label. Specifically, it can also be iterated multiple times to optimize the representation vectors of each label. When the iteration reaches the preset conditions for the representation vectors, the target representation vectors are obtained. The target representation vectors can be used as the feature information of the labels in downstream application tasks. Exemplarily, the word vector generation model (Word2vec model) can be the skip-gram model
[0129] Exemplarily, in each iteration process, a label sequence can be randomly selected from multiple label sequences, and multiple subsequences can be determined in this label sequence using a time sliding window with a preset length. Based on the word vector model and the central words in the subsequences, other words in the subsequences are predicted to obtain the prediction results of other words, and the loss is calculated based on the current representation vectors of other words and the prediction results. Based on the loss value, the current representation vectors are updated. Through iterations with a preset number of learning times, accurate representation vectors of each word (i.e., each label) can be obtained as the target representation vectors.
[0130] Exemplarily, the embodiments of the present application also provide an exemplary application process of the above target representation vectors. Specifically, the above method further includes:
[0131] Based on the target representation vector of the first label among multiple labels and the deep learning model, the content or entity corresponding to the first label is processed to obtain the prediction information corresponding to the content or entity.
[0132] Exemplarily, the first tag can be any one of a plurality of tags. Processing the content corresponding to the first tag, for example, predicting the next content to be recommended to the user based on the content containing the first tag. Processing the entity corresponding to the first tag, for example, predicting the tourism popularity of the location when the first tag is a location entity; predicting the consumption ability corresponding to the person's name when the first tag is a person name entity, etc.
[0133] To more clearly present the technical idea of the present application, taking the first content set as a news article set as an example, a specific application example is provided below to introduce the method of the embodiments of the present application. Figure 5 A schematic diagram of the application example is shown. As Figure 5 shown, in this application example, news articles are used as the content to be processed. After tag extraction (including entity tag extraction and core word tag extraction), the tags are processed through three stages: obtaining graph information, calculating the weight information of tag nodes, and obtaining the representation vectors of tag nodes. It can be understood that in actual applications, the technical details in the following application examples can be flexibly set according to the actual application scenarios, for example, flexibly combining other technical details according to the description of the above embodiments.
[0134] 1. Obtain graph information
[0135] a. First, delimit the time range of the news article source according to timeliness. Generally speaking, news within one year can be selected as the construction source of the tag graph. That is, the first content set in the embodiments of the present application can include news articles within one year.
[0136] b. Summarize the tagging results of each news article within the delimited range. Both entity tags and core word tags are used as ordinary nodes in the tag graph;
[0137] c. Suppose the tagging results of an article are summarized as [(x 1 , c 1 ), (x 2 , c 2 ),..., (x n , c n ), where x i is the i-th tag word, and c i is the tag type of this word (a certain entity type or core word). Traverse this list. When traversing to the i-th tuple, use x i as the starting node, and use the words x j in the remaining n - 1 tuples as the ending nodes (i ≠ j) to establish an edge. And so on, after traversing n tuples, complete the construction of the edges between tag nodes in a single article;
[0138] d. Suppose the time range obtained in step a contains m manuscripts. Traverse these m manuscripts using the method in step c to obtain the statistical results of all nodes and the edges between them on the m manuscripts. This statistical result can be regarded as a kind of graph spectrum information. Suppose a total of n nodes are obtained on the m manuscripts, then the summary result can be represented in the form of the following map:
[0139] map = {x 1 : {x 2 : t 12 , x 3 : t 13 ,...}, x 2 : {x 1 : t 21 , x 3 : t 23 ...},..., x n : {x 1 : t n1 ,...}}.
[0140] Among them, x i is the i-th labeled word node, and t ij represents the number of occurrences of the edge with x i as the starting node and x j as the ending node (i ≠ j). When there is no edge between x i and x j , t ij = 0.
[0141] e. To reduce the impact of "the contingency of two unrelated words co-occurring in a piece of content" on subsequent work, the edges and nodes in the map can be filtered. For example, only retain the edges above t ij in t min and the nodes connected by these edges, that is, for t min less than or equal to t ij , they can be set to zero. Among them, the threshold t min is a positive integer.
[0142] f. Perform relevant tag statistics based on user behavior. In the information flow platform, the consumption behavior of users is another basis for characterizing tag relevance. Since the adjacent content in the user's click sequence often has semantic relevance, associations can also be established between different tags contained in the adjacent content respectively. Let the statistical adjacent range be a positive integer e, then the tags that appear in e articles clicked adjacent by the user can be associated. The click sequence of the user can be randomly selected, and a time sliding window with a width of e is established. Suppose a piece of content in the time sliding window contains p tags, and the remaining e - 1 manuscripts contain q tags (after deduplication). Then, for each of these q tags, a non - co - occurrence set of associations is established with each of the p tags. The above operations are performed for all manuscripts in all time sliding windows in all selected user click sequences, and another set of association statistical results consume = {x 1 : {x 2 : c 12 : x 3 : c 13 :,...}, x 2 : {x 1 : c 21 : x 3 : c 23 ...},...}, x n : {x 1 : c n1 :,...}} can be obtained. This statistical result can be regarded as another type of graph information. Where x i is the i - th tag word, and c mn is the number of times x m and x n co - occur in the time sliding window with a width of e.
[0143] g. To reduce the impact of "the contingency of two unrelated words appearing in the same set of user click sequences" on subsequent work, following step e, only retain the edges where c mn > c min in consume and the nodes connected by these edges, where the threshold c min is a positive integer.
[0144] 2. Calculate the weight information of tag nodes
[0145] In this stage, based on the statistical results in map obtained in the previous stage, through several rounds of iteration of the Pagerank (Page Ranking) algorithm, the importance of various nodes is evaluated, that is, the weight information of each tag node is obtained. Among them, the Pagerank algorithm is a web page propagation algorithm that evaluates the propagation power of a website through website jump relationships. In this application example, its core idea is applied to tag processing.
[0146] a. Let the map obtained after the previous step be map = {x 1 : {x 2 : t 12 ,x 3 : t 13 ,...}, x 2 : {x 1 : t 21 ,x 3 : t 23 ...},...,x n : {x 1 : t n1 ,...}} which contains n nodes. Define the initial weight vector of these n nodes as That is, the initial weight of each label is
[0147] b. Without considering random search, calculate the basic transition probability matrix. Specifically, construct an n*n matrix T according to the map, and the element in the i-th row and j-th column of it represents the probability of transferring from x i to x j . Its initial value is That is, the proportion of the co-occurrence frequency of x i and x j in all co-occurrence frequencies of x i ;
[0148] c. Introduce the random search probability into the basic transition probability matrix, and combine the user behavior data with the co-occurrence relationship skillfully. Let the random search probability be rsp ∈ (0, 1), and its meaning is: when the starting node is x i , assume that there are size(comsume(x i )) nodes that have a consumption relationship with x i , then the probability that its ending node is any one of these size(comsume(x i )) nodes, and the probability of (1 - rsp) is calculated according to the formula in step b. Considering the definition of the transition probability matrix, for the modified transition matrix T′, the element t′ ij in its i-th row and j-th column represents the probability of transferring from x i to x j . When x j is one of the size(comsume(x i )) consumption-related nodes of x i , When x j is not one of the size(comsume(x i )) nodes of x iWhen any one of the consumption-related nodes
[0149] d. The algorithm iteration starts. Let the current be the t-th step. Calculate P t = P t-1 ·T′, where P t and P t-1 represent the weights of n nodes at the t-th step and the (t - 1)-th step respectively. The termination condition of the iteration is ‖P t - P t-1 ‖ ≤ v, where ‖·‖ represents the Euclidean norm and v is a threshold;
[0150] e. When the algorithm iteration ends, assume that the iteration has completed h steps. Then P h is the weight vector of the finally obtained n nodes;
[0151] f. For the i-th node x i among the n nodes, let x i have a weight of h in P x j is any termination node in the map. Then the weight value of the edge connecting from x i to x j is The edge weight result obtained by correcting the map in this way is
[0152] 3. Obtain the representation vector of the label node
[0153] In this stage, a random walk method is adopted to construct a label sequence. First, the starting node is selected with a probability, and jumps are made according to the edge weights in 2.f to generate a label sequence of a certain length. These sequences are input into the skip-gram model to generate the embedding (embedded representation) of each label node.
[0154] a. Calculate the probability of randomly selecting each label node. The i-th value of the vector P h represents the importance of the i-th node x i . Perform a normalization mapping on P h to obtain P g , where Then the normalized importance of the label node x i after normalization is When selecting the initial node of the label sequence, randomly select from all n nodes with a probability of , and select according to P with a probability of gAs the initial selection probability of each node. The purpose of doing this is to avoid almost always selecting high-frequency nodes and reasonably reduce the gap between high-frequency nodes and low-frequency nodes;
[0155] b. First, select an initial node of a label sequence according to the probability in step a;
[0156] c. Select the next-hop result of the walk. Let the current node be x i , normalize the weights of each edge in the normalized_map in 2.f, and select the next-hop node according to this normalized weight as the probability;
[0157] d. Let l be the longest walk length, l≥4. If the length of the sequence obtained after the current jump does not reach l, return to step c; otherwise, terminate and obtain the label node sequence s of length l this time 1 ;
[0158] e. Repeat the operations in b~d to obtain a set S={s 1 , s 2 ,..., s m} containing m node sequences;
[0159] f. Build a skip-gram model to train S to obtain the embedding representation vector of each node. First, randomly initialize the hidden layer vectors of the n label nodes included in the label graph, that is, establish an n*d matrix V. Then the d-dimensional vector v i of the i-th row is the initial hidden layer vector of the i-th label node x i among the n label nodes;
[0160] g. Randomly initialize the output layer parameters W of the neural network, with a dimension of d*n;
[0161] h. Randomly select a label sequence, establish a sliding window with a width of (2ω + 1), where ω is a positive integer and 2ω + 1<l, and let the sliding window slide from left to right on each selected label sequence. Suppose the sliding window is {x a , x b , x k , x c , x d} at a certain moment;
[0162] i. Select the center word v k in the sliding window as the input word, and use the center word v k to predict the other predicted label nodes that appear in the sliding window. For the center word in the sliding window being x k , take the k-th row vector v k in V, and calculate Φ(v k ) = v kW, the calculation result is centered around the word x k Predict the results of other points within the prediction sliding window;
[0163] j. Evaluate other nodes within the sliding window, i.e., calculate the loss loss = -logPr({v a , v b , v c , v d}|Φ(v k ));
[0164] k. Through gradient descent, reduce the loss in j and update the hidden layer matrix V with the learning rate α. Repeat the above processes of sequence selection, sliding window movement, and parameter learning until the preset number of learning rounds is reached;
[0165] l. At the end of learning, the d-dimensional vector v i of the i-th row of the hidden layer matrix V is the embedding vector result of the i-th label node x i among the n label nodes.
[0166] It can be seen that in the above application example, when calculating the importance of each node using the Pagerank algorithm, both ordinary co-occurrence features are considered, and user consumption behavior information is cleverly incorporated through random search. Based on the obtained node importance, a fusion-based calculation method for evaluating the association relationship between labels is proposed, taking both node importance and co-occurrence frequency into account in the calculation formula.
[0167] Furthermore, using the constructed label graph, the embedding representation vectors of each label node are learned using the graphembedding (graph embedding representation) technique for use by various deep learning models in downstream tasks.
[0168] Corresponding to the application scenario and method of the method provided in the embodiments of the present application, the embodiments of the present application also provide a label processing device 600. Refer to Figure 6 , the device 600 may include:
[0169] A sequence determination module 610 for determining a first user behavior sequence; wherein, the first user behavior sequence includes multiple contents continuously processed by the user;
[0170] A jump determination module 620 for determining multiple label jump relationships based on multiple labels corresponding to the multiple contents;
[0171] A first graph determination module 630 for determining first graph information based on the multiple label jump relationships.
[0172] Exemplarily, the tag processing device 600 may further include:
[0173] A content recommendation module, configured to determine a second content to be recommended to a user based on tags of a first content processed by the user and first graph information.
[0174] Exemplarily, the sequence determination module 610 may include:
[0175] A sequence selection unit, configured to select a second user behavior sequence from a plurality of user behavior sequences associated with a first content set;
[0176] A subsequence determination unit, configured to determine at least one subsequence in the second user behavior sequence by using a time sliding window with a preset length; wherein, the at least one subsequence includes a first user behavior sequence;
[0177] The first graph determination module 630 may include:
[0178] A relationship addition unit, configured to add a plurality of tag jump relationships to a tag jump relationship set corresponding to the first content set;
[0179] An information determination unit, configured to determine first graph information corresponding to the first content set based on the tag jump relationship set.
[0180] Exemplarily, the information determination unit is specifically configured to:
[0181] Obtain second graph information based on the tag jump relationship set; wherein, the second graph information includes Y groups of jump times corresponding to Y tags; wherein, the i-th group of jump times in the Y groups of jump times includes the jump times between the i-th tag among the Y tags and each other tag, Y is an integer greater than or equal to 1, and i is a positive integer less than or equal to Y;
[0182] Process the second graph information based on a preset number threshold to obtain the first graph information.
[0183] Exemplarily, the tag processing device may further include:
[0184] A co-occurrence times determination module, configured to determine the co-occurrence times between any two of a plurality of tags corresponding to the first content set;
[0185] A third graph determination module, configured to determine third graph information based on the co-occurrence times; wherein, the third graph information includes X groups of co-occurrence times corresponding to X tags; wherein, the j-th group of co-occurrence times in the X groups of co-occurrence times includes the co-occurrence times between the j-th tag among the X tags and each other tag, X is an integer greater than or equal to 1, and j is a positive integer less than or equal to X;
[0186] The fourth graph determination module is configured to obtain the fourth graph information corresponding to the first content set based on the first graph information and the third graph information.
[0187] Exemplarily, the fourth graph determination module may include:
[0188] The association probability determination unit is configured to determine the association probability between any two of the multiple labels corresponding to the first content set based on the first graph information, the third graph information, and a preset random search probability;
[0189] The weight determination unit is configured to determine the weight information of each label among the multiple labels based on the association probability between any two of the multiple labels;
[0190] The graph processing unit is configured to obtain the fourth graph information based on the weight information of each label.
[0191] Exemplarily, the weight determination unit is specifically configured to:
[0192] Perform multiple iterations based on an initial weight vector and the association probability between any two of the multiple labels to obtain the weight information of each label;
[0193] Wherein, the t-th iteration in the multiple iterations includes:
[0194] Obtain the (t + 1)-th weight vector based on the t-th weight vector and a probability matrix, where the probability matrix is used to represent the association probability between any two of the multiple labels;
[0195] When the (t + 1)-th weight vector meets a preset condition, obtain the weight information of each label based on the (t + 1)-th weight vector, where t is an integer greater than or equal to 1.
[0196] Exemplarily, the graph processing unit is specifically configured to:
[0197] Process the third graph information based on the weight information of each label to obtain the fourth graph information.
[0198] Exemplarily, the label processing device further includes:
[0199] The label sequence determination unit is configured to determine multiple label sequences based on the weight information of the multiple labels and the fourth graph information;
[0200] The representation vector determination unit is configured to process the initial representation vectors of each label among the multiple labels based on the multiple label sequences and a word vector generation model to obtain the target representation vectors of each label.
[0201] Exemplarily, the label processing device further includes:
[0202] A characterization application module is used to process the content or entity corresponding to the first tag based on the target characterization vector of the first tag among multiple tags and a deep learning model, so as to obtain prediction information corresponding to the content or entity.
[0203] Exemplarily, the tag sequence determination unit is specifically configured to:
[0204] Select an initial tag from multiple tags based on the weight information of the multiple tags;
[0205] Select multiple wandering tags from multiple tags based on the initial tag and the fourth atlas information;
[0206] Obtain the first tag sequence among multiple tag sequences based on the initial tag and the multiple wandering tags.
[0207] For the functions of the modules in each device in the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, and they have corresponding beneficial effects, which will not be elaborated here.
[0208] The embodiments of the present application further provide an electronic device for implementing the above method. Figure 7 The structural block diagram of the electronic device according to the embodiments of the present application is shown. As Figure 7 shown, the electronic device includes: a memory 710 and a processor 720. A computer program that can run on the processor 720 is stored in the memory 710. When the processor 720 executes the computer program, the tag processing method in the above embodiments is implemented. The number of the memory 710 and the processor 720 can be one or more.
[0209] The electronic device further includes:
[0210] A communication interface 730, which is used to communicate with external devices and perform data interaction and transmission.
[0211] If the memory 710, the processor 720, and the communication interface 730 are implemented independently, the memory 710, the processor 720, and the communication interface 730 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0212] Optionally, in a specific implementation, if the memory 710, the processor 720, and the communication interface 730 are integrated on a single chip, the memory 710, the processor 720, and the communication interface 730 can communicate with each other through an internal interface.
[0213] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, which when executed by a processor implements the method provided in any embodiment of the present application.
[0214] The embodiments of the present application also provide a computer program product including a computer program, which when executed by a processor implements the method provided in any embodiment of the present application.
[0215] The embodiments of the present application also provide a chip including a processor for calling and running instructions stored in a memory, so that a communication device equipped with the chip executes the method provided in the embodiments of the present application.
[0216] The embodiments of the present application also provide a chip including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is configured to execute code in the memory, and when the code is executed, the processor is configured to execute the method provided in the embodiments of the application.
[0217] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor supporting the advanced reduced instruction set machine (ARM) architecture.
[0218] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may further include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0219] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0220] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0221] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality of" means two or more unless otherwise specifically defined.
[0222] Any process or method description represented in the flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed.
[0223] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatus, or devices.
[0224] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by a program instructing relevant hardware, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiment.
[0225] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. This storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.
[0226] As described above, the above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of various changes or substitutions, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for label processing, including: Determine a first user behavior sequence; wherein, the first user behavior sequence includes a plurality of contents continuously processed by the user; Based on the plurality of labels corresponding to the plurality of contents, determine a plurality of label jump relationships; Based on the plurality of label jump relationships, determine first graph information, including: adding the plurality of label jump relationships to a label jump relationship set corresponding to a first content set processed by the user; based on the label jump relationship set, determine the first graph information corresponding to the first content set; Wherein, the label jump relationship set is obtained by summarizing the label jump relationships corresponding to each subsequence included in the first user behavior sequence.
2. The method according to claim 1, wherein, the method further includes: Based on the label of the first content processed by the user and the first graph information, determine a second content recommended to the user.
3. The method according to claim 1 or 2, wherein, the determining the first user behavior sequence includes: Among the plurality of user behavior sequences associated with the first content set, select a second user behavior sequence; Use a time sliding window with a preset length to determine at least one subsequence in the second user behavior sequence; wherein, the at least one subsequence includes the first user behavior sequence.
4. The method according to claim 3, wherein, the determining the first graph information corresponding to the first content set based on the label jump relationship set includes: Based on the label jump relationship set, obtain second graph information; wherein, the second graph information includes Y groups of jump times corresponding to Y labels; wherein, the i-th group of jump times in the Y groups of jump times includes the jump times between the i-th label among the Y labels and each other label, Y is an integer greater than or equal to 1, and i is a positive integer less than or equal to Y; Process the second graph information based on a preset number threshold to obtain the first graph information.
5. The method according to claim 1 or 2, wherein, the method further includes: Determine the co-occurrence times between any two of the plurality of labels corresponding to the first content set; Based on the co-occurrence times, determine third graph information; wherein, the third graph information includes X groups of co-occurrence times corresponding to X labels; wherein, the j-th group of co-occurrence times in the X groups of co-occurrence times includes the co-occurrence times between the j-th label among the X labels and each other label, X is an integer greater than or equal to 1, and j is a positive integer less than or equal to X; Based on the first graph information and the third graph information, obtain fourth graph information corresponding to the first content set.
6. The method according to claim 5, wherein, the obtaining the fourth graph information corresponding to the first content set based on the first graph information and the third graph information includes: Based on the first graph information, the third graph information, and a preset random search probability, determine the association probability between any two of the plurality of labels corresponding to the first content set; Based on the association probability between any two of the plurality of labels, determine the weight information of each label in the plurality of labels; Based on the weight information of each of the tags, the fourth graph information is obtained.
7. The method according to claim 6, wherein, the determining the weight information of each of the multiple tags based on the association probabilities between the multiple tags pairwise includes: performing multiple iterations based on an initial weight vector and the association probabilities between the multiple tags pairwise to obtain the weight information of each of the tags; wherein the t-th iteration in the multiple iterations includes: obtaining a (t + 1)-th weight vector based on the t-th weight vector and a probability matrix; wherein the probability matrix is used to represent the association probabilities between the multiple tags pairwise; when the (t + 1)-th weight vector meets a preset condition, obtaining the weight information of each of the tags based on the (t + 1)-th weight vector, where t is an integer greater than or equal to 1.
8. The method according to claim 6, wherein, the obtaining the fourth graph information based on the weight information of each of the tags includes: processing the third graph information based on the weight information of each of the tags to obtain the fourth graph information.
9. The method according to claim 6, wherein, the method further includes: determining a plurality of tag sequences based on the weight information of each of the tags and the fourth graph information; processing the initial representation vectors of each of the multiple tags based on the plurality of tag sequences and a word vector generation model to obtain the target representation vectors of each of the tags.
10. The method according to claim 9, wherein, the method further includes: processing the content or entity corresponding to the first tag based on the target representation vector of the first tag in the multiple tags and a deep learning model to obtain the prediction information corresponding to the content or entity.
11. The method according to claim 9, wherein, the determining a plurality of tag sequences based on the weight information of each of the tags and the fourth graph information includes: selecting an initial tag from the multiple tags based on the weight information of each of the tags; selecting a plurality of wandering tags from the multiple tags based on the initial tag and the fourth graph information; obtaining a first tag sequence in the plurality of tag sequences based on the initial tag and the plurality of wandering tags.
12. A tag processing device, comprising: a sequence determining module, configured to determine a first user behavior sequence; wherein the first user behavior sequence includes a plurality of contents continuously processed by a user; a jump determining module, configured to determine a plurality of tag jump relationships based on the plurality of tags corresponding to the plurality of contents; a first graph determining module, configured to determine first graph information based on the plurality of tag jump relationships, including: adding the plurality of tag jump relationships to a tag jump relationship set corresponding to a first content set processed by a user; determining first graph information corresponding to the first content set based on the tag jump relationship set; wherein the tag jump relationship set is obtained by summarizing the tag jump relationships corresponding to each subsequence included in the first user behavior sequence.
13. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, wherein the processor implements the method according to any one of claims 1-11 when executing the computer program.
14. A computer-readable storage medium, in which a computer program is stored, and the computer program implements the method according to any one of claims 1-11 when being executed by a processor.
Citation Information
Patent Citations
Personalized page configuration method and device, electronic equipment, medium and program product
CN114398572A