Information recommendation method, device, computer equipment, storage medium and program product
By training the embedded network to extract information title feature vectors and cluster them, and generate event cluster identifiers, the problem of duplicate content in the information flow content service platform is solved, and the efficiency and accuracy of information recommendation are improved.
Patent Information
- Application Number
- CN202210365522.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-07
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2042-04-07
AI Technical Summary
There is a large amount of duplicate content in the information flow content service platform, and the existing technology is difficult to effectively identify and filter, resulting in low efficiency in information recommendation.
By training the title feature vectors that extract information into the embedding network, perform vector clustering, generate event cluster identification, delete information that meets the similar conditions of event clusters in the candidate recommendation list, and generate a recommendation list.
Improve the efficiency and accuracy of information recommendation, avoid similar information in the same recommendation list, and enhance the richness and diversity of recommendation lists.
Smart Images

Figure CN116932880B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to technical fields such as artificial intelligence and big data, and relates to an information recommendation method, apparatus, computer equipment, storage medium, and program product. Background Art
[0002] In the era of rapid internet development, information flow content services have become widely popular, and a large amount of high-quality original content has emerged on information flow content service platforms. At the same time, some content publishers, driven by profit, plagiarize or even directly copy the content of other original creators, resulting in a large amount of duplicate content on information flow content service platforms. The high volume of duplicate content on information flow content service platforms not only harms the interests of original creators but also negatively impacts the healthy development of the entire content ecosystem.
[0003] In related technologies, deduplication detection is usually performed on the content in the information flow content service platform. Deduplication detection is to detect whether the keywords included in the title of each content are repeated, and filter out the content with repeated title keywords for recommendation.
[0004] However, some reposters can easily circumvent duplicate detection by modifying the title, resulting in the information recommended to the client still containing a lot of duplicate and similar content, making the actual amount of recommended information small. Therefore, the actual recommendation efficiency of the above information recommendation is low. Summary of the Invention
[0005] This application provides a method, apparatus, computer device, storage medium, and program product for information recommendation, which can solve the problem of low actual recommendation efficiency of information recommendation in related technologies. The technical solution is as follows:
[0006] In one aspect, a method for information recommendation is provided, the method comprising:
[0007] In response to a recommendation request from any object, based on event cluster identifiers of at least two pieces of information, delete information that meets the event cluster similarity condition from a candidate recommendation list to obtain a recommendation list, and recommend the to-be-recommended information flow corresponding to the recommendation list to the object;
[0008] The event cluster similarity condition includes being the same as the event cluster identifier of any other information in the candidate recommendation list;
[0009] The method for obtaining the event cluster identifier includes:
[0010] Obtain the title feature vector of each information in the resource pool through the trained embedding network;
[0011] Performing vector clustering on the title feature vectors of each information in the resource pool to obtain at least two event clusters, and marking the information included in each event cluster with an event cluster identifier corresponding to the event cluster, wherein each event cluster includes at least one information belonging to the same event;
[0012] The embedding network is trained based on a first similarity between the anchor sample titles in at least two triples and the positive sample titles, and a second similarity between the anchor sample titles and the negative sample titles;
[0013] Each triplet includes an anchor sample title, a positive sample title, and a negative sample title. The anchor sample title and the positive sample title belong to the same event, and the anchor sample title and the negative sample title belong to different events.
[0014] In a possible implementation, each information in the resource pool includes at least two forms.
[0015] In another aspect, an information recommendation device is provided, comprising:
[0016] A recommendation list determination module is configured to, in response to a recommendation request for any object, delete information that meets the event cluster similarity condition from a candidate recommendation list based on event cluster identifiers of at least two pieces of information to obtain a recommendation list;
[0017] A recommendation module, configured to recommend the information flow to be recommended corresponding to the recommendation list to any of the objects;
[0018] The event cluster similarity condition includes being the same as the event cluster identifier of any other information in the candidate recommendation list;
[0019] The device is further configured to obtain the event cluster identifier. When obtaining the event cluster identifier, the device further includes:
[0020] The title feature vector acquisition module is used to obtain the title feature vector of each information in the resource pool through the trained embedding network;
[0021] a clustering module configured to perform vector clustering on the title feature vectors of each message in the resource pool to obtain at least two event clusters, and mark the information included in each event cluster with an event cluster identifier corresponding to the event cluster, wherein each event cluster includes at least one message belonging to the same event;
[0022] The embedding network is trained based on a first similarity between the anchor sample titles in at least two triples and the positive sample titles, and a second similarity between the anchor sample titles and the negative sample titles;
[0023] Each triplet includes an anchor sample title, a positive sample title, and a negative sample title. The anchor sample title and the positive sample title belong to the same event, and the anchor sample title and the negative sample title belong to different events.
[0024] In one possible implementation, the apparatus is further configured to train the embedding network. When training the embedding network, the apparatus further includes:
[0025] A sample data set acquisition module is used to acquire a sample data set and a label of each sample in the sample data set;
[0026] Each sample includes a base sample title and a candidate sample title, and the label of each sample includes a first event label and a second event label, wherein the first event label indicates whether the candidate sample title and the base sample title belong to the same event, and the second event label indicates whether the text corresponding to the candidate sample title and the base sample title respectively belongs to the same event;
[0027] A triplet construction module, configured to construct the at least two triples based on the label of each sample in the sample data set;
[0028] a similarity determination module, configured to determine, based on the feature vectors of the anchor sample title, the positive sample title, and the negative sample title in each triplet, a first similarity between the anchor sample title and the positive sample title, and a second similarity between the anchor sample title and the negative sample title, respectively, wherein the feature vectors of the anchor sample title, the positive sample title, and the negative sample title are obtained by extracting features from the anchor sample title, the positive sample title, and the negative sample title, respectively, through the initial embedding network;
[0029] A training module is configured to train the initial embedding network based on the difference between the first similarity and the second similarity to obtain the embedding network.
[0030] In one possible implementation, the training module is further configured to iteratively train the initial embedding network when the difference between the second similarity and the first similarity is not higher than a target value, and stop training until the second similarity is higher than the target value of the first similarity, thereby obtaining the embedding model.
[0031] In one possible implementation, the triple construction module includes:
[0032] a title pair acquiring unit, configured to acquire at least two title pairs from the sample dataset based on the label of each sample in the sample dataset, each title pair comprising an anchor sample title and a positive sample title belonging to the same event;
[0033] a distance determining unit configured to obtain, for each title pair, a first distance between the anchor sample title and the positive sample title based on their respective feature vectors, and obtain a second distance between at least one candidate sample title and the anchor sample title based on their respective feature vectors;
[0034] A triplet determination unit, configured to determine a first type of triplet, a second type of triplet, and a third type of triplet based on the first distance and the second distance corresponding to each title pair;
[0035] Among them, the difference between the second distance and the first distance corresponding to the first type triplet is greater than the target value, the first distance corresponding to the second type triplet is less than the second distance, and the sum of the first distance and the target value is greater than the second distance, and the second distance corresponding to the third type triplet is less than the first distance.
[0036] In one possible implementation, the title pair acquiring unit is configured to:
[0037] For each iterative training, obtaining batch sample data used in the current iterative training from the sample data set;
[0038] Based on the label of each sample in the batch sample data, obtaining the at least two title pairs from the batch sample data set;
[0039] Accordingly, the device further includes:
[0040] The candidate sample title obtaining unit is configured to obtain the at least one candidate sample title from the batch sample data.
[0041] In one possible implementation, the triple construction module includes:
[0042] A second title pair acquisition unit, configured to acquire a title pair based on the label of each sample in the sample dataset, wherein the title pair includes an anchor sample title and a positive sample title belonging to the same event;
[0043] A negative sample title sampling unit, configured to randomly sample a target number of sample titles for each title pair as negative sample titles corresponding to the title pair;
[0044] The triplet acquisition unit is configured to obtain the at least two triplets based on each title pair and its corresponding negative sample title.
[0045] In one possible implementation, the sample data set acquisition module includes:
[0046] a duplication detection unit, configured to perform duplication detection on the original information set to obtain at least two information subsets, each information subset including at least one information having a similarity within a target threshold range;
[0047] a sample acquisition unit configured to, for each information subset, use the title of target information in the information subset as a reference sample title, and acquire a candidate sample title corresponding to the reference sample title to obtain a sample;
[0048] A first event label determination unit is configured to determine, for each sample, a first event label corresponding to the benchmark sample title and the corresponding candidate sample title based on the title keywords respectively included in the benchmark sample title and the corresponding candidate sample title;
[0049] The second event label determination unit is configured to determine the second event labels corresponding to the benchmark sample title and its corresponding candidate sample title based on the respective texts corresponding to the benchmark sample title and its corresponding candidate sample title.
[0050] In one possible implementation, the apparatus further includes:
[0051] an adding module, configured to add new information to the candidate recommendation list based on the event cluster identifiers of the remaining information in the candidate recommendation list to obtain the recommendation list;
[0052] The event cluster identifier of the newly added information is different from the remaining information, and the remaining information is information other than the deleted information in the candidate recommendation list.
[0053] In a possible implementation, each information in the resource pool includes at least two forms.
[0054] On the other hand, a computer device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above-mentioned information recommendation method.
[0055] On the other hand, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned information recommendation method is implemented.
[0056] On the other hand, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the computer program implements the above-mentioned information recommendation method.
[0057] The beneficial effects of the technical solution provided by the embodiments of the present application are:
[0058] The information recommendation method provided by the present application obtains a recommendation list by deleting information that meets the event cluster similarity conditions from a candidate recommendation list based on event cluster identifiers of at least two pieces of information, and then makes recommendations. This method filters similar information belonging to the same event cluster through the event cluster identifier, preventing similar information from being concentrated in the same recommendation list, thereby improving recommendation efficiency. Furthermore, the method obtains the title feature vectors of each piece of information through an embedded network, performs vector clustering on the title feature vectors of each piece of information, and obtains at least two event clusters, each of which includes at least one piece of information belonging to the same event. This method divides each piece of information into different event clusters, allowing the use of the event cluster to quickly filter and disperse similar information of the same event in the candidate recommendation list, thereby improving the richness and diversity of the recommendation list and, in turn, improving actual recommendation efficiency.
[0059] Moreover, the embedding network is trained based on the first similarity between the anchor sample titles and the positive sample titles, and the second similarity between the anchor sample titles and the negative sample titles in at least two triples; while the anchor sample titles and the positive sample titles belong to the same event, and the anchor sample titles and the negative sample titles belong to different events. The present application uses the embedding network to perform feature extraction, thereby achieving the effect of effectively bringing the title features of the same event closer and the distinctive features of different events farther apart in the feature space, thereby improving the differentiated expression of the title feature vector for different events, thereby improving the accuracy of clustering, and further improving the accuracy of information recommendation. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0061] Figure 1 A schematic diagram of an implementation environment of an information recommendation method provided in an embodiment of the present application;
[0062] Figure 2 A flow chart of a method for obtaining an event cluster identifier provided in an embodiment of the present application;
[0063] Figure 3 A schematic diagram of a training process framework for an embedded network provided in an embodiment of the present application;
[0064] Figure 4 A schematic diagram of multiple types of triples provided in an embodiment of the present application;
[0065] Figure 5 A schematic diagram of a training process of an embedded network provided in an embodiment of the present application;
[0066] Figure 6 A signaling interaction diagram of an information recommendation method provided in an embodiment of the present application;
[0067] Figure 7 A schematic diagram of an information display page provided in an embodiment of the present application;
[0068] Figure 8 A schematic diagram of a framework of an information recommendation method provided in an embodiment of the present application;
[0069] Figure 9 A schematic diagram of the structure of an information recommendation device provided in an embodiment of the present application;
[0070] Figure 10 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0071] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0072] Those skilled in the art will understand that, unless otherwise specified, the singular forms "a," "an," "the," and "the" used herein may also include the plural forms. The terms "including" and "comprising" used in the embodiments of this application mean that the corresponding features can be implemented as the presented features, information, data, steps, and operations, but do not exclude the implementation of other features, information, data, steps, operations, etc. supported by the technical field.
[0073] Figure 1 This is a schematic diagram of the implementation environment of an information recommendation method provided by this application. Figure 1 As shown, the implementation environment includes: a server 101 and a terminal 102. The server 101 and the terminal 102 can communicate directly or indirectly via wired or wireless communication.
[0074] In an embodiment of the present application, the terminal 102 may be installed with a target application having an information recommendation function, and the server 101 is configured to provide the information recommendation function to the terminal 102. For example, the server 101 may be a backend server of the target application. The terminal 102 may send a recommendation request to the server, and the server 101 may send a recommendation list to the terminal 102 based on the recommendation request, thereby recommending the information flow corresponding to the recommendation list to the terminal 102.
[0075] In one possible example, the information recommendation function can be a function that recommends information streams based on event cluster identifiers for each piece of information in the resource pool. For example, the target application can be a video application, a news application, a social application, an interactive entertainment application, a browser application, a shopping application, a content sharing application, a virtual reality (VR) application, an augmented reality (AR) application, etc., although this embodiment of the present application is not limiting in this regard. Furthermore, different applications may push different information and have different corresponding functions, which can be pre-configured based on actual needs, but this embodiment of the present application is not limiting in this regard. For example, terminal 102 may be running a client for one of the aforementioned applications. The above-mentioned information stream recommendation service covers information from various verticals, including variety shows, film and television, news, finance, sports, entertainment, and games. The information can be in any form, including, but not limited to, videos, articles, and graphics. Through this information stream recommendation service, users can enjoy information recommendations in a variety of formats, including articles, images, short videos, live broadcasts, special topics, and columns.
[0076] Server 101 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, and big data and artificial intelligence platforms. Terminal 102 can be a smartphone, tablet computer, laptop computer, digital broadcast receiver, desktop computer, in-vehicle terminal (such as an in-vehicle navigation terminal, in-vehicle computer, etc.), smart speaker, smartwatch, etc. Terminal 102 and server 101 can be connected directly or indirectly via wired or wireless communication, and can also be determined based on the actual application scenario requirements and are not limited here.
[0077] The similar video detection methods provided in the embodiments of this application involve artificial intelligence (AI) and machine learning technologies. For example, they utilize cloud computing within AI to automatically generate a list of recommended videos. Another example involves training an embedded network using machine learning to extract feature vectors of titles for each piece of information. The following provides a brief description of these techniques to facilitate understanding by those skilled in the art.
[0078] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0079] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and smart transportation.
[0080] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0081] Before introducing the method embodiments provided in the present application, a brief introduction is first given to the application scenarios, relevant terms or nouns that may be involved in the method embodiments of the present application to facilitate understanding by technical personnel in the field of the present application.
[0082] Social networks originated from online social networking, which began with email. The internet is essentially a network of computers. Early email solved the problem of long-distance email transmission and remains the most popular internet application to this day. It also marked the beginning of online social networking. Bulletin Board Systems (BBSs) took this a step further by normalizing group messaging and forwarding, theoretically enabling the ability to publish information and discuss topics to anyone. They became early platforms for the spontaneous generation of internet content. In the past two years, with the widespread adoption of smartphones, the ubiquity of Wi-Fi (wireless communication technology), the generally lower 4G (fourth generation mobile communication technology) rates, and the impending arrival of the 5G (fifth generation mobile communication technology) era, within the current context of the mobile internet, users' demand for information is shifting from text and images to video. Therefore, short videos will gradually become a dominant form of content on the mobile internet, replacing text and images to a certain extent and gradually gaining a dominant position in text-based media such as news and social media platforms. This content is typically displayed in a feed format for users to quickly refresh. The feed consists of your friends or public figures you follow, and the content is their publicly posted updates. When you have a large and active group of friends, you receive a constant stream of updated content, a common form of feed. Time is the ultimate dimension of a feed, as content updates are the result of continuous requests to the server. The timeline is the most primitive, intuitive, and fundamental form of feed display, and building upon it allows for further design of feed flow patterns.
[0083] Short videos, also known as short clips, are a form of internet content dissemination, typically under 30 minutes long, distributed on new media platforms. With the increasing prevalence of mobile devices and faster internet speeds, short, fast-paced, high-volume content is gaining traction with major platforms, fans, and investors. Short videos are frequently broadcast on various new media platforms, suitable for viewing on the go and during short breaks. They range in length from a few seconds to several minutes. They can also incorporate topics such as skill sharing, humor and quirks, fashion trends, social issues, street interviews, public welfare education, advertising creativity, and commercial customization. Due to their short duration, they can be presented as standalone videos or as part of a series. Unlike micro-films and livestreams, short video production doesn't require the same specific forms of expression or team requirements. Instead, it boasts a simple production process, low barriers to entry, and high engagement. It also boasts greater communication value than livestreams. However, the ultra-short production cycle and engaging content present challenges for short video production teams' copywriting and planning skills. Excellent short video production teams typically rely on established self-media platforms or intellectual property (IP). These platforms boast not only a high-frequency and stable output of content, but also a strong fan base. The emergence of short videos has enriched the forms of native advertising in new media. From the initial stages of user-generated content (UGC), post-generated content (PGC), and user uploads, to specialized short video production agencies, MCNs (Multi-Channel Networks), and specialized short video apps, short videos have steadily risen, becoming a crucial communication tool for content startups and social media platforms. While short videos have sparked a frenzy among content entrepreneurs and impacted video media platforms, their influence has further expanded, leading to a fierce competition among major news platforms for short video content. Consequently, the variety and richness of short video content is increasing. Both producers and consumers of short video content have become a vast community.
[0084] MCN: It is a product form of multi-channel network, which combines PGC (Professional Generated Content) content and, with the strong support of capital, ensures the continuous output of content, thereby ultimately achieving stable commercial monetization.
[0085] PGC refers to professionally produced content (e.g., videos on video sites) and expert-produced content (e.g., content on social networks). It broadly refers to personalized content, diversified perspectives, and virtualized social relationships. It's also known as PPC (Professionally Produced Content).
[0086] Figure 2This is a flow chart of a method for obtaining an event cluster identifier provided in an embodiment of the present application. The execution subject of this method may be a server. Figure 2 As shown, the method includes the following steps 201 to 202.
[0087] Step 201: The server obtains the title feature vector of each information in the resource pool through the trained embedding network.
[0088] The embedding network is used to extract the title feature vector of the title of the input information. The server can input the title of each information in the resource pool into the embedding network, and extract the semantic features of the title of each information through the embedding network to obtain the title feature vector of each information. The embedding network is trained based on the first similarity between the anchor sample title and the positive sample title, and the second similarity between the anchor sample title and the negative sample title in at least two triples. Each triple includes an anchor sample title, a positive sample title and a negative sample title. The anchor sample title and the positive sample title belong to the same event, and the anchor sample title and the negative sample title belong to different events. In one possible implementation, the forms of each information in the resource pool include at least two.
[0089] The server can pre-construct at least two triples based on the sample data used for training, and perform similarity calculation based on each sample title included in the triple, so as to train the embedding network based on the calculated similarity. Figure 3 The steps shown illustrate the embedding network training method.
[0090] Figure 3 This is a flow chart of an embedded network training method provided in an embodiment of the present application. The execution subject of this method can be a server. Figure 3 As shown, the method includes the following steps.
[0091] Step 301: The server obtains a sample data set and a label of each sample in the sample data set.
[0092] Among them, each sample includes a baseline sample title and a candidate sample title, and the label of each sample includes a first event label and a second event label. The first event label indicates whether the candidate sample title and the baseline sample title belong to the same event, and the second event label indicates whether the text corresponding to the candidate sample title and the baseline sample title belongs to the same event.
[0093] In one possible implementation, the server may perform deduplication detection on the original information and obtain the reference sample title and candidate sample title, as well as the corresponding label, included in each sample based on the deduplication detection result.
[0094] Step 3011: The server performs duplicate detection on the original information set to obtain at least two information subsets.
[0095] Each information subset includes at least one information whose similarity is within a target threshold range. The server may include an original information set of the target application, and the original information set may be a collection of original information uploaded by an object in the target application.
[0096] In one example, the server may repeatedly check the titles of each piece of information in the original information set, and the title similarity between each piece of information in each information subset may be within a target threshold range. For example, the server may calculate, based on the keywords included in the title of each piece of information, at least one piece of information whose title similarity to the title of that information is within the target threshold range, and obtain an information subset corresponding to that piece of information, the information subset including multiple pieces of information whose similarity to the title of that information is within the target threshold range. The target threshold range may be configured as needed, for example, the target threshold range may be (70%, 90%), (50%, 80%), or (80%, 99%), etc.
[0097] Step 3012: For each information subset, the server uses the title of the target information in the information subset as the benchmark sample title, and obtains the candidate sample title corresponding to the benchmark sample title to obtain a sample.
[0098] Each sample may include a baseline sample title and multiple candidate sample titles corresponding to the baseline sample title. For each information subset, the server may obtain the title of the target information from the information subset as the baseline sample title and the event to which the target information belongs, and then obtain the candidate sample titles corresponding to the baseline sample title from the information subset.
[0099] In one possible implementation, the server can select the benchmark sample title according to certain rules. In one example, for each information subset, the server can obtain the event to which each information in the information subset belongs, and obtain at least one event; the server can count the number of information included in each event in the information subset; the server can use the event with the largest number of information as the target event, select the target information from the information included in the target event, and use the title of the target information as the benchmark sample title. Of course, the server can also use other methods to select the benchmark sample title and its corresponding event. For example, the server can select it by random sampling. The process may include: the server can also randomly select the title of any information from the information subset as the benchmark sample title, and obtain the event to which the any information belongs as the event corresponding to the benchmark sample title.
[0100] Regarding the method for selecting candidate sample titles, the server may select a first number of titles of information from the information subset to which the benchmark sample title belongs as candidate sample titles corresponding to the benchmark sample title. This first number can be configured as needed. For example, the first number can be 3, 5, 8, 10, etc. Alternatively, the first number can be any value within a preconfigured first range, such as (3, 10), meaning that the server can select 3 to 10 candidate sample titles. Furthermore, in one example, based on the event to which the benchmark sample title belongs, the server may select titles of information from the information subset that pertain to the same event as the benchmark sample title as candidate sample titles for the benchmark sample title. Of course, the server may also select candidate sample titles using other methods, such as random sampling. Accordingly, the process may include: the server randomly obtaining a first number of titles of information from the information subset to which the benchmark sample title belongs, and selecting the first number of titles of information as candidate sample titles corresponding to the benchmark sample title.
[0101] It should be noted that since the titles of each information in each information subset have a certain degree of similarity, there is a certain amount of information about the same event in each information subset, and the server has a greater chance of obtaining the benchmark sample titles and candidate sample titles of the same event from the information subset; this facilitates the subsequent construction of anchor sample titles and positive sample titles belonging to the same event in the triplet, thereby improving the efficiency of model training.
[0102] Step 3013: For each sample, the server determines the first event tag corresponding to the benchmark sample title and its corresponding candidate sample title based on the title keywords respectively included in the benchmark sample title and its corresponding candidate sample title.
[0103] The first event tag is used to indicate whether the benchmark sample title and the candidate sample title belong to the same event. Exemplarily, each sample may include a benchmark sample title and a first number of candidate sample titles. The tag for each sample may include a first number of first event tags corresponding to the first number of candidate sample titles, each first event tag being used to indicate whether the corresponding candidate sample title and the benchmark sample title belong to the same event.
[0104] As for the method of determining whether the candidate sample title and the benchmark sample title are about the same event, the server can determine whether the two titles belong to the same event based on multiple angles, such as the main events of the two titles, the descriptions of the events in the two titles, the objects listed in the main events of the titles, and the types of keywords included in the titles.
[0105] In Example 1, the server can make a judgment based on the subject events of the two titles. For example, if the base sample title and the candidate sample title correspond to the same subject event, the server determines that the candidate sample title and the base sample title belong to the same event. Two titles describing the same subject event can be considered to belong to the same event even if they have different perspectives on the event.
[0106] For example 1, Title 1 is "xx study says that floor height affects life expectancy! How to choose a floor?", and Title 2 is "Floor height affects life expectancy". The main events of both Title 1 and Title 2 are "Floor height affects life expectancy", so Title 1 and Title 2 can be considered as the same event.
[0107] In Example 1, if the subject matter of the two titles is the same, even if the viewpoints of the description of the subject matter are different, such as if the title contains doubts, denials, affirmations, refutations, or rumors about the subject matter, they can still be considered titles about the same event.
[0108] Example 2: Title 1 is "Is it really not possible to use the air conditioner during the summer confinement period?", and Title 2 is "Don't disbelieve it, you can use the air conditioner during the summer confinement period." The main events of both Title 1 and Title 2 are "Using the air conditioner during the summer confinement period." Although Title 1 is in an interrogative tone and Title 2 is in an affirmative tone, they can be regarded as the same event.
[0109] Example 3: Title 1 is an entertainment news title "K's acting skills are unexpectedly well received? The "xxx" program takes you to know the children from the stars!", and Title 2 is an entertainment news title "The audience cried so loudly! The "xxx" program made K's acting skills a hot topic." The main events of both Title 1 and Title 2 are "The "xxx" program, K's acting skills." Although one of Title 1 and Title 2 received positive reviews and the other was on the hot topic, "The "xxx" program, K's acting skills" is sufficient to determine that the two titles refer to the same event, and therefore can be regarded as the same event.
[0110] In Example 2, the server can make a judgment based on the relationship between the main events of the two titles. For example, when the main events corresponding to the benchmark sample title and the candidate sample title are included in a relationship, the server determines that the candidate sample title and the benchmark sample title belong to the same event. For example, if the overlapping main events of the two titles can form a complete event, they can be considered titles belonging to the same event.
[0111] Example 4: Title 1 is "After B made C cry, he was praised by D. Netizens: A talented actor who was delayed by his humor", and Title 2 is "Behind making C cry, it is not only the acting skills but also the efforts of B". Title 1 has "praised by D" more than Title 2, but the repeated part of the main event of the two titles "After B made C cry" can constitute a complete event. Therefore, Title 1 and Title 2 can be considered as the same event.
[0112] In Example 3, the server can make a judgment based on the objects in the main events described by the two titles. For example, if the objects in the main events corresponding to the benchmark sample title and the candidate sample title are the same, the server determines that the candidate sample title and the benchmark sample title belong to the same event. For example, if the inventory objects in the main events of the two titles are the same, then the titles can be considered to belong to the same event.
[0113] Example 5: Title 1 is "Three tips for drinking a thousand cups without getting drunk, have you learned them?", and Title 2 is "5 tips for drinking without harming your body, drink a thousand cups without getting drunk, the more you drink the more sober you will be, prepare for the New Year!". The main events of Title 1 and Title 2 are both reviewing "Tips for drinking a thousand cups without getting drunk". Since the review objects are the same, Title 1 and Title 2 can be considered to be the same event.
[0114] In Example 4, the server can make a judgment based on the keyword types included in the two titles. For example, when the benchmark sample title and the candidate sample title include the same keyword type and the content of the same keyword type is also the same, the server determines that the candidate sample title and the benchmark sample title belong to the same event. For example, keyword types can include but are not limited to numerical values, locations, time and date, etc.; if the content of the same keyword type in the two titles is completely inconsistent and is a key attribute of the event, then the events are different.
[0115] Example 6: Only one of the titles "The latest statistics of the number of people taking the XXX exam" and "The latest statistics of the number of people taking the XXX exam in 2018" contains "2018", so these two titles belong to the same event.
[0116] Example 7: The titles "The latest statistics of the number of people taking the XXX exam in 2019" and "The latest statistics of the number of people taking the XXX exam in 2018" both have inconsistent years, one is 2018 and the other is 2019. Therefore, these two titles do not belong to the same event.
[0117] Example 8: The title "The latest statistics of the number of people taking the XXX exam is more than in 2017" and the title "The latest statistics of the number of people taking the XXX exam in 2018". Both titles have years. Although they are inconsistent, 2017 in the first title is not closely related to the main event, so Title 1 and Title 2 can belong to the same event.
[0118] In Example 5, the server can make a judgment based on the subject details of the subject events described by the two titles. For example, when the subject details of the subject events described by the benchmark sample title and the candidate sample title are the same, the server determines that the candidate sample title and the benchmark sample title belong to the same event.
[0119] Example 9: Title 1 is "The TV series "XXX" was launched in a low-key manner, and characters A and B in the series drive the direction of the plot"; Title 2 is "Hermit actress H from the TV series "XXX" on the xx film and television platform appeared at the xxx party." The events described in Title 1 and Title 2 are different. Therefore, Title 1 and Title 2 cannot be attributed to the same event.
[0120] Furthermore, if the subject event described in a title is unclear and cannot be confirmed as the same event, it can be marked as "Unclear Event Description." For example, Title 1 reads, "Starting in April, three good things will happen in the XX region. People in the XX region should be aware of this in advance so they can prepare!", while Title 2 reads, "Friends in the XX region, please note: Two good things will happen in May. Every household will get the first one, so please check." Neither title clearly states what the good things are. Therefore, Title 1 and Title 2 should not be attributed to the same event. It is important to note that the corresponding content of Title 1 and Title 2 may refer to the same event.
[0121] Step 3014: The server determines the second event tags corresponding to the benchmark sample title and its corresponding candidate sample title based on the text corresponding to each of the benchmark sample title and its corresponding candidate sample title.
[0122] The server can determine whether the samples of the benchmark sample title and the candidate sample title belong to the same event based on the text content corresponding to each of the benchmark sample title and the candidate sample title. The text corresponding to the title can include one or more forms of information such as articles, images, videos, and audio; of course, the server can determine whether the two titles belong to the same event based on multiple angles such as the main events described in the texts of the two titles, the descriptive viewpoints of the events described in the texts of the two titles, the objects listed in the main events described in the texts of the titles, and the types of keywords included in the texts of the titles. This judgment method can be similar to the process of step 3013 above, and no further examples will be given here.
[0123] Step 302: The server constructs the at least two triples based on the label of each sample in the sample data set.
[0124] Based on the label of each sample, the server can select two titles belonging to the same event in the sample dataset as a title pair, and select a title that is different from the event to which the title pair belongs as the negative sample title corresponding to the title pair. Based on the title pair and the negative sample title corresponding to the title pair, the server constructs a triplet. The server constructs multiple triplets in this manner. A title pair can include an anchor sample title and a positive sample title.
[0125] In one possible implementation, for each sample's reference sample title and candidate sample title, the server can treat the reference sample title and candidate sample title belonging to the same event as a title pair in a triplet. Then, a negative sample title is selected from the sample dataset, which belongs to a different event than the anchor sample title, to obtain a triplet. For each sample, the server can determine whether the candidate sample title and the reference sample title in the sample belong to the same event based on the sample's first event label and second event label. In one example, the server can make a judgment based on either the first event label or the second event label. For example, when either the first event label or the second event label indicates that the candidate sample title and the reference sample title are the same event, the server treats the candidate sample title and the reference sample title as a title pair belonging to the same event. For example, if the titles of candidate title 1 and reference title 2 do not belong to the same event, but the text does, the server can also treat candidate title 1 and reference title 2 as a title pair, such as using reference title 2 as the anchor sample title and candidate title 1 as the positive sample title. In another example, the server may also use other methods based on the first event tag and the second event tag to determine whether two tags are the same event. For example, the server may further restrict the conditions for determining whether two titles are the same event. This process may include: when both the first event tag and the second event tag indicate that the candidate sample title and the benchmark sample title are the same event, the server may treat the candidate sample title and the benchmark sample title as a title pair belonging to the same event.
[0126] In one example, a triple can be represented as<a,p,n> , where a is the anchor and represents the anchor sample title; p (positive, representing the positive sample title) represents the title that is close to or similar to the anchor in the feature space. In the embodiment of the present application, p and a are sample titles belonging to the same event; n (negative, representing the negative sample title) represents the sample that is far away from the anchor in the feature space. In the embodiment of the present application, n and a are sample titles that do not belong to the same event; that is, p and n are positive samples and negative samples compared with the anchor, respectively.
[0127] In this step, the server can construct triples based on samples in the sample dataset using a random sampling method. Alternatively, the server can construct multiple types of triples based on the first and second distances between the anchor sample title and the positive and negative sample titles, respectively, with each type of triple having a different range of first and second distances. Accordingly, step 302 can include the following methods 1 and 2.
[0128] Method 1: The server obtains at least two title pairs based on the label of each sample in the sample data set, and constructs three types of triples based on the first distance between the anchor sample title and the positive sample title in each title pair, and the second distance between the anchor sample title and at least one alternative title.
[0129] The three types of triples include: first-class triples, second-class triples, and third-class triples; wherein the difference between the second distance corresponding to the first-class triple and the first distance is greater than the target value, the first distance corresponding to the second-class triple is less than the second distance, and the sum of the first distance and the target value is greater than the second distance, and the second distance corresponding to the third-class triple is less than the first distance. In one example, the server can first calculate the first distance and the second distance, and then further filter out triples of each type based on the first distance and the second distance. Accordingly, step 302 in method one can include the following steps 3021a to 3023a:
[0130] Step 3021a: The server may obtain at least two title pairs from the sample dataset based on the label of each sample in the sample dataset.
[0131] Each title pair includes an anchor sample title and a positive sample title belonging to the same event. For each sample, the server can use the benchmark sample title and the candidate sample title belonging to the same event as a title pair in a triplet based on the label of the sample.
[0132] Step 3022a. For each title pair, the server obtains a first distance between the anchor sample title and the positive sample title based on the feature vectors of the anchor sample title and the positive sample title in the title pair, and obtains a second distance between at least one alternative sample title and the anchor sample title based on the feature vectors of at least one alternative sample title and the anchor sample title.
[0133] In one example, the server may calculate a first distance between the feature vector of the anchor sample title and the feature vector of the positive sample title using a method such as Hamming distance or cosine distance; and calculate at least one second distance between the feature vector of the anchor sample title and the feature vector of at least one candidate sample title. Before obtaining the first and second distances, the server may extract feature vectors for each title using an initial embedding network. This process may include: the server inputting the anchor sample title, the positive sample title, and at least one candidate sample title in the sample dataset into the initial embedding network, and outputting, through the initial embedding network, a feature vector for each anchor sample title, a feature vector for the positive sample title, and a feature vector for each candidate sample title.
[0134] Step 3023a: The server determines first-category triples, second-category triples, and third-category triples based on the first distance and the second distance corresponding to each title pair.
[0135] The difference between the second distance and the first distance corresponding to the first type triplet is greater than the target value, the first distance corresponding to the second type triplet is less than the second distance, and the sum of the first distance and the target value is greater than the second distance, and the second distance corresponding to the third type triplet is less than the first distance.
[0136] In one example, the server can pre-configure the proportion of each type of triple, for example, the first type of triple accounts for 50%, the second type of triple, and the third type of triple each accounts for 25%. For example, the server can obtain a first proportion of title pairs, and obtain the second distance corresponding to each title pair in the first proportion of title pairs, to obtain the first type of triples of the first proportion. The difference between the second distance corresponding to each title pair in the first proportion of title pairs and the first distance of the title pair is greater than the target value. Similarly, the server obtains a second proportion of title pairs, and obtains the second distance corresponding to each title pair of the second proportion, to obtain the second type of triples of the second proportion. The server obtains a third proportion of title pairs, and obtains the second distance corresponding to each title pair of the third proportion, to obtain the third type of triples of the third proportion.
[0137] In an example, the first type of triplet, the second type of triplet, and the third type of triplet can be represented as easytriplet (triplet that is easy to distinguish), semi-hard triplet (triplet of intermediate difficulty), and hard triplet (triplet that is difficult to distinguish). The following is an introduction to each type of triplet:
[0138] easy triplet: d(a,p) + margin < d(a, n), where d(a,p) represents the first distance between the anchor sample title a and the positive sample title p, d(a, n) represents the second distance between the anchor sample title a and the negative sample title n, and the value of margin (margin parameter) can be the target value. At this time, the difference between the second distance and the first distance is greater than margin.
[0139] semi-hard triplet: Samples that are semi-easy to distinguish, d(a,p) + margin > d(a, n) > d(a,p). At this time, the difference between the second distance and the first distance is within margin.
[0140] hard triplet: Samples that are difficult to distinguish, d(a, p) > d(a, n).
[0141] such as Figure 4 shown, the relationship between each type of triplet and margin is as Figure 4 shown, Figure 4 In the figure, the center position represents the anchor. The positive sample title p can be located on the contour line of the small circle with a smaller radius centered on the anchor. Among them, the second distance between the negatives in the easy triplet and the anchor is larger, and the difference between the second distance minus the first distance exceeds margin. That is, the negative sample title negatives is located outside the contour line of the large circle with a larger radius centered on the anchor; while the difference between the second distance and the first distance of the negatives in the semi-hard triplet is in the middle position. That is, the negative sample title negatives is located between the large circle and the small circle centered on the anchor; and the negative sample title negatives in the hard triplet is located inside the small circle centered on the anchor, and the second distance between the negatives and the anchor is the smallest, and the second distance is instead smaller than the first distance.
[0142] In this step, by calculating the first distance between the anchor sample title and the positive sample title in the title pair, and the second distance between the anchor sample title and the alternative sample title, various types of triplets are constructed, so that the network training process can learn more and richer information during the training process using various types of triplets, thereby improving the effect of machine learning, and finally training a more accurate embedding network.
[0143] In one possible implementation, the server may construct the triples used in each training iteration from the batch of samples used in the current training iteration. In one example, the step of the server obtaining at least two title pairs from the sample dataset based on the label of each sample in the sample dataset may include: for each training iteration, the server obtaining a batch of sample data used in the current training iteration from the sample dataset; and obtaining the at least two title pairs from the batch of sample data based on the label of each sample in the batch of sample data.
[0144] In one example, if the triples are constructed in real time during each iterative training using the above method, the server can obtain the at least one candidate sample title from the batch of sample data before obtaining the second distance. Of course, the server can also extract the feature vector of the at least one candidate sample title through the initial embedding network.
[0145] Method 2: The server obtains at least two title pairs based on the label of each sample in the sample data set, and randomly obtains the negative sample title corresponding to each title pair to obtain at least two triplets.
[0146] In this step, the server obtains a title pair based on the label of each sample in the sample dataset, and the title pair includes an anchor sample title and a positive sample title belonging to the same event; for each title pair, the server randomly samples a target number of sample titles as the negative sample titles corresponding to the title pair; the server obtains the at least two triples based on each title pair and its corresponding negative sample title. For example, for each title pair<a,p> , the server can randomly sample K titles as the<a,p> For example, K can be configured based on needs, and this application does not limit this. For example, K can be set to 3, 4, 10, etc.
[0147] The present application utilizes the labels of pre-labeled samples to iteratively train the initial embedding network. The iterative training process may include the following steps 303 and 304 .
[0148] Step 303: The server determines a first similarity between the anchor sample title and the positive sample title, and a second similarity between the anchor sample title and the negative sample title based on the respective feature vectors of the anchor sample title, the positive sample title, and the negative sample title in each triplet.
[0149] Among them, the respective feature vectors of the anchor sample title, the positive sample title and the negative sample title are obtained by extracting features of the anchor sample title, the positive sample title and the negative sample title respectively through the initial embedding network. The server can use the vector distance between the feature vectors to measure the degree of difference between the two titles. The larger the vector distance, the greater the difference between the two titles and the smaller the similarity. For example, for each triple, the server can calculate the first vector distance between the anchor sample title and the positive sample title based on the feature vector of the anchor sample title and the feature vector of the positive sample title in the triple; and the server can calculate the second vector distance between the anchor sample title and the positive sample title based on the feature vector of the anchor sample title and the feature vector of the negative sample title in the triple.
[0150] The larger the first vector distance, the greater the difference between the anchor sample title and the positive sample title, indicating a smaller first similarity. The larger the second vector distance, the greater the difference between the anchor sample title and the positive sample title, indicating a smaller second similarity.
[0151] Step 304: The server trains the initial embedding network based on the difference between the first similarity and the second similarity to obtain the embedding network.
[0152] The server may determine a loss value based on the first similarity and the second similarity, and iteratively train the initial embedding network based on the loss value until the training is stopped when a target condition is met, thereby obtaining the embedding network.
[0153] In one possible implementation, the server may measure the difference between the two similarities based on a target value, and the target condition may include that the second similarity is higher than the target value of the first similarity. Accordingly, the iterative training process may include: when the difference between the second similarity and the first similarity is not higher than the target value, the server iteratively trains the initial embedding network until the second similarity is higher than the target value of the first similarity, and stops training to obtain the embedding model. The computer device may repeatedly perform the above steps 302-304, and optimize the model parameters in the initial embedding network based on the difference between the first similarity and the second similarity of each repeated process, and stops iteration when the second similarity is higher than the target value of the first similarity, and obtains the embedding model.
[0154] In one possible example, the computer device may determine the loss value of this iterative training based on the feature vector of the anchor sample title, the feature vector of the positive sample title, and the feature vector of the negative sample title using the triple loss function shown in Formula 1 below. When the loss value reaches a target condition, the iteration is stopped to obtain the embedding model:
[0155] Formula 1: Triplet loss ;
[0156] Among them, triplet loss is the loss value of the triplet loss function, and d(x, y) represents the distance between x and y, that is, represents the first distance between the anchor sample title a and the positive sample title p. A larger first distance indicates a smaller first similarity. d(a, n) represents the second distance between the anchor sample title a and the negative positive sample title n. A larger second distance indicates a smaller second similarity. For example, this distance can be represented by Hamming distance, cosine distance, or Manhattan distance, and this is not limited in this application. margin is an interval value and a hyperparameter. The value of margin can be a target value.
[0157] It should be noted that the initial embedding network can be a network with feature extraction capabilities. In one example, the initial embedding network can be a network structure based on the Bert model. BERT is a deep bidirectional language representation model based on Transformer. In essence, it uses the Transformer structure to construct a multi-layer bidirectional Encoder network. This application uses the BERT model to extract semantic features of the title. The full name of BERT is Bidirectional Encoder Representation from Transformers, and the core of BERT is the bidirectional Transformer Encoder. In one example, the server can use the pre-trained Bert network and the variable layer as the initial embedding network.
[0158] In one example, the embedding network can be a network layer for feature extraction in a trained target model. For example, the server can input the sample data set into the initial model, use the initial embedding network in the initial model to extract the feature vector of each sample, and output the event judgment result of each sample based on the extracted feature vector. The event judgment result may include a first judgment result and a second judgment result. The first judgment result is used to indicate whether the candidate sample title and the benchmark sample title belong to the same event, and the second judgment result indicates whether the text corresponding to the candidate sample title and the benchmark sample title belongs to the same event. The server can train the initial model based on the event judgment results of each sample and the label of the sample, and construct and train triples based on the above steps 302-304 to obtain a target model. The server can extract the trained embedding network in the target model.
[0159] For example, the server can fix the first (n-1) layers of the initial embedding network consisting of n layers, and set the last layer of the initial embedding network and the FC (Fully Connected Layers) as optimizable network layers, and adjust the network parameters of the last layer and the FC layer during iterative training. Of course, the server can also use other networks with feature extraction functions as the initial network for training to obtain the final feature extraction network. This application only uses the above-mentioned initial embedding network as an example for illustration, but does not limit the network structure, network type, etc. adopted.
[0160] It should be noted that, through the process of steps 302-304 above, the positive sample title and negative sample title corresponding to the anchor sample title are determined, and the difference between the two similarities between the anchor sample title and the positive and negative sample titles is further used to iteratively train the initial embedding network, so that the initial embedding network can perform comparative learning through positive and negative sample pairs, optimize the expression of the network's feature vector, thereby shortening the distance between the feature vector of the anchor sample title and the feature vector of the positive sample title in the feature space, and at the same time widening the distance between the feature vector of the anchor sample title and the feature vector of the negative sample title. The pre-trained Bert network can well extract the semantic features of the title, and through the above iterative training process, it can make the features of the positive sample closer to the features of the anchor point, while making the features of the negative sample farther away from the features of the anchor point in the feature space. Therefore, the network training method of steps 301-304 above is used in this application to enable the trained embedding network to achieve the purpose of extracting features that better characterize similar titles, thereby improving the accuracy of the feature vectors extracted by the trained embedding network.
[0161] like Figure 5 As shown, the initial embedding network includes a network layer for extracting feature vectors of anchor sample titles, a network layer for extracting feature vectors of positive sample titles, and a network layer for extracting feature vectors of negative sample titles. These network layers can share network parameters (share weight). The server can directly input the triples into the initial embedding network, extract the feature vectors of each sample title in the triples through the initial embedding network, and calculate the loss value using the triple loss function shown in Formula 1 above. Based on this, the initial embedding network is iteratively trained to optimize the network parameters of the initial embedding network, thereby obtaining the embedding network.
[0162] Step 202: The server performs vector clustering on the title feature vectors of each information in the resource pool to obtain at least two event clusters, and marks the information included in each event cluster with an event cluster identifier corresponding to the event cluster.
[0163] Each event cluster includes at least one information belonging to the same event.
[0164] The server can Figure 3 After extracting the title feature vectors of each message in the resource pool using the embedding network trained using the method, the server can cluster the title feature vectors of each message to obtain at least two event clusters. For each event cluster, the server labels the information included in the event cluster with an event cluster identifier. Clustering refers to the process of dividing a collection of physical or abstract objects into multiple classes consisting of similar objects. A cluster generated by clustering is a collection of data objects. Objects within the same cluster are similar to each other and different from objects in other clusters.
[0165] The present application can use a hierarchical clustering method to cluster the title feature vectors of each message. In one example, the server can first classify the title feature vectors of all messages in the resource pool into a cluster, and then gradually split the cluster according to pre-configured rules until the splitting end condition is met, at which point the splitting stops and at least two event clusters are obtained.
[0166] In one possible implementation, each information in the resource pool includes at least two forms. By clustering information of at least two forms, the information in the resource pool can be divided into cross-form events, and information of different forms can be marked with the event cluster identifier to which they belong. This facilitates the deletion and filtering of similar information of the same event in different forms in subsequent recommendation lists, thereby improving the form richness of information in the final recommended information flow and thereby improving the actual recommendation efficiency.
[0167] In one possible executable algorithm, the clustering process can be expressed as follows:
[0168] Input: The original information set D to be clustered, and the pre-configured splitting end condition (the splitting end condition may include but is not limited to: the pre-configured number of target categories, reaching the sample clustering threshold, etc.). Here, the sample distance threshold is used as an example.
[0169] Output: Clustering results
[0170] 1. Classify the title feature vectors of all information in the information set D into a cluster;
[0171] repeat (repeat the following steps):
[0172] 2. Calculate the distance between the title feature vectors of each message in the same cluster (denoted as c) and find the title feature vectors a and b of the two messages with the farthest distance.
[0173] 3. Assign the title feature vectors a and b of the information to different clusters c1 and c2;
[0174] 4. Calculate the distances between the title feature vectors of the remaining other information in the original cluster (c) and the title feature vectors a and b of the information respectively. If dis(a) < dis(b), then classify the title feature vectors of the other information into c1; otherwise, classify them into c2. Here, dis(a) represents the distance between the title feature vector of the other information and the title feature vector a of the information, and dis(b) represents the distance between the title feature vector of the other information and the title feature vector b of the information;
[0175] Until: Reach the pre-configured splitting end condition, such as stopping clustering when the distance is less than a certain fixed threshold.
[0176] The server can perform vector clustering on the title feature vectors of all enabled information in the resource pool based on the above executable algorithm to obtain at least two event clusters. It should be noted that the distance here is calculated between the title feature vectors of two pieces of information through the title feature vectors of each piece of information, so as to aggregate the graphic, text, and video content of similar titles together. By learning the representation of the title feature vectors of each title through the above embedding network, the title feature vectors of each piece of information can be differentially expressed based on the events to which they belong. Then, based on the above step 202, event clustering is performed to obtain the IDs of the event clusters, and each piece of information in the resource pool is marked with the ID of the event cluster to which it belongs. Finally, when generating a candidate recommendation list during recommendation recall, the information of the same event in the candidate recommendation list is scattered into different recommendation lists, thereby realizing the repeated scattering distribution of the information flow content, which can largely avoid the problem of concentrated appearance of similar titles in the same recommendation list, bring a better content consumption and reading experience, and ensure a higher level of richness and diversity of the recommended information flow for readers.
[0177] The following provides an information recommendation method, which adopts the above Figure 2 shown method for obtaining event cluster identifiers to obtain the event cluster identifiers of each piece of information in the resource pool, and perform information recommendation based on the event cluster identifiers of each piece of information.
[0178] Figure 6 FIG. is a signaling interaction diagram of an information recommendation method provided by an embodiment of the present application. This method can be implemented through the interaction between a server and a terminal, and the terminal can be the terminal of any object. As Figure 6 shown, the method includes the following steps.
[0179] Step 601: The terminal displays an information display page of the target application.
[0180] The target application may be an application with an information recommendation function. The information display page is used to display a recommended information stream, and the form of each information in the information stream may include but is not limited to: text, image, video, short video, audio, column, etc.
[0181] It should be noted that the information flow in the information display page can be displayed in the form of Feeds, where, for example Figure 7 Several feed-style display pages are available. Feeds, also known as source material, feed, information provider, feed, summary, source, news subscription, or web source, are a data format through which websites communicate the latest information to users, typically arranged in a timeline. The timeline is the most primitive, intuitive, and basic form of feed display. A prerequisite for users to subscribe to a website is that the website provides a source.
[0182] Step 602: The terminal sends a recommendation request to the server based on the target operation in the information display page.
[0183] The target operation is used to trigger the information recommendation function of the target application. For example, the target operation can be a pull-down operation, a downward sliding operation, a double-click operation, a refresh operation, etc.
[0184] Step 603: The server receives a recommendation request for any object.
[0185] The terminal may be the terminal where any object is located.
[0186] Step 604: In response to a recommendation request for any object, the server deletes information that meets the event cluster similarity condition from the candidate recommendation list based on the event cluster identifiers of at least two pieces of information to obtain a recommendation list.
[0187] In which, the event cluster similarity condition includes the same event cluster identifier as any other information in the candidate recommendation list. The event cluster identifier is obtained based on the method of steps 201-202 above. In one example, the server can also complete the candidate recommendation list after deletion. The process may include: the server adds new information to the candidate recommendation list based on the event cluster identifier of the remaining information in the candidate recommendation list to obtain the recommendation list. In which, the event cluster identifier of the added new information is different from the remaining information, and the remaining information is the information in the candidate recommendation list other than the first information. In one example, the candidate recommendation list is a list generated based on a preconfigured recommendation strategy, and the method in the present application can be used to filter it to increase the diversity of information in the final recommended list. In one example, the corresponding information in the recommendation list includes at least two forms. For example, the corresponding information in the list includes graphic information, video information, audio information and other forms of information.
[0188] By identifying an event cluster based on at least two pieces of information, the information in the candidate recommendation list that meets the similarity conditions of the event cluster is deleted, thereby filtering the candidate recommendation list for similar information of the same event, and minimizing similar information of the same event in the recommendation list. In addition, at least two forms of information in the resource pool are clustered, so that the recommendation list can support the filtering and deletion of information of multiple forms for the same event, thereby distinguishing and quantifying the similarity of the content presentation level across forms, improving the diversity of information forms in the recommendation list, expanding the applicable scenarios of the present application, and improving the applicability of the recommendation method. In addition, by clustering the title feature vectors extracted based on the embedded network, the same event information in the candidate recommendation list is scattered by the clustered event clusters, which can improve the reading experience; especially for the intensive exposure of the information flow of the current hot events, it can effectively change the consumption motivation of readers to read information and increase the consumption time of readers. On the basis of deduplication of recommended information, it can distinguish and measure similar content of the same event in a more refined granularity from the perspective of the event to which it belongs, which can better meet a variety of consumer needs.
[0189] Step 605: The server recommends the information flow to be recommended corresponding to the recommendation list to the any object.
[0190] In one example, the server may return a recommendation list to the any object, and the recommendation list may include information identifiers of each information in the information flow to be recommended.
[0191] Step 606: The terminal receives the recommendation list and displays the to-be-recommended information flow corresponding to the recommendation list on the information display page.
[0192] The terminal can receive the recommendation list, and based on the information identifier of at least one information included in the recommendation list, obtain each information in the information flow to be recommended from the resource server storing the resource data of each information, and display each information in the information flow to be recommended on the information display page.
[0193] The information recommendation method provided by the present application obtains a recommendation list by deleting information that meets the event cluster similarity conditions from a candidate recommendation list based on event cluster identifiers of at least two pieces of information, and then makes recommendations. This method filters similar information belonging to the same event cluster through the event cluster identifier, preventing similar information from being concentrated in the same recommendation list, thereby improving recommendation efficiency. Furthermore, the method obtains the title feature vectors of each piece of information through an embedded network, performs vector clustering on the title feature vectors of each piece of information, and obtains at least two event clusters, each of which includes at least one piece of information belonging to the same event. This method divides each piece of information into different event clusters, allowing the use of the event cluster to quickly filter and disperse similar information of the same event in the candidate recommendation list, thereby improving the richness and diversity of the recommendation list and, in turn, improving actual recommendation efficiency.
[0194] Moreover, the embedding network is trained based on the first similarity between the anchor sample titles and the positive sample titles, and the second similarity between the anchor sample titles and the negative sample titles in at least two triples; while the anchor sample titles and the positive sample titles belong to the same event, and the anchor sample titles and the negative sample titles belong to different events. The present application uses the embedding network to perform feature extraction, thereby achieving the effect of effectively bringing the title features of the same event closer and the distinctive features of different events farther apart in the feature space, thereby improving the differentiated expression of the title feature vector for different events, thereby improving the accuracy of clustering, and further improving the accuracy of information recommendation.
[0195] Figure 8 A framework diagram of an information recommendation method provided in an embodiment of the present application, such as Figure 8 As shown, the modules and services in the information recommendation framework may include: 1. Content production and consumption end, 2. Uplink and downlink content interface server, 3. Content database, 4. Scheduling center & manual review system, 5. Content storage service, 6. Title same event sample library and title same event model, 7. Title event vector generation service, 8. Title vector library, 9. Title vector clustering service.
[0196] Below is Figure 8 The framework diagram shown here introduces the modules and related services in the framework of the information recommendation method of this application:
[0197] 1. Content production and consumption
[0198] (1) GC or UGC, MCN content producers provide publishing portals for video content and graphic content through mobile terminals or back-end API (Application Programming Interface) systems. The multimedia content obtained by the above publishing portals is the main content source of information flow content services.
[0199] (2) The content production end uploads graphic and text content through communication with the upstream and downstream content interface services. The source of graphic and text content is usually a lightweight publishing end and editing content portal. Video content publishing is usually a shooting and photography end. During the shooting process, the local video content can choose matching music, filter templates and video beautification functions, etc.
[0200] (3) The content consumer communicates with the upstream and downstream content interface servers to obtain content index information, and then obtains the content source file from the content storage service based on the content index information, and then loads the content source file to display it to the client. The above index information can be the index information of the content subscribed by the client. The content storage server stores content entities such as video source files and cover image source files. The content metadata, such as title, author, cover image, category, tag information, etc., is stored in the content database.
[0201] (4) The content production end and the content consumption end simultaneously report the user playback behavior data, freezes, loading time, playback clicks, etc. during the upload and download process to the upstream and downstream content interface servers or other background servers for subsequent data statistical analysis.
[0202] (5) The content consumption end usually displays content data to the object through the feed stream. The object continuously refreshes by sliding to obtain more recommended information, similar to a waterfall flow.
[0203] 2. Uplink and Downlink Content Interface Server
[0204] (1) Communicate directly with the content production end to obtain content submitted by the content production end. Usually, the content metadata such as the title, publisher, summary, cover image, and release time.
[0205] (2) Write content metadata into the content database, such as file size, cover image link, title, release time, author, and other information into the content database.
[0206] (3) Synchronize the content published and submitted by the content production end to the dispatch center server (which can be referred to as content entering the dispatch center) so that the dispatch center server can perform subsequent content processing and circulation.
[0207] 3. Content Database
[0208] (1) The content database is the core database of content. The metadata of all content published by all content production ends is stored in this content database. The focus is on storing metadata of the content itself, such as file size, cover image link, bit rate, file format, title, release time, author, video file size, video format, original mark, first release mark, and classification label information of the content during the manual review process. The above classification label information includes first-level, second-level, and third-level classification and label information. For example, a video explaining a certain brand of mobile phone has a first-level classification of technology, a second-level classification of smart phones, and a third-level classification of domestic mobile phones. The label information is a certain brand and model.
[0209] (2) The manual review system will read the information in the content database during the manual review process (in Figure 10 At the same time, the manual review results and status will also be sent back by the manual review system to the content database.
[0210] (3) The content processing by the dispatch center server mainly includes machine processing and manual review. The core process of machine processing includes various methods for judging video quality, such as filtering low-quality videos and adding content tags, such as adding video classification information and tag information. In addition, it also includes content similarity screening, the results of which are written to the content database. Completely duplicate content will not be repeatedly processed by humans, saving human resources for review.
[0211] (4) When recommending subsequent information, the meta-information of the content can be read from the content database. In addition, when extracting the title feature vector of each information title, the title of the text and video content that has been enabled in the content database can be read and the title feature vector of each information title in the resource pool can be output through the trained embedding network.
[0212] 4. Dispatch Center and Manual Review System
[0213] (1) The dispatch center server is responsible for the entire dispatch process of content flow, receives the content that needs to be stored through the upstream and downstream content interface servers, and then obtains the metadata of the content from the content database.
[0214] (2) Scheduling manual review system and precision verification service to control the order and priority of scheduling.
[0215] (3) The dispatch center starts distributing the content and provides it directly to the content consumers at the terminal through the content distribution service (usually a recommendation engine, search engine, or operation). That is, the content consumer obtains content index information, such as URL address information.
[0216] (4) The manual review system is the carrier of manual service capabilities. It is mainly used to review and filter some sensitive content that needs to be filtered. It can also label and re-confirm information.
[0217] 5. Content Storage Service
[0218] (1) Content entity information other than the metadata of the stored content, such as video source files and image source files of graphic content.
[0219] (2) By communicating with the upstream and downstream interface servers to store source files, the content source files are provided to the content consumption end, and the download file system can download video files from the content storage service.
[0220] 6. Sample Library and Embedding Model for Events with the Same Title
[0221] (1) According to the sample dataset acquisition method in this application, a sample dataset including multiple samples is obtained, and the label of each sample is obtained using the baseline sample title and candidate sample title of each sample to label each sample. Samples labeled as the same event can also be stored in a sample library with the same title and event.
[0222] (2) According to the training method described above, construct triples, calculate the first similarity and the second similarity, and further use triplet loss to train the embedding model.
[0223] 7. Title Event Vector Generation Service
[0224] (1) According to the embedding network training method of the present application, the trained embedding network is used as the title feature vector extraction network for all information to obtain the basic model of the title vector of the same event. In this way, the semantic information of the title can be extracted well, and the vector distance of the same event can be shortened in the feature space, and the distance vectors of different events can be separated.
[0225] (2) The embedding network trained by this application is used to vectorize the titles of all enabled contents in the resource pool, and the processing results are stored in the title vector library.
[0226] 8. Title Vector Library
[0227] (1) According to the embedded network trained in this application, extract and save the feature vectors of all enabled content titles in the resource pool, including graphic, text and video content.
[0228] (2) It can communicate with the title vector clustering service and provide all enabled content in the resource pool of each information used by clustering.
[0229] 9. Title Vector Clustering Service
[0230] (1) Using the hierarchical clustering method, vector clustering is performed based on the feature vectors of all enabled content titles in the resource pool to obtain the ID of each event cluster, which can eventually be used to delete and disperse the same event information in the candidate recommendation list at the recommendation recall layer, thereby achieving the repeated distribution of information flow content.
[0231] (2) Read the same event vectors of titles in the title vector library and cluster them.
[0232] Figure 9 This is a schematic diagram of the structure of an information recommendation device provided in an embodiment of the present application. Figure 9 As shown, the device includes:
[0233] A recommendation list determination module 901 is configured to, in response to a recommendation request for any object, delete information that meets the event cluster similarity condition from a candidate recommendation list based on event cluster identifiers of at least two pieces of information to obtain a recommendation list;
[0234] A recommendation module 902 is configured to recommend the information flow to be recommended corresponding to the recommendation list to the any object;
[0235] The similarity condition of the event cluster includes being the same as the event cluster identifier of any other information in the candidate recommendation list;
[0236] The device is also used to obtain the event cluster identifier. When obtaining the event cluster identifier, the device further includes:
[0237] The title feature vector acquisition module 903 is used to acquire the title feature vector of each information in the resource pool through the trained embedding network;
[0238] Clustering module 904 is configured to perform vector clustering on the title feature vectors of each message in the resource pool to obtain at least two event clusters, and mark the message included in each event cluster with the event cluster identifier corresponding to the event cluster, wherein each event cluster includes at least one message belonging to the same event;
[0239] The embedding network is trained based on a first similarity between the anchor sample titles in at least two triples and the positive sample titles, and a second similarity between the anchor sample titles and the negative sample titles;
[0240] Each triplet includes an anchor sample title, a positive sample title, and a negative sample title. The anchor sample title and the positive sample title belong to the same event, and the anchor sample title and the negative sample title belong to different events.
[0241] In one possible implementation, the apparatus is further configured to train the embedding network. When training the embedding network, the apparatus further includes:
[0242] A sample data set acquisition module is used to obtain a sample data set and a label of each sample in the sample data set;
[0243] Each sample includes a base sample title and a candidate sample title, and the label of each sample includes a first event label and a second event label, wherein the first event label indicates whether the candidate sample title and the base sample title belong to the same event, and the second event label indicates whether the text corresponding to the candidate sample title and the base sample title respectively belongs to the same event;
[0244] A triplet construction module, configured to construct the at least two triples based on the label of each sample in the sample data set;
[0245] a similarity determination module, configured to determine, based on the feature vectors of the anchor sample title, the positive sample title, and the negative sample title in each triplet, a first similarity between the anchor sample title and the positive sample title, and a second similarity between the anchor sample title and the negative sample title, respectively, wherein the feature vectors of the anchor sample title, the positive sample title, and the negative sample title are obtained by extracting features from the anchor sample title, the positive sample title, and the negative sample title, respectively, using the initial embedding network;
[0246] A training module is configured to train the initial embedding network based on a difference between the first similarity and the second similarity to obtain the embedding network.
[0247] In one possible implementation, the training module is further used to iteratively train the initial embedding network when the difference between the second similarity and the first similarity is not higher than a target value, and stop training until the second similarity is higher than the target value of the first similarity to obtain the embedding model.
[0248] In one possible implementation, the triple construction module includes:
[0249] a title pair acquisition unit, configured to acquire at least two title pairs from the sample dataset based on the label of each sample in the sample dataset, wherein each title pair includes an anchor sample title and a positive sample title belonging to the same event;
[0250] a distance determining unit configured to obtain, for each title pair, a first distance between the anchor sample title and the positive sample title based on their respective feature vectors, and obtain a second distance between at least one candidate sample title and the anchor sample title based on their respective feature vectors;
[0251] A triplet determination unit, configured to determine a first type of triplet, a second type of triplet, and a third type of triplet based on the first distance and the second distance corresponding to each title pair;
[0252] Among them, the difference between the second distance and the first distance corresponding to the first type triplet is greater than the target value, the first distance corresponding to the second type triplet is less than the second distance, and the sum of the first distance and the target value is greater than the second distance, and the second distance corresponding to the third type triplet is less than the first distance.
[0253] In one possible implementation, the title pair acquisition unit is used to:
[0254] For each iterative training, the batch sample data used in the current iterative training is obtained from the sample dataset;
[0255] Based on the label of each sample in the batch of sample data, obtaining the at least two title pairs from the batch of sample data sets;
[0256] Accordingly, the device further comprises:
[0257] The candidate sample title obtaining unit is configured to obtain the at least one candidate sample title from the batch of sample data.
[0258] In one possible implementation, the triple construction module includes:
[0259] A second title pair acquisition unit is used to acquire a title pair based on the label of each sample in the sample dataset, where the title pair includes an anchor sample title and a positive sample title belonging to the same event;
[0260] A negative sample title sampling unit is used to randomly sample a target number of sample titles for each title pair as negative sample titles corresponding to the title pair;
[0261] The triplet acquisition unit is configured to obtain the at least two triplets based on each title pair and its corresponding negative sample title.
[0262] In one possible implementation, the sample data set acquisition module includes:
[0263] a duplication detection unit, configured to perform duplication detection on the original information set to obtain at least two information subsets, each information subset including at least one information having a similarity within a target threshold range;
[0264] A sample acquisition unit is configured to, for each information subset, use the title of the target information in the information subset as a reference sample title, and acquire a candidate sample title corresponding to the reference sample title to obtain a sample;
[0265] A first event label determination unit is configured to determine, for each sample, a first event label corresponding to the benchmark sample title and the corresponding candidate sample title based on the title keywords respectively included in the benchmark sample title and the corresponding candidate sample title;
[0266] The second event label determination unit is configured to determine the second event labels corresponding to the benchmark sample title and its corresponding candidate sample title based on the respective texts corresponding to the benchmark sample title and its corresponding candidate sample title.
[0267] In one possible implementation, the device further includes:
[0268] An adding module, configured to add new information to the candidate recommendation list based on the event cluster identifier of the remaining information in the candidate recommendation list to obtain the recommendation list;
[0269] The event cluster identifier of the newly added information is different from the remaining information, and the remaining information is information other than the deleted information in the candidate recommendation list.
[0270] The information recommendation device provided in this application obtains a recommendation list by deleting information that meets the event cluster similarity conditions from a candidate recommendation list based on event cluster identifiers of at least two pieces of information, and then makes recommendations. This filter uses the event cluster identifiers to filter similar information belonging to the same event cluster, preventing similar information from being concentrated in the same recommendation list and improving recommendation efficiency. Furthermore, the device obtains the title feature vectors of each piece of information through an embedded network and performs vector clustering on the title feature vectors of each piece of information to obtain at least two event clusters, each of which includes at least one piece of information belonging to the same event. This divides each piece of information into different event clusters, allowing the event clusters to be utilized to quickly filter and disperse similar information from the same event in the candidate recommendation list, thereby increasing the richness and diversity of the recommendation list and, in turn, improving actual recommendation efficiency.
[0271] Moreover, the embedding network is trained based on the first similarity between the anchor sample titles and the positive sample titles, and the second similarity between the anchor sample titles and the negative sample titles in at least two triples; while the anchor sample titles and the positive sample titles belong to the same event, and the anchor sample titles and the negative sample titles belong to different events. The present application uses the embedding network to perform feature extraction, thereby achieving the effect of effectively bringing the title features of the same event closer and the distinctive features of different events farther apart in the feature space, thereby improving the differentiated expression of the title feature vector for different events, thereby improving the accuracy of clustering, and further improving the accuracy of information recommendation.
[0272] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, and will not be repeated here.
[0273] Figure 10 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 10 As shown, the computer device includes: a memory, a processor, and a computer program stored in the memory. The processor executes the above computer program to implement the steps of the information recommendation method. Compared with the related art, it can achieve:
[0274] The information recommendation method provided by the present application obtains a recommendation list by deleting information that meets the event cluster similarity conditions from a candidate recommendation list based on event cluster identifiers of at least two pieces of information, and then makes recommendations. This method filters similar information belonging to the same event cluster through the event cluster identifier, preventing similar information from being concentrated in the same recommendation list, thereby improving recommendation efficiency. Furthermore, the method obtains the title feature vectors of each piece of information through an embedded network, performs vector clustering on the title feature vectors of each piece of information, and obtains at least two event clusters, each of which includes at least one piece of information belonging to the same event. This method divides each piece of information into different event clusters, allowing the use of the event cluster to quickly filter and disperse similar information of the same event in the candidate recommendation list, thereby improving the richness and diversity of the recommendation list and, in turn, improving actual recommendation efficiency.
[0275] Moreover, the embedding network is trained based on the first similarity between the anchor sample titles and the positive sample titles, and the second similarity between the anchor sample titles and the negative sample titles in at least two triples; while the anchor sample titles and the positive sample titles belong to the same event, and the anchor sample titles and the negative sample titles belong to different events. The present application uses the embedding network to perform feature extraction, thereby achieving the effect of effectively bringing the title features of the same event closer and the distinctive features of different events farther apart in the feature space, thereby improving the differentiated expression of the title feature vector for different events, thereby improving the accuracy of clustering, and further improving the accuracy of information recommendation.
[0276] In an alternative embodiment, a computer device is provided, such as Figure 10 As shown, Figure 10The computer device 1000 shown includes a processor 1001 and a memory 1003. The processor 1001 and the memory 1003 are connected, for example, via a bus 1002. Optionally, the computer device 1000 may also include a transceiver 1004, which can be used for data exchange between the computer device and other computer devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 1004 is not limited to one, and the structure of the computer device 1000 does not constitute a limitation on the embodiments of this application.
[0277] Processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, or a combination of a DSP and a microprocessor.
[0278] Bus 1002 may include a path for transmitting information between the above components. Bus 1002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. Bus 1002 may be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 10 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0279] The memory 1003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.
[0280] The memory 1003 is used to store the computer program for executing the embodiments of the present application, and the execution is controlled by the processor 1001. The processor 1001 is used to execute the computer program stored in the memory 1003 to implement the steps shown in the above method embodiments.
[0281] Among them, electronic equipment includes but is not limited to: servers, terminals or cloud computing center equipment, etc.
[0282] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0283] An embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiment when executed by a processor.
[0284] It can be understood that in the specific implementation of this application, any object-related data such as information uploaded by producers in the resource pool, the title of the information, the graphics, text, video and other content included in the information, etc., when the above embodiments of this application are applied to specific products or technologies, it is necessary to obtain the object's permission or consent, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0285] Those skilled in the art will understand that, unless otherwise specified, the singular forms "a," "an," "the," and "the" used herein may also include the plural forms. The terms "including" and "comprising" used in the embodiments of this application mean that the corresponding features can be implemented as the presented features, information, data, steps, and operations, but do not exclude the implementation of other features, information, data, steps, operations, etc. supported by the technical field.
[0286] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," "fourth," "1," "2," and so on (if any) are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the present application described herein can be practiced in an order other than that shown or described.
[0287] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times respectively. Under different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0288] The above description is only an optional implementation method for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, the use of other similar implementation methods based on the technical ideas of this application also falls within the protection scope of the embodiments of this application.
Claims
1. An information recommendation method, characterized in that: The method comprises: In response to a recommendation request from any object, based on the event cluster identifiers of at least two pieces of information, information that meets the event cluster similarity condition is deleted from a candidate recommendation list to obtain a recommendation list, and the to-be-recommended information stream corresponding to the recommendation list is recommended to the object; wherein the event cluster similarity condition includes being identical to the event cluster identifier of any other information in the candidate recommendation list; The method for obtaining the event cluster identifier includes: Obtain the title feature vector of each information in the resource pool through the trained embedding network; Performing vector clustering on the title feature vectors of each information in the resource pool to obtain at least two event clusters, and marking the information included in each event cluster with an event cluster identifier corresponding to the event cluster, wherein each event cluster includes at least one information belonging to the same event; The embedding network is trained in the following way: Obtain a sample data set and a label for each sample in the sample data set; wherein each sample includes a base sample title and a candidate sample title, and the label for each sample includes a first event label and a second event label, wherein the first event label indicates whether the candidate sample title and the base sample title belong to the same event, and the second event label indicates whether the text corresponding to the candidate sample title and the base sample title respectively belongs to the same event; Based on the label of each sample in the sample data set, construct at least two triplets; each triplet includes an anchor sample title, a positive sample title, and a negative sample title, the anchor sample title and the positive sample title belong to the same event, and the anchor sample title and the negative sample title belong to different events; Based on the respective feature vectors of the anchor sample title, the positive sample title, and the negative sample title in each triplet, respectively determining a first similarity between the anchor sample title and the positive sample title, and a second similarity between the anchor sample title and the negative sample title, wherein the respective feature vectors of the anchor sample title, the positive sample title, and the negative sample title are obtained by performing feature extraction on the anchor sample title, the positive sample title, and the negative sample title respectively through the initial embedding network; When the difference between the second similarity and the first similarity is not higher than a target value, the initial embedding network is iteratively trained until the second similarity is higher than the target value of the first similarity, and the training is stopped to obtain the embedding network.
2. The information recommendation method according to claim 1, characterized in that: The step of constructing at least two triples based on the label of each sample in the sample data set includes: Based on the label of each sample in the sample dataset, obtaining at least two title pairs from the sample dataset, each title pair including an anchor sample title and a positive sample title belonging to the same event; For each title pair, obtaining a first distance between the anchor sample title and the positive sample title based on their respective feature vectors in the title pair, and obtaining a second distance between at least one candidate sample title and the anchor sample title based on their respective feature vectors; Determine first-category triples, second-category triples, and third-category triples based on the first distance and the second distance corresponding to each title pair; Among them, the difference between the second distance and the first distance corresponding to the first type triplet is greater than the target value, the first distance corresponding to the second type triplet is less than the second distance, and the sum of the first distance and the target value is greater than the second distance, and the second distance corresponding to the third type triplet is less than the first distance.
3. The information recommendation method according to claim 2, characterized in that: The acquiring at least two title pairs from the sample dataset based on the label of each sample in the sample dataset comprises: For each iterative training, obtaining batch sample data used in the current iterative training from the sample data set; Obtaining the at least two title pairs from the batch sample data set based on a label of each sample in the batch sample data; Before obtaining the second distance between the at least one candidate sample title and the anchor sample title based on their respective feature vectors, the method further includes: The at least one candidate sample title is then obtained from the batch sample data.
4. The information recommendation method according to claim 1, characterized in that The step of constructing at least two triples based on the label of each sample in the sample data set includes: Based on the label of each sample in the sample dataset, obtaining a title pair, wherein the title pair includes an anchor sample title and a positive sample title belonging to the same event; For each title pair, randomly sample the target number of sample titles as the negative sample titles corresponding to the title pair; Based on each title pair and its corresponding negative sample title, at least two triplets are obtained.
5. The information recommendation method according to claim 1, characterized in that: The obtaining of a sample data set and a label of each sample in the sample data set includes: Repeatedly detecting the original information set to obtain at least two information subsets, each information subset including at least one information having a similarity within a target threshold range; For each information subset, taking the title of the target information in the information subset as the benchmark sample title, and obtaining the candidate sample title corresponding to the benchmark sample title to obtain a sample; For each sample, determining the first event label corresponding to the benchmark sample title and the corresponding candidate sample title based on the title keywords respectively included in the benchmark sample title and the corresponding candidate sample title; Based on the respective texts corresponding to the benchmark sample title and the corresponding candidate sample title, the second event labels corresponding to the benchmark sample title and the corresponding candidate sample title are determined.
6. The information recommendation method according to claim 1, characterized in that: After deleting the information in the candidate recommendation list that meets the event cluster similarity condition based on the event cluster identification of at least two pieces of information, the method further includes: adding new information to the candidate recommendation list based on the event cluster identifiers of the remaining information in the candidate recommendation list to obtain the recommendation list; The event cluster identifier of the newly added information is different from the remaining information, and the remaining information is information other than the deleted information in the candidate recommendation list.
7. The information recommendation method according to claim 1, characterized in that: The forms of each information in the resource pool include at least two.
8. An information recommendation device, characterized in that: The device comprises: a recommendation list determination module, configured to, in response to a recommendation request for any object, delete information from a candidate recommendation list that meets an event cluster similarity condition based on event cluster identifiers of at least two pieces of information, thereby obtaining a recommendation list; wherein the event cluster similarity condition includes being identical to an event cluster identifier of any other information in the candidate recommendation list; A recommendation module, configured to recommend the information flow to be recommended corresponding to the recommendation list to any of the objects; The device is further configured to obtain the event cluster identifier. When obtaining the event cluster identifier, the device further includes: The title feature vector acquisition module is used to obtain the title feature vector of each information in the resource pool through the trained embedding network; a clustering module configured to perform vector clustering on the title feature vectors of each message in the resource pool to obtain at least two event clusters, and mark the information included in each event cluster with an event cluster identifier corresponding to the event cluster, wherein each event cluster includes at least one message belonging to the same event; Wherein, when training the embedding network, the device further includes: A sample data set acquisition module is configured to acquire a sample data set and a label for each sample in the sample data set; wherein each sample includes a base sample title and a candidate sample title, and each sample label includes a first event label and a second event label, wherein the first event label indicates whether the candidate sample title and the base sample title belong to the same event, and the second event label indicates whether the text corresponding to the candidate sample title and the base sample title respectively belongs to the same event; A triplet construction module is configured to construct at least two triples based on the label of each sample in the sample dataset; each triplet includes an anchor sample title, a positive sample title, and a negative sample title, wherein the anchor sample title and the positive sample title belong to the same event, and the anchor sample title and the negative sample title belong to different events; a similarity determination module, configured to determine, based on the feature vectors of the anchor sample title, the positive sample title, and the negative sample title in each triplet, a first similarity between the anchor sample title and the positive sample title, and a second similarity between the anchor sample title and the negative sample title, respectively, wherein the feature vectors of the anchor sample title, the positive sample title, and the negative sample title are obtained by extracting features from the anchor sample title, the positive sample title, and the negative sample title, respectively, using an initial embedding network; A training module is used to iteratively train the initial embedding network when the difference between the second similarity and the first similarity is not higher than a target value, and stop training until the second similarity is higher than the target value of the first similarity to obtain the embedding network.
9. The information recommendation device according to claim 8, characterized in that The triple construction module includes: a title pair acquiring unit, configured to acquire at least two title pairs from the sample dataset based on the label of each sample in the sample dataset, each title pair comprising an anchor sample title and a positive sample title belonging to the same event; a distance determining unit configured to obtain, for each title pair, a first distance between the anchor sample title and the positive sample title based on their respective feature vectors, and obtain a second distance between at least one candidate sample title and the anchor sample title based on their respective feature vectors; A triplet determination unit, configured to determine a first type of triplet, a second type of triplet, and a third type of triplet based on the first distance and the second distance corresponding to each title pair; Among them, the difference between the second distance and the first distance corresponding to the first type triplet is greater than the target value, the first distance corresponding to the second type triplet is less than the second distance, and the sum of the first distance and the target value is greater than the second distance, and the second distance corresponding to the third type triplet is less than the first distance.
10. The information recommendation device according to claim 9, characterized in that: The title pair acquiring unit is configured to: For each iterative training, obtaining batch sample data used in the current iterative training from the sample data set; Obtaining the at least two title pairs from the batch sample data set based on a label of each sample in the batch sample data; The apparatus further includes a candidate sample title obtaining unit, wherein the candidate sample title obtaining unit is configured to: The at least one candidate sample title is then obtained from the batch sample data.
11. The information recommendation device according to claim 8, characterized in that: The triple construction module includes: A second title pair acquisition unit, configured to acquire a title pair based on the label of each sample in the sample dataset, wherein the title pair includes an anchor sample title and a positive sample title belonging to the same event; A negative sample title sampling unit, configured to randomly sample a target number of sample titles for each title pair as negative sample titles corresponding to the title pair; The triplet acquisition unit is configured to obtain at least two triplets based on each title pair and its corresponding negative sample title.
12. The information recommendation device according to claim 8, characterized in that The sample data set acquisition module includes: a duplication detection unit, configured to perform duplication detection on the original information set to obtain at least two information subsets, each information subset including at least one information having a similarity within a target threshold range; a sample acquisition unit configured to, for each information subset, use the title of target information in the information subset as a reference sample title, and acquire a candidate sample title corresponding to the reference sample title to obtain a sample; A first event label determination unit is configured to determine, for each sample, a first event label corresponding to the benchmark sample title and the corresponding candidate sample title based on the title keywords respectively included in the benchmark sample title and the corresponding candidate sample title; The second event label determination unit is configured to determine the second event labels corresponding to the benchmark sample title and its corresponding candidate sample title based on the respective texts corresponding to the benchmark sample title and its corresponding candidate sample title.
13. The information recommendation device according to claim 8, characterized in that The device further comprises: an adding module, configured to add new information to the candidate recommendation list based on the event cluster identifiers of the remaining information in the candidate recommendation list to obtain the recommendation list; The event cluster identifier of the newly added information is different from the remaining information, and the remaining information is information other than the deleted information in the candidate recommendation list.
14. The information recommendation device according to claim 8, characterized in that The forms of each information in the resource pool include at least two.
15. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the information recommendation method according to any one of claims 1 to 7.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the information recommendation method according to any one of claims 1 to 7 are implemented.
17. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the information recommendation method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Information recommendation method and device
CN109886353A
Information recommendation method and device
CN111125460A