News recommendation method based on image-text combination
By constructing a keyword lexicon and user interest map, combining the relevance of graphic and text and news popularity, personalized news recommendations are achieved, solving the problem that the relevance of graphic and text in the existing technology has not been effectively analyzed, and improving the accuracy and user experience of news recommendations.
Patent Information
- Application Number
- CN202510092623.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing news recommendation methods fail to effectively analyze the relevance of pictures and texts, resulting in the unrelated pictures and text combinations bringing users a bad visual experience, reducing the accuracy of news recommendations and the user's information acquisition efficiency.
By constructing keyword lexicon, user interest map and graphic relevance analysis, combined with news popularity and recommendation index, personalized recommendations for news that users are interested in.
It improves the accuracy and user experience of news recommendations, reduces the time for users to screen irrelevant news, and improves the efficiency of information acquisition.
Smart Images

Figure CN120011634A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer computing, and in particular to a news recommendation method based on the combination of images and texts. Background Art
[0002] With the rapid development of digital technology, the news dissemination pattern has undergone tremendous changes. Today, the amount of news on the Internet has exploded, and various news platforms are generating massive amounts of news information every day. This phenomenon has led to a serious information overload problem. Users are faced with the dilemma of finding interesting and valuable information from a large amount of news. An ordinary news aggregation platform may update tens of thousands of news articles every day. In the absence of an effective recommendation mechanism, users need to spend a lot of time and energy to browse these news.
[0003] The existing Chinese patent with application number 202210145950.8 discloses a news recommendation method and a news recommendation system. The scheme obtains current user information, user behavior information and news data, pre-screens candidate news data based on this information, and then uses the feature cross-model of the multi-head attention mechanism to calculate the news recommendation index, and determines the active and inactive news data through the steps of acquisition, splicing, screening, and elimination of outliers. Finally, the candidate news with the highest recommendation index is pushed to the user, thereby improving the accuracy of news recommendation.
[0004] However, the above patent has the following problems: the solution does not involve the analysis of the relevance of pictures and texts when analyzing the news recommendation index. Irrelevant pictures and texts will bring a bad visual experience to users. If the pictures cannot well assist in explaining the news content, even if the news topic may be of interest to users, the overall news presentation effect may be greatly reduced, resulting in reduced user interest in the news. At the same time, irrelevant pictures and texts will increase the difficulty for users to understand the news, causing users to spend more time to identify the news content, reducing the efficiency of information acquisition. Summary of the invention
[0005] In order to overcome the shortcomings of the background technology, the embodiment of the present invention provides a news recommendation method based on the combination of pictures and texts, which can effectively solve the problems involved in the above-mentioned background technology.
[0006] The purpose of the present invention can be achieved through the following technical solutions: a news recommendation method based on the combination of pictures and texts, the method comprising the following steps: S1. Keyword thesaurus construction: using web crawler technology to obtain each type of news corresponding to each piece of news, and extracting initial keywords from them to form a keyword thesaurus for each type of news.
[0007] S2. User behavior data collection: Obtain the user's behavior data news, and determine the news category to which the news corresponding to each user's behavior data belongs. Each behavior data news includes each browsing interest news, each liked news, each commented news, and each collected news.
[0008] S3. Construction of user interest graph: Obtain the total interest weight of users in each news category to build a user interest graph framework.
[0009] S4. Retrieve news interest matching: retrieve each preliminary search news according to each search term of the user, and obtain the interest weight of each preliminary search news according to the user interest graph framework.
[0010] S5. Image-text correlation analysis: Analyze the image-text consistency of the text information images and non-text information images of each user's preliminary search news to obtain the image-text correlation of each user's preliminary search news, and use this to filter out each user's searched news to obtain the image-text correlation of each user's searched news.
[0011] S6. News popularity analysis: Obtain the popularity parameters of each retrieved news in each time range, and then analyze the news popularity of each retrieved news. The popularity parameters include page views, likes, shares, and comments.
[0012] S7. News recommendation index analysis: Filter the interest weights Q of each retrieved news from the interest weights of each preliminary retrieved news i , combined with the image-text relevance β i 、News popularity γ i The news recommendation index of each retrieved news of the user is obtained by analysis, i represents the number of the i-th retrieved news, i=1,2,...,n, the news recommendation index of each retrieved news of the user is sorted and pushed to the user.
[0013] Preferably, the specific operation method for constructing the keyword thesaurus is: traverse various news websites through web crawler technology to obtain various news in various news websites, record them as various news corresponding to each news, use natural language processing technology to analyze various news corresponding to each news, extract initial keywords from them respectively, and reorganize them into keyword thesaurus for various news through decomposition.
[0014] Preferably, the specific analysis method for each user's browsing interest news is as follows: the first step is to read the user's browsing history data on the news platform from the management database, obtain the user's browsing time record for each news, and obtain the user's browsing time for each news.
[0015] In the second step, by comparing the browsing time of each news item of the user with the preset browsing time threshold, the news items whose browsing time of the user exceeds the preset browsing time threshold are screened out and recorded as the browsing interest news of the user.
[0016] Preferably, the specific analysis method for collecting the user behavior data is as follows: in the first step, the keywords of the titles of the news of interest browsed by the users are extracted through text cleaning technology, and they are matched with the keyword thesaurus of each type of news respectively to obtain the number of keywords of the titles of the news of interest browsed by the users in the keyword thesaurus of each type of news.
[0017] In the second step, the number of keywords in each news keyword dictionary of each user's browsing interest news is arranged in order from large to small, and the news category corresponding to the first keyword dictionary is recorded as the news category to which the browsing interest news belongs, thereby obtaining the news category to which each user's browsing interest news belongs.
[0018] The third step is to obtain the user's likes, comments and collection records on the news platform from the news platform backend, and record the corresponding news as the user's liked news, commented news, and collected news. According to the method of analyzing the news categories to which the user's browsed interest news belongs, the news categories to which the user's liked news, commented news, and collected news belong are analyzed respectively.
[0019] Preferably, the specific operation method for constructing the user interest graph is as follows: the first step is to respectively read the news categories to which each user's browsed interest news belongs, the news categories to which each user's liked news, commented news, and collected news belongs, and assign corresponding weights to the user's browsed interest news, liked news, commented news, and collected news according to the set weights to obtain the weights of each user's browsed interest news, liked news, commented news, and collected news.
[0020] In the second step, the total weight of each news category is obtained by accumulating the weights of each user's browsing interest news, each liked news, each commented news, and each collected news in each news category, which is recorded as the user's total interest weight in each news category.
[0021] The third step is to construct a user interest graph framework using each news category as different dimensions of the graph, and mark the user's total interest weight in each news category on the corresponding dimension.
[0022] Preferably, the specific operation method for retrieving news interest matches is: the first step is to obtain the user's search terms by obtaining the search terms input by the user during the search, obtain the news items containing the user's search terms in the news website by searching the user's search terms, record them as preliminary search news, and obtain the user's search terms contained in the preliminary search news, record them as search keywords for the preliminary search news.
[0023] In the second step, each search keyword of each preliminary search news is queried in each type of news keyword thesaurus to obtain the category news keyword thesaurus corresponding to each search keyword of each preliminary search news. At the same time, the total interest weight of the user in each news category is read from the user interest graph framework, and the total interest weight of the user in the category news keyword thesaurus corresponding to each search keyword of each preliminary search news is obtained by screening, which is recorded as the interest weight of each search keyword of each preliminary search news of the user, and the interest weight Q of each preliminary search news is obtained by accumulating it. j , where j represents the number of the jth preliminary retrieved news, j = 1, 2, ..., g.
[0024] Preferably, the specific analysis method for the text-image consistency of the text information images and non-text information images of each user's preliminary search news is as follows: the first step is to read each user's preliminary search news, obtain each image in each user's preliminary search news, and divide it into each text information image of each user's preliminary search news and each non-text information image of each user's preliminary search news.
[0025] In the second step, natural language processing technology is used to extract keywords from the text information images of each preliminary search news by the user, and the news category of each preliminary search news and the news category of each text information image of each preliminary search news by the user are obtained by analyzing the news category of each news of interest to which the user browses.
[0026] In the third step, the news category to which each text information image of each preliminary search news of the user belongs is compared with the news category to which each preliminary search news belongs. If a text information image of a preliminary search news of the user corresponds to the news category to which the preliminary search news belongs, the image-text consistency of the text information image of the preliminary search news of the user is recorded as 1, otherwise the image-text consistency of the text information image of the preliminary search news of the user is recorded as 0. Thus, the image-text consistency of each text information image of each preliminary search news of the user is obtained by analysis, and the image-text consistency of the text information image of each preliminary search news of the user is obtained by accumulation, which is recorded as
[0027] In the fourth step, the pre-trained convolutional neural network model is used to classify the non-text information images of each preliminary search news by the user, and the image type of each non-text information image of each preliminary search news by the user is obtained. According to the method of analyzing the image-text consistency of the text information image of each preliminary search news by the user, the image-text consistency of the non-text information image of each preliminary search news by the user is obtained, which is recorded as
[0028] Preferably, the specific analysis method of the image-text relevance analysis is as follows: the first step is to read the image-text consistency of the text information image and the non-text information image of each preliminary searched news by the user. By formula Get the image and text relevance β of each user's initial search news j , where φ1 and φ2 represent the weight factors of the preset text information image's text-image consistency and the non-text information image's text-image consistency, respectively, and e represents a natural constant.
[0029] In the second step, the text-image relevance of each preliminary search news of the user is compared with the preset text-image relevance threshold, and all preliminary search news with text-image relevance greater than or equal to the preset text-image relevance threshold are screened out, which are recorded as each search news of the user. The text-image relevance of each search news of the user is screened out, which is recorded as β i , where i represents the number of the i-th retrieved news, i=1,2,...,n.
[0030] Preferably, the specific analysis method of the news heat analysis is: select different time ranges according to the settings, record them as each time range, obtain the number of times the news page of each searched news by the user in each time range is visited from the news website, record them as the number of page views L of each searched news by the user in each time range im , where m represents the number of the mth time range, m = 1, 2, ..., q, and the number of likes, shares, and comments of each time range of each retrieved news by the user is obtained, which are recorded as D im 、F im , P im , the average page views, likes, shares, and comments of each retrieved news are obtained by average calculation, recorded as Substituting this into the formula
[0031] Get the news popularity of each searched newsγ i ,in They represent the weight factors of the preset page views, likes, shares, and comments respectively, n represents the number of retrieved news, and q represents the number of time ranges.
[0032] Preferably, the specific analysis method of the news recommendation index analysis is: reading the interest weight of each preliminary searched news, and screening out the interest weight of each searched news, which is recorded as Q i , and read the image and text relevance β of each retrieved news separately i 、News popularity γ i , substituting it into the formula Get the news recommendation index of each user's searched news Among them, w1, w2, and w3 represent the preset interest weight, image-text relevance, and news popularity weight factors respectively, and e represents a natural constant. The news recommendation index of each user's searched news is sorted in descending order, and the sorted searched news is pushed to the user.
[0033] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects: 1. The present invention can widely collect keywords in the news and construct a comprehensive keyword thesaurus for various types of news by extracting keywords corresponding to each piece of news from each type of news, covering the key information of various types of news.
[0034] 2. The present invention constructs a user interest graph framework by analyzing the total interest weights of users in each news category, from which the interest weights of each retrieved news are obtained, and recommendations are made based on each user's unique interest pattern, which can reduce the time users spend screening irrelevant news.
[0035] 3. The present invention obtains the image-text relevance of each news searched by the user according to the image-text relevance screening, which is helpful for news screening. When searching for news, news with high image-text relevance can be obtained first, thereby improving the efficiency of obtaining effective information.
[0036] Fourth, the present invention obtains the news popularity of each retrieved news according to the popularity parameters of each retrieved news in each time range, provides a reference for news timeliness, and can judge which news is the current hot topic according to the news popularity, and obtain hot information in time.
[0037] 5. The present invention obtains the news recommendation index of each retrieved news by the user according to the interest weight, image-text relevance and news popularity analysis of each retrieved news, sorts them and pushes them to the user, so that news can be pushed in a targeted manner, thereby improving the utilization rate of news resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The present invention is further described using the accompanying drawings, but the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative work.
[0039] Figure 1 It is a schematic diagram of the method flow of the present invention.
[0040] Figure 2 for Figure 1 Schematic diagram of the process of browsing news of interest in step S1.
[0041] Figure 3 for Figure 1 Schematic diagram of the process of step S3 in FIG. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] See also Figure 1 As shown, the present invention provides a news recommendation method based on the combination of pictures and texts, and the method includes the following steps: S1. Keyword word library construction: using web crawler technology to obtain various types of news corresponding to each piece of news, and extracting initial keywords from them to form various types of news keyword word libraries.
[0044] The specific operation method of constructing the keyword word library is as follows: traverse various news websites through web crawler technology to obtain various news in various news websites, record them as various news corresponding to each news, use natural language processing technology to analyze various news corresponding to each news, extract initial keywords from them respectively, and form various news keyword word libraries through decomposition; it is possible to widely collect keywords in news and construct comprehensive keyword word libraries of various news, covering key information of various types of news.
[0045] S2. User behavior data collection: Obtain the user's behavior data news, and determine the news category to which the news corresponding to each user's behavior data belongs. Each behavior data news includes each browsing interest news, each liked news, each commented news, and each collected news.
[0046] See also Figure 2 As shown, the specific analysis method for each user's browsing interest news is as follows: the first step is to read the user's browsing history data on the news platform from the management database, obtain the user's browsing time record for each news, and obtain the user's browsing time for each news; it can quantify the user's attention to different news, and more truly reflect the user's interest in news inclination in an unconscious state.
[0047] In the second step, by comparing the user's browsing time of each news with the preset browsing time threshold, the news whose browsing time exceeds the preset browsing time threshold is screened out and recorded as the user's browsing interest news; by setting the threshold, the news that the user is really interested in is screened out, and the news that the user may just click on accidentally or browse quickly is excluded.
[0048] See also Figure 3As shown, the specific analysis method for collecting user behavior data is as follows: the first step is to extract the keywords of the titles of the user's browsing interest news through text cleaning technology, and match them with the keyword thesaurus of each type of news respectively to obtain the number of keywords of the user's browsing interest news in the keyword thesaurus of each type of news; through the matching of keywords, it is possible to preliminarily determine the news category to which the user's interest news belongs, laying the foundation for further determining the user's overall news browsing interest.
[0049] In the second step, the number of keywords in each type of news keyword dictionary of each user's browsing interest news is arranged in descending order, and the news category corresponding to the first keyword dictionary is recorded as the news category to which the browsing interest news belongs, thereby obtaining the news category to which each user's browsing interest news belongs; it is possible to more accurately classify the user's browsing interest news and provide specific classification results for a comprehensive understanding of the user's news browsing interest preferences.
[0050] It should be noted that, in a specific embodiment, the user browsing history is read from the database. The user browsed news A for 10 minutes, news B for 15 minutes, and news C for 3 minutes and seconds, and the preset browsing time threshold is 5 minutes, that is, news A and news B are determined to be the user's browsing interest news. The keywords extracted from the title of news A through text cleaning technology are "technology company", "artificial intelligence", and "algorithm". There are 3 keyword matches in the keyword vocabulary of technology news and 1 keyword match in the keyword vocabulary of financial news. The keywords extracted from the title of news B are "exchange rate" and "import and export enterprise". There are 2 keyword matches in the keyword vocabulary of financial news and 1 keyword match in the keyword vocabulary of technology news. The number of keywords of news A in each type of news keyword vocabulary is arranged from large to small: technology (3) > finance (1), so the news category to which news A belongs is technology news. The number of keywords of news B in each type of news keyword vocabulary is arranged from large to small: finance (2) > technology (1), so the news category to which news B belongs is finance news.
[0051] The third step is to obtain the user's likes, comments and collection records on the news platform from the news platform backend, and record the corresponding news as the user's liked news, commented news, and collected news. According to the method of analyzing the news categories to which the user's browsing interest news belongs, the news categories to which the user's liked news, commented news, and collected news are analyzed respectively; this can more directly reflect the user's preferences, and can more completely and accurately construct the user's news interest portrait, which will help the news platform to perform more accurate personalized recommendations and other operations.
[0052] It should be noted that, in a specific embodiment, the news categories of the news platform include "culture and art" and "science and technology". The keyword vocabulary corresponding to culture and art news includes: painting, sculpture, music performance, drama, etc., and the keyword vocabulary corresponding to science and technology news includes: electronic equipment, software development, scientific and technological innovation, scientific research results, etc. The user's likes, comments and collection records on the news platform are obtained from the background of the news platform, among which the commented news includes news C (titled "A grand concert was held in the capital"), and the liked news and collection news include news D (titled "A technology company launches a new smart phone"). The keywords extracted from news C are "concert" and "capital", and there is 1 keyword match in the keyword vocabulary of culture and art news, and 0 keyword matches in the keyword vocabulary of science and technology news, so news C belongs to culture and art news, and the keywords extracted from news D are "technology company" and "smart phone", and there are 2 keyword matches in the keyword vocabulary of science and technology news, and 0 keyword matches in the keyword vocabulary of culture and art news, so news D belongs to science and technology news.
[0053] S3. Construction of user interest graph: Obtain the total interest weight of users in each news category to build a user interest graph framework.
[0054] See also Figure 3 As shown, the specific operation method for constructing the user interest graph is: the first step is to read the news categories to which each user's browsed interest news belongs, the news categories to which each user's liked news, commented news, and collected news belong, and assign corresponding weights to the user's browsed interest news, liked news, commented news, and collected news according to the set weights, so as to obtain the weights of each user's browsed interest news, liked news, commented news, and collected news; the importance of different types of news to the user's interest representation can be distinguished, and the user's interest in different news can be more accurately reflected.
[0055] In the second step, the total weight of each news category is obtained by accumulating the weights of the corresponding users' browsed interest news, liked news, commented news, and collected news in each news category, which is recorded as the user's total interest weight in each news category. This can comprehensively evaluate the user's comprehensive interest in each news category, avoid the one-sidedness of judging the user's interest due to a single behavior, and thus more accurately grasp the user's overall preference tendencies in different news categories.
[0056] It should be noted that, in a specific embodiment, according to the set weights, the weight of browsing interest news is 3, the weight of liking news is 2, the weight of commenting on news is 2, and the weight of collecting news is 3. The news categories to which news A, B, C, and D belong are technology, finance, culture and art, and technology, respectively. Browsing interest news includes news A and news B, commenting on news includes news C, and liking news and collecting news include news D. Therefore, the total interest weight of the user in technology news is 3+2+3=8, the total interest weight of the user in finance news is 3, and the total interest weight of the user in culture and art news is 2.
[0057] The third step is to build a user interest graph framework using each news category as different dimensions of the graph, and mark the user's total interest weight in each news category on the corresponding dimension; this helps to quickly understand the user's interest preferences, such as which news categories they have a strong interest in and which categories they have a weaker interest in, making it easier to perform personalized recommendations, content screening and other operations.
[0058] S4. Retrieve news interest matching: retrieve each preliminary search news according to each search term of the user, and obtain the interest weight of each preliminary search news according to the user interest graph framework.
[0059] The specific operation method for retrieving news interest matching is as follows: the first step is to obtain the user's search terms by obtaining the search terms input by the user during the search, obtain the news on the news website containing the user's search terms by searching the user's search terms, record them as preliminary search news, obtain the user's search terms contained in the preliminary search news, record them as search keywords for the preliminary search news; the news related to the user's search terms can be quickly located to meet the user's search needs for specific information, and basic data can be provided for subsequent more accurate screening and sorting by determining the preliminary search news and its search keywords.
[0060] In the second step, each search keyword of each preliminary search news is queried in each type of news keyword thesaurus to obtain the category news keyword thesaurus corresponding to each search keyword of each preliminary search news. At the same time, the total interest weight of the user in each news category is read from the user interest graph framework, and the total interest weight of the user in the category news keyword thesaurus corresponding to each search keyword of each preliminary search news is obtained by screening, which is recorded as the interest weight of each search keyword of each preliminary search news of the user, and the interest weight Q of each preliminary search news is obtained by accumulating it. j , where j represents the number of the jth preliminary retrieved news, j = 1, 2, ..., g; this can make the search results more in line with the user's interest preferences and improve the accuracy of the search results and user satisfaction.
[0061] The search keywords extracted from the two preliminary search news include "artificial intelligence", "medical applications" and "algorithm improvement" in news 1, and "artificial intelligence" and "algorithm improvement" in news 2. There are keyword lexicons for science and technology and medical news. "Artificial intelligence" and "algorithm improvement" are in the keyword lexicon for science and technology news, and "medical applications" are in the keyword lexicon for medical news. The total interest weight of users in science and technology news is 0.5, the total interest weight in transportation news is 0.2, and the total interest weight in medical news is 0.3. For news 1, because it contains three search keywords, two of which correspond to the science and technology lexicon and one corresponds to the medical lexicon, its search news interest weight is 0.5+0.5+0.3=1.3. For news 2, the two search keywords it contains are both in the science and technology lexicon, and its search news interest weight is 0.5+0.5=1.
[0062] S5. Image-text correlation analysis: Analyze the image-text consistency of the text information images and non-text information images of each user's preliminary search news to obtain the image-text correlation of each user's preliminary search news, and use this to filter out each user's searched news to obtain the image-text correlation of each user's searched news.
[0063] The specific analysis method for the text-image consistency of the text information images and non-text information images of each user's preliminary search news is as follows: the first step is to read each user's preliminary search news, obtain each image in each user's preliminary search news, and divide it into each text information image of each user's preliminary search news and each non-text information image of each user's preliminary search news; this helps to subsequently adopt different processing methods for different types of images, thereby improving the pertinence and efficiency of information processing.
[0064] In the second step, natural language processing technology is used to extract keywords from the text information images of each preliminary search of news by users, and the news categories of each preliminary search news and the news categories of each text information image of each preliminary search of news by users are obtained by analyzing the news categories of each news of interest to which users browse, thereby achieving a deeper understanding of the content of the text information images and incorporating them into the news category system, providing a basis for subsequent text-image consistency analysis.
[0065] In the third step, the news category to which each text information image of each preliminary search news of the user belongs is compared with the news category to which each preliminary search news belongs. If a text information image of a preliminary search news of the user corresponds to the news category to which the preliminary search news belongs, the image-text consistency of the text information image of the preliminary search news of the user is recorded as 1, otherwise the image-text consistency of the text information image of the preliminary search news of the user is recorded as 0. Thus, the image-text consistency of each text information image of each preliminary search news of the user is obtained by analysis, and the image-text consistency of the text information image of each preliminary search news of the user is obtained by accumulation, which is recorded as It can intuitively reflect the adaptability of text information images in news, and help evaluate the quality of news content.
[0066] In the fourth step, the pre-trained convolutional neural network model is used to classify the non-text information images of each preliminary search news by the user, and the image type of each non-text information image of each preliminary search news by the user is obtained. According to the method of analyzing the image-text consistency of the text information image of each preliminary search news by the user, the image-text consistency of the non-text information image of each preliminary search news by the user is obtained, which is recorded as It is able to process non-text information images, associate them with news content, measure the appropriateness of non-text information images in news in a unified way, and improve the comprehensive evaluation of the consistency of initial retrieved news images and texts.
[0067] The specific analysis method of the image-text relevance analysis is as follows: the first step is to read the image-text consistency of the text information image and the non-text information image of each preliminary searched news by the user. By formula Get the image and text relevance β of each user's initial search news j , where φ1 and φ2 represent the weight factors of the text-image consistency of the preset text information image and the text-image consistency of the non-text information image, respectively, and e represents a natural constant; the contribution of the text-image consistency of different types of images to the overall news text-image relationship is comprehensively considered, so that the text-image relationship can be represented by a numerical value, which is convenient for subsequent comparison and screening.
[0068] It should be noted that, in a specific embodiment, φ1 can be set to 0.7, and φ2 can be set to 0.3. In news, text information images (such as images with explanatory text, charts, etc. that are closely related to the text content) play an important role in supplementing, explaining or summarizing the text content. If the text information image has a high degree of consistency with the news text, it will greatly enhance the user's understanding of the news. Although non-text information images (such as news scene illustrations, character photos, etc. that only serve as auxiliary display) are also related to the news, their influence on the interpretation and understanding of the core content of the news is relatively weak. For example, in a news article about a meeting, even if a photo of the meeting site is not completely consistent with the text content of the news (such as the details such as the position of the characters in the photo are different from the text description), as long as it does not affect the understanding of the core event of the meeting, the influence on the user's initial judgment of whether the news is relevant is relatively small. Therefore, the weight corresponding to the consistency of the text information image is relatively high.
[0069] In the second step, the text-image relevance of each preliminary search news of the user is compared with the preset text-image relevance threshold, and all preliminary search news with text-image relevance greater than or equal to the preset text-image relevance threshold are screened out, which are recorded as each search news of the user. The text-image relevance of each search news of the user is screened out, which is recorded as β i , where i represents the number of the i-th retrieved news, i=1,2,...,n; news with low text-image relevance can be filtered out to improve the quality of news finally presented to users, ensuring that the news users see meets certain requirements in terms of text-image relevance, thereby improving the efficiency and experience of users in obtaining effective information.
[0070] S6. News popularity analysis: Obtain the popularity parameters of each retrieved news in each time range, and then analyze the news popularity of each retrieved news. The popularity parameters include page views, likes, shares, and comments.
[0071] The specific analysis method of the news heat analysis is: select different time ranges according to the settings, record them as each time range, obtain the number of times the news page of each searched news by the user in each time range from the news website, record them as the number of page views L of each searched news by the user in each time range im , where m represents the number of the mth time range, m = 1, 2, ..., q, and the number of likes, shares, and comments of each time range of each retrieved news by the user is obtained, which are recorded as D im 、F im , P im , the average page views, likes, shares, and comments of each retrieved news are obtained by average calculation, recorded as Substituting this into the formula
[0072] Get the news popularity of each searched newsγ i ,in They represent the weight factors of the preset web page views, likes, shares, and comments respectively, n represents the number of retrieved news, and q represents the number of time ranges; news popularity is an indicator that comprehensively evaluates the popularity and influence of news; it can accurately measure the overall popularity of news among the audience, which is helpful for application scenarios such as news recommendation and news value assessment.
[0073] It should be noted that, in a specific embodiment, It can be set to 0.3. Can be set to 0.1, It can be set to 0.3. It can be set to 0.3. Page views are an intuitive indicator of news popularity. A large number of views indicates that many users have viewed the news, which means that the news has wide appeal and is an important manifestation of news popularity. The number of likes represents a user's positive attitude towards the news, but it may be affected by many factors. The like behavior is relatively subjective, and there may be some users who just like it habitually or are attracted by the title without reading the news content in depth. The number of shares indicates that users think the news has sufficient value and are willing to spread it to more people. This not only reflects the user's recognition of the news, but also means that the news has the potential to spread further and can be spread on a wider range. The number of comments reflects the degree to which the news triggers user discussion. News with a large number of comments is usually controversial or has a great impact on users. If a news article has a large number of comments, it means that the news is very topical among the audience. Therefore, the corresponding weights of page views, shares, and comments are higher.
[0074] S7. News recommendation index analysis: Filter the interest weights Q of each retrieved news from the interest weights of each preliminary retrieved news i , combined with the image-text relevance β i 、News popularity γ i The news recommendation index of each retrieved news of the user is obtained by analysis, i represents the number of the i-th retrieved news, i=1,2,...,n, the news recommendation index of each retrieved news of the user is sorted and pushed to the user.
[0075] The specific analysis method of the news recommendation index analysis is: reading the interest weight of each preliminary searched news, and filtering out the interest weight of each searched news, denoted as Q i , and read the image and text relevance β of each retrieved news separately i 、News popularity γ i , substituting it into the formula Get the news recommendation index of each user's searched news Among them, w1, w2, and w3 represent the preset interest weight, image-text relevance, and news popularity weight factors respectively, and e represents a natural constant. The news recommendation index of each user's searched news is sorted in descending order, and the sorted searched news is pushed to the user; sorting by recommendation index can ensure that the user first receives the news that best meets the comprehensive evaluation criteria, improve the efficiency of users in obtaining valuable news, enhance user experience, and enable the recommendation system to more accurately meet users' news needs.
[0076] It should be noted that, in a specific embodiment, w1 can be set to 0.4, w2 can be set to 0.2, and w3 can be set to 0.4. The interest weight reflects the degree of match between news and user's personal interests. One of the core goals of the news recommendation system is to provide users with content that meets their interest preferences. If a piece of news is highly relevant to the user's interests, then the user is more likely to pay attention to it and further interact with it, even if the picture-text relevance or news popularity of this news is not very prominent. The picture-text relevance mainly reflects the completeness and accuracy of the news content. Although it has an important impact on the quality and readability of the news, compared with the interest weight and news popularity, its influence range is relatively narrow. News popularity reflects the degree of attention paid to the news by the public. Even if a piece of news is not particularly matched with the user's interests or the picture-text relevance is not optimal, a higher news popularity may also attract users to view it, because these news are often the focus of social public opinion or have a wide influence. Therefore, the weights corresponding to the interest weight and news popularity are higher.
[0077] The present invention extracts keywords corresponding to each piece of news from each type of news to form a keyword thesaurus for each type of news, constructs a user interest graph framework by analyzing the total interest weight of the user in each news category, obtains the interest weight of each retrieved news, obtains the image-text relevance of each retrieved news by screening according to the image-text relevance, obtains the news popularity of each retrieved news according to the popularity parameters of each time range of each retrieved news, and then comprehensively analyzes the news recommendation index of each retrieved news by the user, sorts and pushes them to the user, so that news can be pushed in a targeted manner, thereby improving the utilization rate of news resources.
[0078] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention and they are still covered by the protection scope of the present invention.
Claims
1. A news recommendation method based on the combination of images and texts, characterized in that: The steps include: S1. Keyword word library construction: Use web crawler technology to obtain the corresponding news of various types of news, and extract the initial keywords from them to form a keyword word library of various types of news; S2. User behavior data collection: Obtain the user's behavior data news, determine the news category to which the user's behavior data corresponds, and the behavior data news includes browsing interest news, like news, comment news, and collection news; S3. User interest graph construction: Obtain the total interest weight of users in each news category to build a user interest graph framework; S4. Retrieve news interest matching: retrieve each preliminary search news according to each search term of the user, and obtain the interest weight of each preliminary search news according to the user interest graph framework; S5. Image-text relevance analysis: Analyze the image-text consistency of the text information image and non-text information image of each preliminary searched news by the user to obtain the image-text relevance of each preliminary searched news by the user, and use this to filter out each searched news by the user to obtain the image-text relevance of each searched news by the user; S6. News popularity analysis: Obtain the popularity parameters of each retrieved news in each time range, and then analyze and obtain the news popularity of each retrieved news. The popularity parameters include page views, likes, shares, and comments. S7. News recommendation index analysis: Filter the interest weights Q of each retrieved news from the interest weights of each preliminary retrieved news i , combined with the image-text relevance β i 、News popularity γ i The news recommendation index of each retrieved news of the user is obtained by analysis, i represents the number of the i-th retrieved news, i=1,2,...,n, the news recommendation index of each retrieved news of the user is sorted and pushed to the user.
2. The news recommendation method based on the combination of pictures and texts according to claim 1, characterized in that: The specific operation method of constructing the keyword word library is as follows: Through web crawler technology, various news websites are traversed to obtain various news in various news websites, which are recorded as various news corresponding to each news. Natural language processing technology is used to analyze various news corresponding to each news, and initial keywords are extracted from them respectively, and then they are reorganized into various news keyword libraries through decomposition.
3. The news recommendation method based on the combination of pictures and texts according to claim 2 is characterized by: The specific analysis method of each user's browsing interest news is: The first step is to read the browsing history data of the user on the news platform from the management database, obtain the browsing time record of each news by the user, and obtain the browsing time of each news by the user; In the second step, by comparing the browsing time of each news item of the user with the preset browsing time threshold, the news items whose browsing time of the user exceeds the preset browsing time threshold are screened out and recorded as the browsing interest news of the user.
4. The news recommendation method based on the combination of pictures and texts according to claim 3 is characterized by: The specific analysis method for collecting the user behavior data is as follows: The first step is to extract the keywords in the titles of the news that the user browses by using text cleaning technology, and match them with the keyword word library of each type of news to obtain the number of keywords in the keyword word library of each type of news that the user browses by using text cleaning technology; The second step is to arrange the keyword numbers of each user's browsing interest news in each news keyword word library in descending order, and record the news category corresponding to the first keyword word library as the news category to which the browsing interest news belongs, thereby obtaining the news category to which each user's browsing interest news belongs; The third step is to obtain the user's likes, comments and collection records on the news platform from the news platform backend, and record the corresponding news as the user's liked news, commented news, and collected news. According to the method of analyzing the news categories to which the user's browsed interest news belongs, the news categories to which the user's liked news, commented news, and collected news belong are analyzed respectively.
5. The news recommendation method based on the combination of pictures and texts according to claim 4 is characterized by: The specific operation method of constructing the user interest graph is as follows: The first step is to read the news categories of the user's browsed news of interest, the news categories of the user's liked news, commented news, and collected news, and assign corresponding weights to the user's browsed news of interest, liked news, commented news, and collected news according to the set weights, so as to obtain the weights of the user's browsed news of interest, liked news, commented news, and collected news; The second step is to accumulate the weights of the users' browsing interest news, liking news, commenting news, and collecting news in each news category to obtain the total weight of each news category, which is recorded as the user's total interest weight in each news category. The third step is to construct a user interest graph framework using each news category as different dimensions of the graph, and mark the user's total interest weight in each news category on the corresponding dimension.
6. The news recommendation method based on the combination of pictures and texts according to claim 1, characterized in that: The specific operation method of retrieving news interest matches is: In the first step, the user's search terms are obtained by obtaining the search terms input by the user during the search, and each news item containing the user's search terms in the news website is obtained by searching for each user's search terms, which are recorded as each preliminary search news item, and each user's search terms contained in each preliminary search news item are obtained, which are recorded as each search keyword of each preliminary search news item; In the second step, each search keyword of each preliminary search news is queried in each type of news keyword thesaurus to obtain the category news keyword thesaurus corresponding to each search keyword of each preliminary search news. At the same time, the total interest weight of the user in each news category is read from the user interest graph framework, and the total interest weight of the user in the category news keyword thesaurus corresponding to each search keyword of each preliminary search news is obtained by screening, which is recorded as the interest weight of each search keyword of each preliminary search news of the user, and the interest weight Q of each preliminary search news is obtained by accumulating it. j , where j represents the number of the jth preliminary retrieved news, j = 1, 2, ..., g.
7. The news recommendation method based on the combination of pictures and texts according to claim 6, characterized in that: The specific analysis method of the text-image consistency of the text information image and the non-text information image of each preliminary searched news by the user is as follows: The first step is to read each preliminary search news of the user, obtain each image in each preliminary search news of the user, and divide it into each text information image of each preliminary search news of the user and each non-text information image of each preliminary search news of the user; The second step is to extract keywords from each text information image of each preliminary searched news by the user through natural language processing technology, and to obtain the news category of each preliminary searched news and the news category of each text information image of each preliminary searched news by the user according to the method of analyzing the news category to which each browsing news of interest belongs; In the third step, the news category to which each text information image of each preliminary search news of the user belongs is compared with the news category to which each preliminary search news belongs. If a text information image of a preliminary search news of the user corresponds to the news category to which the preliminary search news belongs, the image-text consistency of the text information image of the preliminary search news of the user is recorded as 1, otherwise the image-text consistency of the text information image of the preliminary search news of the user is recorded as 0. Thus, the image-text consistency of each text information image of each preliminary search news of the user is obtained by analysis, and the image-text consistency of the text information image of each preliminary search news of the user is obtained by accumulation, which is recorded as In the fourth step, the pre-trained convolutional neural network model is used to classify the non-text information images of each preliminary search news by the user, and the image type of each non-text information image of each preliminary search news by the user is obtained. According to the method of analyzing the image-text consistency of the text information image of each preliminary search news by the user, the image-text consistency of the non-text information image of each preliminary search news by the user is obtained, which is recorded as 8. The news recommendation method based on the combination of pictures and texts according to claim 7 is characterized by: The specific analysis method of the image-text relevance analysis is: The first step is to read the text information image and non-text information image of each preliminary search news of the user. By formula Get the image and text relevance β of each user's initial search news j , where φ1 and φ2 represent the weight factors of the preset text information image's text-image consistency and the non-text information image's text-image consistency, respectively, and e represents a natural constant; In the second step, the text-image relevance of each preliminary search news of the user is compared with the preset text-image relevance threshold, and all preliminary search news with text-image relevance greater than or equal to the preset text-image relevance threshold are screened out, which are recorded as each search news of the user. The text-image relevance of each search news of the user is screened out, which is recorded as β i , where i represents the number of the i-th retrieved news, i=1,2,...,n.
9. The news recommendation method based on the combination of pictures and texts according to claim 1, characterized in that: The specific analysis method of the news heat analysis is as follows: According to the settings, different time ranges are selected, recorded as each time range, and the number of times the news page of each searched news by the user in each time range is accessed is obtained from the news website, recorded as the number of page views L for each searched news by the user in each time range im , where m represents the number of the mth time range, m = 1, 2, ..., q, and the number of likes, shares, and comments of each time range of each retrieved news by the user is obtained, which are recorded as D im 、F im , P im , the average page views, likes, shares, and comments of each retrieved news are obtained by average calculation, recorded as Substituting this into the formula Get the news popularity of each searched newsγ i ,in They represent the weight factors of the preset page views, likes, shares, and comments respectively, n represents the number of retrieved news, and q represents the number of time ranges.
10. The news recommendation method based on the combination of pictures and texts according to claim 1, characterized in that: The specific analysis method of the news recommendation index analysis is as follows: Read the interest weights of each preliminary search news, and filter out the interest weights of each search news, denoted as Q i , and read the image and text relevance β of each retrieved news separately i 、News popularity γ i , substituting it into the formula Get the news recommendation index of each user's searched news Among them, w1, w2, and w3 represent the preset interest weight, image-text relevance, and news popularity weight factors respectively, and e represents a natural constant. The news recommendation index of each user's searched news is sorted in descending order, and the sorted searched news is pushed to the user.
Citation Information
Patent Citations
News recommendation method and news recommendation system
CN114519141A
Cited By
News data classification method and device and electronic equipment
CN120179823A
News data classification method and apparatus, and electronic device
CN120179823B
News transmission state analysis method and system
CN120578817A
Software information processing method based on big data
CN120804430A
A software information processing method based on big data
CN120804430B