A Multimodal Information Bias Removal Recommendation Method for Breaking Through the Information Cocoon
By analyzing the user's multimodal information and cognitive characteristics, combining matrix decomposition and collaborative filtering methods, information recommendation and de-biasing processing are carried out, the information cocoon problem is solved and high-quality information recommendation is achieved.
Patent Information
- Application Number
- CN202310125986.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-02-17
AI Technical Summary
How to design a multimodal information de-biased recommendation method to break through the constraints of information cocoons and realize the de-biased processing of user information recommendations.
By analyzing the user's historical behavior data, combining multimodal information (text, images, videos) and user cognitive characteristics, a matrix decomposition recommendation algorithm and collaborative filtering method are used to obtain recommendation results, and through correlation analysis and priority sorting, we recommend information with strong correlation.
It realizes the de-biasing processing of user information recommendations, breaks through the constraints of the information cocoon, provides relevant information content that is more in line with the current status of the user, and avoids information narrowing and bias in public opinion.
Smart Images

Figure CN116150487B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network space cognitive domain security research, and specifically relates to a multi-modal information debiasing recommendation method for breaking through information cocoons. Background Art
[0002] The concept of "information cocoon" originated from the book "Information Utopia" published by Cass R. Sunstein, a law professor at Harvard University, in 2006. He believes that in the dissemination of Internet information, due to the fact that the public's own information needs are not all-round, they only pay attention to selecting the things they like or the information fields that can make them happy, and are less likely to actively search for other information. Over time, the breadth and depth of the information that individuals come into contact with become more and more limited, thus imprisoning themselves in a closed space like a silkworm cocoon.
[0003] Negative assertions about the "information cocoon" such as the original sin of algorithm technology and the restrictive effect of the emergence of the "information cocoon" on the audience's active selection need to be reexamined and analyzed in depth from an objective perspective. Correctly analyzing the essence and characteristic features of the "information cocoon" and breaking the existing cognitive biases have profound theoretical significance for exploring the relationship between technology and communication in the field of communication, and have practical significance for the future governance of network information space and the guidance of network public opinion direction. Therefore, on the basis of personalized preference recommendation, introducing cognitive characteristic factors that are more relevant to the user himself and making information recommendations based on the user's cognitive characteristics can avoid the negative impact of network users being overly concentrated on a certain type of recommended information, thus leading to information narrowing, so as to break out of the "information cocoon", and further avoid the cognitive bias of public opinion caused by users concentrating on a certain public opinion topic, and correctly guide the emotions of network users. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] The technical problem to be solved by the present invention is: how to design a multi-modal information debiasing recommendation method to improve the content screening and pushing mechanism, perform debiasing processing while realizing user information recommendation, and break through the bondage of the information cocoon.
[0006] (2) Technical Solutions
[0007] To solve the above technical problems, the present invention provides a multi-modal information debiasing recommendation method for breaking through information cocoons, including the following steps:
[0008] (1) Information recommendation based on user multi-modal information: By analyzing the historical behavior data of users, analyzing and representing the text information, image information, and video information data that users have searched and browsed, capturing the preferences of users, and completing the personalized modeling of users, so as to actively recommend information to different user groups;
[0009] (2) Information recommendation integrating user cognitive characteristics: Combining the user's historical behavior data, analyzing the user's cognitive characteristics through data mining algorithms, and proposing a matrix factorization recommendation algorithm integrating user cognitive characteristics. It uses matrix factorization technology and collaborative filtering methods to obtain recommendation results, thereby recommending relevant information content that better suits the user's current state to the user;
[0010] (3) Bias-removing recommendation: Conduct a relevance analysis on the recommended content in steps 1 and 2, and prioritize the relevance of the recommended content. Give priority to recommending information with a relevance greater than the preset threshold to the user.
[0011] Preferably, in step (1), analyze the multi-modal information in the user's search, and the multi-modal information includes text, image, and video information.
[0012] Preferably, in step (1), encode the user's multi-modal information exchange into the representation learning of each modality, and predict each user's preference by fusing the representation learning of multiple modalities.
[0013] Preferably, the specific implementation steps of step (1) are as follows:
[0014] 1) For the user nodes in the bipartite graph G m , the aggregation layer uses the aggregation function f(·) to quantify the influence of its neighbors and output, and the calculation method is shown in the following formula:
[0015] h m =f(N u ) (1)
[0016] where m ∈ M = {t, i, v} represents the modality, t, i, and v represent the text, image, and video modalities respectively, N u represents the neighbor of user u, and f(·) represents the aggregation function;
[0017] 2) Adopt a new combination layer to fuse three pieces of information. It integrates the structural information h m , the intrinsic information u m and the modality connection u id into a unified representation form, which is specifically expressed as:
[0018]
[0019] where g() represents the combination function;
[0020] 3) Collect the information propagated by the l-hop neighbors through the modality m to simulate the user's exploration process. Formally, the l-hop neighbor representation of user u and the output recursive representation of the l-th multi-modal combination layer are:
[0021]
[0022] That is, it is the representation learning of the modality; among them, is the representation generated by the previous layer, recording the information from the l-1 hop neighbors, which is set to u in the initial iteration m ; among them, the user u is associated with the trainable vector u m , associated; describes the user's preference for the content information features in the modality m, and considers the influence of the modality interaction reflecting the potential relationship between modalities;
[0023] After stacking L single-modal aggregation layers and a multi-modal combination layer, through the linear combination of the multi-modal representations, the final representations of the user u and the predicted text information i are obtained:
[0024]
[0025] Preferably, in step (2), the user's cognitive characteristics are incorporated into the information prediction score, and it is assumed that each cognitive characteristic has the same influence on the content information of the same category. Among them, matrix factorization refers to decomposing a matrix into the product of two or more matrices, and the decomposed matrices are used in the recommendation calculation. The idea of the collaborative filtering method is to discover the user's interests and hobbies based on the mining of the user's historical behavior data, divide the users based on different interests and hobbies, and recommend similar-interest products to the users.
[0026] Preferably, in step (2), let b c be the cognitive bias, representing the influence of the cognitive characteristic on a type of predicted content:
[0027]
[0028] where w is the category to which the content belongs, represents the influence of the cognitive characteristic j on the content information of the w category, and there are a total of k cognitive characteristics; then the score prediction model incorporating cognitive characteristics is as follows:
[0029]
[0030] where each user has a user embedding, that is, p i , P is the set of all users; each type of information has an embedding, that is, q j , Q is the set of all content; P T Q represents the interest scores of all users for all content;
[0031] Let the objective function of the score prediction model be:
[0032]
[0033] Among them, r ij is the true score of user i for content j, indicating the degree of interest of user i in content j, b i is the user bias score, b j is the content information category bias score; the b i value indicates whether this user's score for this type of content is high or low, and the b j value indicates whether all users' scores for this content information are high or low, and λ is a preset constant; to optimize the objective function F, the stochastic gradient descent method is used for optimization training:
[0034]
[0035]
[0036] Among them,
[0037] γ is the update step size. Equations (8) to (12) are the calculation formulas for the negative gradient direction, and equations (13) to (17) are the parameter update formulas. Substitute the learned correct parameters into the scoring prediction model of formula 6 to obtain the predicted scoring matrix of users;
[0038] Use the established scoring prediction matrix to predict scores, and perform a centering operation on the predicted score results as follows:
[0039]
[0040] Among them, represents the predicted score result of the user for the content information, μ i represents the average score value of user i, μ j represents the average score value of content information j;
[0041] Then use the following cosine similarity formula to calculate the similarity between users, and find the nearest neighbor user set of the target user u in the centered scoring matrix;
[0042]
[0043] Among them, represents the predicted score of the target user u for content j;
[0044] Through calculation, obtain the similar user matrix, and select N users from it to form the nearest neighbor set. v represents the neighbor user of the target user, v ∈ N; finally, calculate the predicted score of the target user u for the content information j according to the nearest neighbor set of the target user, that is, the recommended result, and the formula is as follows:
[0045]
[0046] Among them, represents the predicted score of neighbor user v for content j, represents the predicted average score of user v for content j, represents the predicted average score of target user u for content j.
[0047] Preferably, in step (3), the relevance calculation is performed on the information recommended in step (1) and step (2), and the Pearson correlation coefficient is used to rank the relevance of the recommended content in priority.
[0048] The present invention also provides a system for implementing the above method.
[0049] The present invention also provides an application of the above method in the technical field of network space cognitive domain security research.
[0050] The present invention also provides an application of the above system in the technical field of network space cognitive domain security research.
[0051] (III) Beneficial effects
[0052] The present invention is implemented through three parts: a precise information recommendation algorithm based on multimodal information, an information recommendation algorithm based on user cognitive characteristics, and debiasing processing. To provide high-quality information recommendation services, it is crucial to consider the interaction between users and the network and the content from various modalities (such as text, images, videos, etc.). Therefore, a multimodal graph convolutional network framework is designed to better capture user preferences and then better recommend content information to users; the present invention also integrates user cognitive characteristics into the matrix factorization model and proposes a matrix factorization recommendation algorithm that fuses user cognitive characteristics. This algorithm uses user cognitive characteristics to predict the interest value of users in the recommended content and adds user cognitive characteristic regularization to the matrix factorization model, thereby obtaining more targeted recommendation information; during the debiasing processing, the present invention also performs a relevance analysis on the recommended content information, that is, the Pearson correlation coefficient can be used for analysis, and the relevance ranking is performed to preferentially recommend content information with stronger relevance. While fully considering user interest preferences, user cognitive characteristics are considered as recommendation factors, thereby avoiding the limitation that the algorithm only recommends based on the user's search content. User cognitive characteristics (such as knowledge background and cognitive ability) can effectively expand the recommendation scope and break through the "information cocoon". Description of the drawings
[0053] Figure 1 is the design schematic diagram of the present invention. Detailed implementation manners
[0054] To make the objectives, content, and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention with reference to the accompanying drawings and embodiments.
[0055] The embodiment of the present invention provides a multi-modal information de-biasing recommendation method for breaking through the information cocoon. In this embodiment, first, the traditional information recommendation algorithm is improved. Based on the user's multi-modal information, user personalization modeling is performed, and the user is characterized in multiple aspects, so as to more accurately capture the user's interest preferences and recommend information to the user. Secondly, based on the traditional recommendation algorithm, the user's cognitive characteristics are mined and analyzed, and a user information recommendation method based on the user's cognitive characteristics is proposed, and then content information that better fits the user's current cognitive state is recommended. Finally, combining the recommended content from the two aspects, correlation analysis is performed on it, and the recommended information with a high correlation coefficient is the information that not only meets the user's preferences but also has the user's cognitive characteristics, and is given higher priority to be recommended to the user, thereby breaking through the information cocoon constraint caused by only relying on the user's historical search content to recommend information.
[0056] Specifically, referring to Figure 1 , a multi-modal information de-biasing recommendation method for breaking through the information cocoon provided by the embodiment of the present invention includes the following steps:
[0057] (1) Precise information recommendation based on user multi-modal information: In this step, by analyzing the user's historical behavior data, the text information, image information, and video information data browsed by the user are analyzed and represented to more completely capture the user's preferences and interests, complete the user's personalization modeling, and thus achieve the purpose of actively and precisely recommending information to different user groups.
[0058] (2) Information recommendation integrating user cognitive characteristics: Combining the user's historical behavior data, the user's cognitive characteristics are analyzed through data mining algorithms, and a matrix factorization recommendation algorithm integrating user cognitive characteristics is proposed. Matrix factorization technology and collaborative filtering methods are used to obtain the recommendation results, so as to recommend relevant information content that better fits the user's current state to the user.
[0059] (3) De-biasing recommendation: Perform correlation analysis on the recommended content in steps (1) and (2), and use the Pearson correlation coefficient to rank the relevance of the recommended content in order of priority. Information with a stronger correlation is preferentially recommended to the user. In this step, on the premise of recommending information based on the user's search content, information recommended based on the user's cognitive characteristics is added, integrating the user's current cognitive state, and breaking the "information cocoon" effect caused by only recommending information based on the user's search content.
[0060] In step (1), the precise information recommendation mainly analyzes the multi-modal information (including text, images, and videos) in the user's search, and then enhances the user representation to better capture the user's preferences and implement precise recommendation strategies for different user groups. The details are as follows:
[0061] This step (1) is implemented by three components: an aggregation layer, a combination layer, and a prediction layer. By stacking multiple aggregation layers and combination layers, the multi-modal information exchange of the user is encoded into the representation learning of each modality, where the multi-modal refers to the three modalities of text, images, and videos. Finally, the prediction layer predicts the preferences of each user by fusing the representation learning of multiple modalities. The specific implementation steps are as follows:
[0062] 1) Inspired by the GCN message passing mechanism, for the user nodes in the bipartite graph G m , the aggregation layer uses the aggregation function f(·) to quantify the influence of its neighbors and output, and the calculation method is shown in the following formula:
[0063] h m = f(N u ) (1)
[0064] where m ∈ M = {t, i, v}, representing the modality indicator, and t, i, v represent the textual, image, and visual modalities respectively, N u represents the neighbors of user u, and f(·) represents the aggregation function.
[0065] 2) A new combination layer is adopted to fuse three pieces of information. It integrates the structural information h m , the intrinsic information u m and the modality connection u id into a unified representation form, which is specifically expressed as:
[0066]
[0067] g() is the combination function.
[0068] 3) The information propagated by the l-hop neighbors is collected through the modality m to simulate the user's exploration process. Formally, the l-hop neighbor representation of user u and the output recursive representation of the l-th multi-modal combination layer are:
[0069]
[0070] That is the representation learning of the modality. Among them, is the representation generated by the previous layer, recording the information from the (l - 1)-hop neighbors, is set to u m in the initial iteration; among them, user u and the trainable vector um , is associated, and this vector is randomly initialized; In modality m, the user's preference for the content information features is described, and the influence of modality interaction reflecting the potential relationship between modalities is considered, so as to better realize the prediction function.
[0071] After stacking L single-modal aggregation layers and a multi-modal combination layer, the prediction layer obtains the final representation of user u and the predicted text information i through a linear combination of multi-modal representations:
[0072]
[0073] In step (2), referring to the idea proposed by Baltrunas in the article, the user's cognitive characteristics are incorporated into the information prediction score, and it is assumed that each cognitive characteristic has the same influence on the content information of the same category. In this step, matrix factorization refers to decomposing a matrix into the product of two or more matrices. In actual recommendation calculation, the large matrix is no longer used, but two small matrices obtained by decomposition are used. The following process of matrix deduction is the process of matrix factorization. The basic idea of the collaborative filtering method is to discover the user's interests and hobbies based on the mining of the user's historical behavior data, divide the users based on different interests and hobbies, and recommend similar-interest products to the users. Below, the collaborative filtering is reflected in incorporating the user's cognitive characteristics into the information prediction score.
[0074] Let b c be the cognitive bias, representing the influence of cognitive characteristics on a certain type of predicted content:
[0075]
[0076] where w is the category to which the content belongs, represents the influence of cognitive characteristic j on the content information of category w, and there are a total of k cognitive characteristics; then the score prediction model incorporating cognitive characteristics is as follows:
[0077]
[0078] where each user has a user embedding, which is p i , and P is the set of all users; each type of information has an embedding, which is q j , and Q is the set of all contents. represents the degree of interest of user i in content j, and P T Q represents the interest scores of all users in all contents.
[0079] The objective function of the score prediction model is:
[0080]
[0081] Among them, r ij is the true score of user i for content j, represents the degree of interest of user i in content j, b i is the user bias score, b j is the content information category bias score; b i value indicates whether the score of this user for this type of content is generally high or low, b j value indicates whether the scores of all users for this content information are generally high or low, and λ represents a constant. To optimize the objective function F, the stochastic gradient descent method is used for optimization training:
[0082]
[0083] Among them,
[0084] γ is the update step size. Equations (8) to (12) are the negative gradient direction calculation formulas, and equations (13) to (17) are the parameter update formulas. Substituting the learned correct parameters into the score prediction model of formula 6 can obtain the user's predicted score matrix.
[0085] Use the score prediction model established above to predict the score, and perform the centering operation on the predicted score results as follows:
[0086]
[0087] Among them, represents the predicted score result of the user for the content information, μ i represents the average score value of user i, μ j represents the average score value of content information j.
[0088] Then use the following cosine similarity formula to calculate the similarity between users, and find the nearest neighbor user set of the target user u in the centered score matrix.
[0089]
[0090] Among them, represents the predicted score of the target user u for content j.
[0091] Through the above formula calculation, the similar user matrix is obtained, and N users are selected from it to form the nearest neighbor set. v represents the neighbor user of the target user, v ∈ N; finally, according to the nearest neighbor set of the target user, calculate the predicted score of the target user u for the content information j, that is, the recommended result, and the formula is as follows:
[0092]
[0093] Among them, represents the predicted score of neighbor user v for content j, represents the predicted average score of user v for content j, represents the predicted average score of target user u for content j.
[0094] In step (3), the relevance calculation is performed on the information recommended in steps (1) and (2). For example, the Pearson correlation coefficient is used for analysis, and the relevance of the recommended information is sorted. The information with stronger relevance is preferentially recommended. On the basis of the traditional recommendation algorithm, an information recommendation algorithm based on cognitive characteristics is added, so that the recommended information is no longer single, breaking through the information cocoon bondage brought by the traditional algorithm.
[0095] The above is only the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. A multi-modal information de-biasing recommendation method for breaking through the information cocoon, characterized in that, It includes the following steps: (1) Information recommendation based on user multimodal information: By analyzing the historical behavior data of users, analyze and represent the text information, image information, and video information data that users have searched and browsed, capture the preferences of users, complete the personalized modeling of users, and thus realize the active recommendation of information for different user groups; (2) Information recommendation integrating user cognitive characteristics: Combine the historical behavior data of users, analyze the cognitive characteristics of users through data mining algorithms, and propose a matrix factorization recommendation algorithm integrating user cognitive characteristics. It uses matrix factorization technology and collaborative filtering methods to obtain recommendation results, so as to recommend relevant information content that better matches the current state of users for users; (3) Bias removal recommendation: Conduct a correlation analysis on the recommendation content in steps 1 and 2, and prioritize the correlation of the recommendation content. Give priority to recommending information with a correlation greater than a preset threshold to users; The specific implementation steps of step (1) are as follows: 1) For the user nodes in the bipartite graph G m in the aggregation layer, the aggregation function f(·) is used to quantify the influence of its neighbors and output, and the calculation method is shown as follows: h m = f(N u ) (1) where, m ∈ M = {t, i, v}, representing modality, and t, i, v respectively represent text, image, and video modalities, N u represents the neighbors of user u, and f(·) represents the aggregation function; 2) A new combined layer is adopted to fuse three pieces of information, which combines the structural information h m , the intrinsic information u m and the modal connection u id into a unified representation form, specifically expressed as: where g() represents a combination function; 3) Collect the information propagated by l-hop neighbors through modality m to simulate the exploration process of users. Formally, the l-hop neighbor representation of user u and the output recursive representation of the l-th multimodal combination layer are: Namely, it is the representation learning of modalities; among which, is the representation generated by the previous layer, recording the information from the l-1 hop neighbors, which is set to u in the initial iteration m ; among which, the user u is associated with the trainable vector u m , is associated; describes the user's preference for the content information features in modality m and takes into account the influence of modality interaction reflecting the potential relationship between modalities; After stacking L single-modal aggregation layers and multimodal combination layers, through the linear combination of multimodal representations, the final representations of user u and the predicted text information i are obtained:
2. The method according to claim 1, characterized in that, In step (1), analyze the multimodal information in the user search. The multimodal information includes text, image, and video information.
3. The method according to claim 2, characterized in that, In step (1), exchange and encode the multimodal information of users into the representation learning of each modality, and predict the preferences of each user by fusing the representation learning of multiple modalities.
4. The method according to claim 1, characterized in that, In step (2), integrate the user cognitive characteristics into the information prediction score, and assume that each cognitive characteristic has the same impact on the content information of the same category. Among them, matrix factorization refers to decomposing a matrix into the product of two or more matrices, and the decomposed matrices are used in the recommendation calculation. The idea of the collaborative filtering method is to discover the interests and hobbies of users based on the mining of the historical behavior data of users, divide users based on different interests and hobbies, and recommend similar-interest products to users.
5. The method according to claim 4, characterized in that, In step (2), let b c be the cognitive bias, representing the influence of cognitive characteristics on a type of prediction content: where w is the category to which the content belongs, represents the influence of the cognitive characteristic j on the content information of category w, and there are k cognitive characteristics in total; then the scoring prediction model incorporating cognitive characteristics is as follows: Among them, each user has a user embedding, which is p i , where P is the set of all users; each type of information has an embedding, which is q j , where Q is the set of all content; P T Q represents the interest scores of all users for all content; Let the objective function of the rating prediction model be: Among them, r ij is the true score of user i for content j, indicating the degree of interest of user i in content j, b i is the user bias score, b j is the content information category bias score; b i value indicates whether this user's score for this type of content is high or low, b j value indicates whether all users' scores for this content information are high or low, and λ is a preset constant; to optimize the objective function F, the stochastic gradient descent method is used for optimization training: Among them, γ is the update step size. Formulas (8) to (12) are the calculation formulas for the negative gradient direction, and formulas (13) to (17) are the parameter update formulas. Substitute the learned correct parameters into the rating prediction model of formula 6 to obtain the predicted rating matrix of users; Use the established rating prediction matrix to predict the ratings, and perform a centering operation on the predicted rating results as follows: Among them, represents the predicted scoring result of the user on the content information, μ i represents the average scoring value of user i, μ j represents the average scoring value of content information j; Then use the following cosine similarity formula to calculate the similarity between users, and find the nearest neighbor user set of the target user u in the centered rating matrix; Among them, represents the predicted score of the target user u for the content j; Through calculation, obtain the similar user matrix, and select N users from it to form the nearest neighbor set. v represents the neighbor user of the target user, v ∈ N; Finally, calculate the predicted rating of the target user u for the content information j according to the nearest neighbor set of the target user, that is, the said recommendation result. The formula is as follows: Among them, represents the predicted score of neighbor user v for content j, represents the predicted average score of user v for content j, represents the predicted average score of target user u for content j.
6. The method according to claim 1, characterized in that, In step (3), the relevance of the information recommended in steps (1) and (2) is calculated, and the Pearson correlation coefficient is used to prioritize the relevance of the recommended content.
7. A system for implementing the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Trust degree and metric factor matrix decomposition fused interest point recommendation method and system
CN110955829A
Efficient non-sampling graph convolutional network recommendation method
CN114282122A