Graphic and text information browsing and pushing method based on user preference setting
By constructing a multi-dimensional user behavior feature matrix and designing a dual matching algorithm, combined with a dynamic push sequence generation mechanism, the problems of inaccurate user preference modeling, single content matching measurement and solidified push strategy in the existing technology are solved, and more accurate user preference modeling and content matching are achieved, improving user experience and push effect.
Patent Information
- Application Number
- CN202510136530.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
The existing graphics and text information push technology has shortcomings in user preference modeling, content matching and dynamic adjustment of push strategies, resulting in poor user experience and push effects.
By constructing a multi-dimensional user behavior feature matrix, designing a double matching algorithm (combination of cosine similarity and deep semantic matching) and introducing a dynamic push sequence generation mechanism, the accuracy of user preference modeling and the accuracy of content matching are achieved.
It improves the accuracy of user preference modeling, improves the accuracy of content matching, dynamically adjusts push strategies, ensures the diversity and timeliness of pushed content, thereby improving user experience and push effects.
Smart Images

Figure CN120067443A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information push, and particularly to a method for browsing and pushing graphic and text information based on user preference settings. Background Art
[0002] Traditional graphic and text information push technologies mainly rely on collaborative filtering algorithms and content-based recommendation algorithms, and perform similarity calculation and matching recommendation by analyzing user historical behavior data and item feature information. However, these methods often overly rely on explicit user feedback data, such as clicks, favorites and other behaviors, while ignoring the rich implicit feedback information generated by users during browsing, such as dwell time, browsing trajectory, etc. At the same time, traditional push algorithms have problems such as insufficient feature extraction and in-depth semantic understanding when processing multi-modal information, and it is difficult to accurately capture the dynamic interest changes and deep preference features of users.
[0003] The existing graphic and text information push technologies also face the following technical difficulties in practical applications: First, the representation method of user behavior characteristics is relatively single, and it is difficult to comprehensively depict the interaction mode between users and content; Second, multi-dimensional behavior characteristics are not effectively integrated in the user interest modeling process, resulting in one-sided results of user preference analysis; Third, the matching algorithms for push content generally adopt a single similarity measurement method, and cannot take into account both the surface feature similarity and the deep semantic relevance of the content at the same time; Finally, the construction of the push sequence lacks a dynamic adjustment mechanism, and it is difficult to ensure the diversity and timeliness of the push content. These technical problems seriously affect the user experience and push effect of the graphic and text information push service.
[0004] Based on the above problems, the present invention provides a method for browsing and pushing graphic and text information based on user preference settings. This method effectively solves the technical problems in the prior art, such as inaccurate user preference modeling, single content matching measurement, and fixed push strategies, by constructing a multi-dimensional user behavior feature matrix, designing a dual matching algorithm, and a dynamic push sequence generation mechanism. Summary of the Invention
[0005] In view of the problems existing in the existing graphic and text information push technologies, such as single representation of user behavior characteristics, incomplete user interest modeling, simple content matching algorithms, and lack of dynamic adjustment of the push sequence, the present invention is proposed.
[0006] Therefore, the problem to be solved by the present invention is how to accurately model user interests, improve content matching degree, and dynamically adjust push strategies by comprehensively considering multi-dimensional user behavior characteristics, designing a dual matching algorithm, and introducing a dynamic push sequence generation mechanism, so as to improve the deficiencies of traditional push algorithms in terms of personalization, accuracy, and timeliness, and enhance the user experience and push effect.
[0007] To solve the above technical problems, the present invention provides the following technical solutions:
[0008] In a first aspect, an embodiment of the present invention provides a method for browsing and pushing graphic and text information based on user preference settings, which includes obtaining interaction data when the user is currently browsing graphic and text information, and establishing a user behavior feature matrix; calculating the user's interest in the current graphic and text information through the feature matrix and a preset weight vector, and performing clustering analysis on the graphic and text information in the user's historical browsing records to generate a set of user preference categories; based on the set of user preference categories, using a dual matching algorithm that combines cosine similarity and deep semantic matching, screening out candidate graphic and text information with a matching score exceeding a second threshold from the graphic and text information library to form a push candidate set; sorting the graphic and text information in the push candidate set in ascending order according to the matching degree with the set of user preference categories, and selecting graphic and text information to construct a push sequence, and pushing the push sequence to the user terminal.
[0009] As a preferred solution of the method for browsing and pushing graphic and text information based on user preference settings of the present invention, wherein: sorting the graphic and text information in the push candidate set in descending order according to the matching degree with the set of user preference categories, and selecting graphic and text information to construct a push sequence, and pushing the push sequence to the user terminal, includes: obtaining the comprehensive matching score and the corresponding category attribution information of the graphic and text information in the push candidate set; based on the interest threshold of each category in the set of user preference categories, performing category weighting adjustment on the comprehensive matching score to obtain a sorting score; sorting the graphic and text information in the push candidate set in ascending order according to the sorting score to generate an initial push sequence; selecting graphic and text information from the initial push sequence according to a preset time window size to construct a push sequence, wherein the category identifiers of adjacent graphic and text information in the push sequence are not repeated; encapsulating the push sequence into a push message, and sending the push message to the user terminal through the push interface of the user terminal, wherein the push message includes the title, abstract, thumbnail and category identifier of the graphic and text information.
[0010] As a preferred solution of the graphic and text information browsing and pushing method based on user preference settings according to the present invention, wherein: the method for forming the push candidate set is as follows: based on the user preference category set, extract the text feature vector through the TF-IDF algorithm, and use the pre-trained image feature extraction model to obtain the picture feature vector; splice the text feature vector and the picture feature vector to form a hybrid feature vector of graphic and text information, calculate the cosine similarity between the hybrid feature vector and the feature vectors of each category in the user preference category set to obtain the surface matching score; at the same time, adopt a deep semantic matching model based on the Transformer architecture to calculate the semantic-level similarity between the graphic and text information content and the sample content of each category in the user preference category set to obtain the deep matching score, wherein the deep semantic matching model captures the context correlation of the content through the attention mechanism; sort the graphic and text information in descending order according to the deep matching score, and select the top N pieces of graphic and text information to form the push candidate set; perform weighted fusion on the surface matching score and the deep matching score to obtain the matching degree score; when the matching degree score of a certain piece of graphic and text information and any category in the user preference category set exceeds the second threshold, add this piece of graphic and text information to the push candidate set.
[0011] As a preferred solution of the graphic and text information browsing and pushing method based on user preference settings according to the present invention, wherein: the specific formula of the deep semantic matching model is as follows:
[0012]
[0013] Wherein, S deep (Q, D) is the matching degree score, Q is the query text feature vector, D is the target text feature vector, n is the feature dimension, w i is the weight coefficient of the i-th dimension feature, α is the fusion parameter, H(·) is the semantic encoding function, and σ is the Gaussian kernel parameter.
[0014] As a preferred embodiment of the method for browsing and pushing graphic and text information based on user preferences according to the present invention, wherein: the method for generating the set of user preference categories is to normalize the feature matrix and set a preset weight vector, where the components in the preset weight vector correspond to the weight coefficients of the interaction data; perform a dot product operation on the interaction feature vector of the normalized feature matrix and the preset weight vector to obtain the user's interest degree in the graphic and text information; based on the interest degree, use an improved K-means algorithm to perform topic clustering on the graphic and text information in the user's historical browsing records, and extract the topic feature vectors of the graphic and text information through a text topic model; perform clustering division through the cosine similarity of the topic feature vectors, determine the topic categories, and extract the topic features of each category to form category feature vectors; calculate the mapping relationship between the category feature vectors and the predefined topic system, count the number and interest degree of the graphic and text information associated with the semantic categories, and generate category weights; sort the semantic categories according to the category weights, and the categories whose category weights exceed the first threshold form the set of user preference categories.
[0015] As a preferred embodiment of the method for browsing and pushing graphic and text information based on user preferences according to the present invention, wherein: the method for establishing the feature matrix is to normalize the interaction data and combine them into behavior feature vectors, and arrange them in the order of the identification ID of the graphic and text information to form a user behavior feature matrix, where the row vectors of the feature matrix represent the user interaction data of the graphic and text information, and the column vectors represent the interaction feature dimensions.
[0016] As a preferred embodiment of the method for browsing and pushing graphic and text information based on user preferences according to the present invention, wherein: the interaction data includes the residence duration, click frequency, favorite mark, sharing behavior, and semantic score of the comment content; the residence duration is to record the start time point and end time point of the user's residence on the graphic and text information page, and obtain the residence duration by calculating the time difference between the start time point and the end time point; the click frequency is to count the number of clicks of the user on the interaction elements in the graphic and text information within a specified time window, and calculate the click frequency per unit time in combination with the duration of the time window; the favorite mark is to record the user's favorite operation on the graphic and text information, and set the favorite mark as a binary feature, where the favorite state is recorded as 1 and the non-favorite state is recorded as 0; the sharing behavior is to detect the behavior of the user sharing the graphic and text information to different social platforms, record the type of the shared target platform and the number of shares, and calculate the comprehensive sharing index according to the weight coefficients of different social platforms; the semantic score of the comment content is to perform sentiment analysis on the comment content published by the user using a natural language processing model, extract the sentiment polarity and sentiment intensity of the comment, and calculate the semantic comprehensive score of the comment content in combination with the number of comment words and the comment time.
[0017] In a second aspect, an embodiment of the present invention provides a graphic and text information browsing and pushing system based on user preference settings, which includes: an acquisition module for acquiring interaction data when a user browses graphic and text information currently and establishing a user behavior feature matrix; a clustering analysis module for calculating the user's interest in the current graphic and text information through the feature matrix and a preset weight vector, and performing clustering analysis on the graphic and text information in the user's historical browsing records to generate a set of user preference categories; a screening module, based on the set of user preference categories, using a dual matching algorithm that combines cosine similarity and deep semantic matching to screen out candidate graphic and text information with a matching score exceeding a second threshold from a graphic and text information library to form a push candidate set; a push module for sorting the graphic and text information in the push candidate set in ascending order according to the matching degree with the set of user preference categories, and selecting graphic and text information to construct a push sequence and pushing the push sequence to a user terminal.
[0018] In a third aspect, an embodiment of the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program instructions are executed by the processor, the steps of the graphic and text information browsing and pushing method based on user preference settings as described in the first aspect of the present invention are implemented.
[0019] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program instructions are executed by the processor, the steps of the graphic and text information browsing and pushing method based on user preference settings as described in the first aspect of the present invention are implemented.
[0020] The beneficial effects of the present invention are as follows: By constructing a multi-dimensional user behavior feature matrix, a three-dimensional representation of the user's browsing behavior is realized, avoiding the problem of serious information loss in traditional feature extraction methods; adopting feature matrix normalization processing and an improved K-means clustering algorithm, combined with the extraction of feature vectors of a text topic model, accurately quantifying the user's interest distribution and improving the accuracy of user preference modeling; introducing a dual matching algorithm that combines cosine similarity and deep semantic matching based on the Transformer architecture, and through the collaborative calculation of surface features and deep semantics, improving the accuracy of content matching; through a category-weighted sorting mechanism and a time window control strategy, realizing the dynamic optimization of the push sequence and ensuring the diversity and timeliness of the pushed content. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings. Among them:
[0022] Figure 1 It is a flowchart of the graphic and text information browsing and pushing method set based on user preferences in Embodiment 1. Specific implementation manners
[0023] To make the above objects, features and advantages of the present invention more obvious and understandable, the specific implementation manners of the present invention will be described in detail below with reference to the accompanying drawings of the specification.
[0024] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0025] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other from other embodiments.
[0026] Embodiment 1
[0027] Referring to Figure 1 , which is the first embodiment of the present invention. This embodiment provides a graphic and text information browsing and pushing method set based on user preferences, including
[0028] S1: Obtain the interaction data when the user is currently browsing graphic and text information, and establish a user behavior feature matrix.
[0029] Specifically, the interaction data includes the residence duration, click frequency, favorite mark, sharing behavior, and semantic score of the comment content.
[0030] Further, the residence duration records the start time point and end time point of the user's stay on the graphic and text information page, and the residence duration is obtained by calculating the time difference between the start time point and the end time point; the click frequency is to count the number of clicks of the user on the interaction elements in the graphic and text information within a specified time window, and calculate the click frequency per unit time in combination with the duration of the time window; the favorite mark records the favorite operation of the user on the graphic and text information, and the favorite mark is set as a binary feature, where the favorite state is recorded as 1 and the non-favorite state is recorded as 0; the sharing behavior is to detect the behavior of the user sharing the graphic and text information to different social platforms, record the type of the shared target platform and the number of shares, and calculate the comprehensive sharing index according to the weight coefficients of different social platforms; the semantic score of the comment content is to perform sentiment analysis on the comment content published by the user using a natural language processing model, extract the sentiment polarity and sentiment intensity of the comment, and calculate the semantic comprehensive score of the comment content in combination with the number of comment words and the comment time.
[0031] It should be noted that the construction method of the natural language processing model is as follows: the BERT pre-trained natural language processing model is used to encode the review text to obtain the context semantic representation of the text; a multi-layer perceptron network is connected to extract the sentiment features, including a sentiment classification layer and a sentiment intensity regression layer; the cross-entropy loss function is used to optimize the sentiment polarity classification task, and the mean squared error loss function is used to optimize the sentiment intensity prediction task, and joint training is carried out through the method of multi-task learning; the predicted sentiment polarity and sentiment intensity scores are combined with the review word count and review time features, and the semantic comprehensive score of the review content is calculated by weighted summation.
[0032] Furthermore, the method for establishing the feature matrix is as follows: the interaction data is normalized and combined into a behavioral feature vector, and arranged in the order of the identification ID of the graphic and text information to form a user behavior feature matrix, where the row vector of the feature matrix represents the user interaction data of the graphic and text information, and the column vector represents the interaction feature dimension.
[0033] S2: Calculate the user's interest in the current graphic and text information through the feature matrix and the preset weight vector, and perform clustering analysis on the graphic and text information in the user's historical browsing records to generate a set of user preference categories.
[0034] Specifically, the method for generating the set of user preference categories is as follows: the feature matrix is normalized, and a preset weight vector is set, where the components in the preset weight vector W correspond to the weight coefficients of the interaction data; the dot product operation is performed on the interaction feature vector of the normalized feature matrix and the preset weight vector to obtain the user's interest in the graphic and text information.
[0035] Further, the specific formula for the interest degree is as follows:
[0036]
[0037] where I is the user's interest in the graphic and text information, A ′ is the interaction feature vector in the normalized feature matrix, W is the preset weight vector, a' i is the i-th normalized interaction feature value, δ i is the weight coefficient corresponding to the i-th interaction feature, and m is the number of dimensions of the interaction feature.
[0038] Furthermore, based on the interest degree, an improved K-means algorithm is used to perform topic clustering on the graphic and text information in the user's historical browsing records, and the topic feature vector of the graphic and text information is extracted through the text topic model; clustering division is performed through the cosine similarity of the topic feature vectors to determine the topic categories, and the topic features of each category are extracted to form the category feature vectors.
[0039] It should be noted that the method for constructing the text topic model is to perform word segmentation and preprocessing on the text and image content, and establish a term frequency-inverse document frequency (TF-IDF) feature matrix; through the Gibbs sampling method, the document-topic distribution and topic-term distribution are iteratively optimized to obtain the topic distribution representation of the document; an attention mechanism is introduced to capture the semantic associations between words, and combined with the context semantic features extracted by the natural language processing model, a topic feature vector integrating global semantic information is generated.
[0040] Specifically, calculate the mapping relationship between the category feature vector and the predefined topic system, count the number and interest degree of text and image information associated with semantic categories, and generate category weights; sort the semantic categories according to the category weights, and the categories whose category weights exceed the first threshold form the user preference category set.
[0041] It should be noted that the first threshold is a dynamic threshold optimized based on the clustering analysis of the historical user preference category distribution and the ROC curve of the interaction behavior.
[0042] S3: Based on the user preference category set, a dual matching algorithm combining cosine similarity and deep semantic matching is used to screen out candidate text and image information with a matching score exceeding the second threshold from the text and image information library to form a push candidate set.
[0043] Specifically, the method for forming the push candidate set is to extract the text feature vector through the TF-IDF algorithm based on the user preference category set, and use the pre-trained image feature extraction model to obtain the image feature vector.
[0044] It should be noted that the method for constructing the image feature extraction model is to adopt a deep convolutional neural network architecture based on ResNet-50. After pre-training on a large-scale image data set, multi-level visual features are extracted. The low-level features of the image are extracted through the convolutional layer and the pooling layer; the high-level semantic features are extracted through the residual connection and multi-layer stacking; the feature map is compressed into a feature vector with a fixed dimension through the global average pooling layer.
[0045] Furthermore, the text feature vector and the image feature vector are concatenated to form a hybrid feature vector of the text and image information, and the cosine similarity between the hybrid feature vector and the category feature vectors in the user preference category set is calculated to obtain the surface layer matching score; at the same time, a deep semantic matching model based on the Transformer architecture is used to calculate the semantic-level similarity between the text and image information content and the sample content of each category in the user preference category set to obtain the deep layer matching score, where the deep semantic matching model captures the context associations of the content through the attention mechanism.
[0046] Even further, the specific formula of the deep semantic matching model is as follows:
[0047]
[0048] Among them, S deep (Q, D) is the matching degree score, Q is the query text feature vector, D is the target text feature vector, n is the feature dimension, w i is the weight coefficient of the i-th dimensional feature, α is the fusion parameter, H(·) is the semantic encoding function, and σ is the Gaussian kernel parameter.
[0049] Specifically, sort the graphic and text information in descending order according to the deep matching score, and select the top N pieces of graphic and text information to form a push candidate set; perform weighted fusion on the surface matching score and the deep matching score to obtain the matching degree score.
[0050] Furthermore, when the matching degree score of a piece of graphic and text information with any category in the user preference category set exceeds the second threshold, this piece of graphic and text information is added to the push candidate set.
[0051] It should be noted that the second threshold is determined by statistically analyzing the matching degree scores in the historical user interaction data and finding the optimal classification threshold through the ROC curve.
[0052] S4: Sort the graphic and text information in the push candidate set in ascending order according to the matching degree with the user preference category set, and select the graphic and text information to construct a push sequence, and push the push sequence to the user terminal.
[0053] Specifically, obtain the comprehensive matching degree score and the corresponding category attribution information of the graphic and text information in the push candidate set; based on the interest degree thresholds of each category in the user preference category set, perform category weighted adjustment on the comprehensive matching degree score to obtain the sorting score.
[0054] Furthermore, sort the graphic and text information in the push candidate set in ascending order according to the sorting score to generate an initial push sequence; select the graphic and text information from the initial push sequence according to the preset time window size to construct a push sequence, where the category identifiers of adjacent graphic and text information in the push sequence are not repeated; encapsulate the push sequence into a push message, and send the push message to the user terminal through the push interface of the user terminal, where the push message includes the title, abstract, thumbnail, and category identifier of the graphic and text information.
[0055] Furthermore, this embodiment also provides a graphic and text information browsing and pushing system based on user preference settings, including: an acquisition module, configured to acquire interaction data when the user browses graphic and text information currently, and establish a user behavior feature matrix; a clustering analysis module, configured to calculate the user's interest in the current graphic and text information through the feature matrix and a preset weight vector, and perform clustering analysis on the graphic and text information in the user's historical browsing records to generate a set of user preference categories; a screening module, based on the set of user preference categories, adopts a dual matching algorithm combining cosine similarity and deep semantic matching to screen out candidate graphic and text information with a matching score exceeding a second threshold from a graphic and text information library to form a push candidate set; a pushing module, configured to sort the graphic and text information in the push candidate set in ascending order according to the matching degree with the set of user preference categories, and select graphic and text information to construct a push sequence, and push the push sequence to a user terminal.
[0056] This embodiment also provides a computer device applicable to the case of a graphic and text information browsing and pushing method based on user preference settings, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the graphic and text information browsing and pushing method based on user preference settings as proposed in the above embodiment.
[0057] This computer device may be a terminal, and this computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of this computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of this computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of this computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0058] The present embodiment also provides a storage medium having a computer program stored thereon, which implements the following steps when executed by a processor: obtaining interaction data of a user when currently browsing graphic and text information, and establishing a user behavior feature matrix; calculating the user's interest in the current graphic and text information through the feature matrix and a preset weight vector, and performing cluster analysis on the graphic and text information in the user's historical browsing record to generate a user preference category set; based on the user preference category set, a dual matching algorithm combining cosine similarity and deep semantic matching is used to screen out candidate graphic and text information with a matching score exceeding a second threshold from a graphic and text information library to form a push candidate set; sorting the graphic and text information of the push candidate set in ascending order according to the matching degree with the user preference category set, selecting graphic and text information to construct a push sequence, and pushing the push sequence to the user terminal.
[0059] In summary, the present invention realizes the three-dimensional characterization of user browsing behavior by constructing a multi-dimensional user behavior feature matrix, avoiding the serious information loss problem of traditional feature extraction methods; adopts feature matrix normalization processing and improved K-means clustering algorithm, combined with feature vector extraction of text topic model, accurately quantifies user interest distribution and improves the accuracy of user preference modeling; introduces a dual matching algorithm combining cosine similarity and deep semantic matching based on Transformer architecture, and improves the accuracy of content matching through the collaborative calculation of surface features and deep semantics; realizes dynamic optimization of push sequence through category weighted sorting mechanism and time window control strategy, ensuring the diversity and timeliness of pushed content.
[0060] Example 2
[0061] Referring to Table 1, which is a second embodiment of the present invention, this embodiment provides a method for browsing and pushing graphic information based on user preference settings. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0062] Specifically, 10,000 active users of a news information platform were selected as experimental subjects to collect browsing behavior data; the experimental environment adopted a Linux server cluster configured with Intel Xeon CPU E7-8880v4, 256GB RAM, and NVIDIA Tesla V100 GPU; the user behavior log system was used to collect the interaction data between users and graphic information, including page dwell time, click frequency, favorite tags, sharing behavior, and comment content semantic scores; the collected raw data was cleaned and preprocessed to remove outliers and invalid data, and finally 8,732 valid samples were obtained.
[0063] Furthermore, the processed interaction data is normalized to construct a 300×5-dimensional user behavior feature matrix, and a preset weight vector with a dwell time weight of 0.3, a click frequency weight of 0.25, a favorite mark weight of 0.2, a sharing behavior weight of 0.15, and a comment semantic score weight of 0.1 is set; through the dot product operation of the feature matrix and the weight vector, the interest score of the user for each piece of graphic and text information is calculated; based on the calculated interest, the improved K-means algorithm (K = 8) is used to perform clustering analysis on the user's historical browsing records, and the topic feature vector is extracted through the text topic model, and finally the user preference category set is formed.
[0064] Furthermore, in the content matching stage, the TF-IDF algorithm and the pre-trained ResNet-50 model are used to extract the text feature vector and the image feature vector respectively, and the two feature vectors are concatenated to form a 1024-dimensional hybrid feature vector. At the same time, a 12-layer Transformer encoder is used to construct a deep semantic matching model, with the number of attention heads set to 8 and the hidden layer dimension to 512. The surface matching score (weight 0.4) and the deep matching score (weight 0.6) are weighted and fused, and a matching degree threshold of 0.75 is set to screen and form a push candidate set. The graphic and text information in the push candidate set is sorted and screened. A time window of 30 minutes is set to ensure that the categories of adjacent pushed graphic and text information do not repeat, and the final push sequence is constructed; the push message is sent to the user terminal through the REST API interface, and the user's feedback data is recorded.
[0065] Specifically, as shown in Table 1, in terms of the average click-through rate, the average click-through rate of the method of the present invention reaches 35.9%, which is significantly increased by 20.7 percentage points compared with 15.2% of the traditional method, and the increase rate reaches 136.2%. This increase is mainly due to the multi-dimensional user behavior feature matrix and the dual matching algorithm adopted by the present invention, which can more accurately capture the user's interest characteristics and thus push content more in line with the user's interests.
[0066] Table 1. Comparison table between the method of the present invention and the traditional method
[0067] Evaluation Index Traditional Method Method of the Present Invention Average Click-Through Rate (%) 15.2 35.9 User Stay Duration (s) 45.3 142.8 Content Relevance 0.68 0.94 User Feedback Satisfaction 3.2 4.7 7-Day Retention Rate (%) 45.6 75.8 Calculation Time Consumption (ms) 850 595 Push Conversion Rate (%) 8.4 20.1
[0068] Furthermore, in terms of the user stay duration metric, the method of the present invention reached 142.8 seconds, an increase of 97.5 seconds compared to the 45.3 seconds of the traditional method, with a promotion rate as high as 215.2%. This indicates that through the push mechanism of the present invention, users obtain more valuable and interesting content, so they are willing to spend more time reading in depth, reflecting the high quality of the pushed content; the content relevance of the method of the present invention reached 0.94, an increase of 0.26 compared to the 0.68 of the traditional method, with a promotion rate of 38.2%. This benefits from the dual matching algorithm that combines cosine similarity and deep semantic matching adopted by the present invention, which not only considers the matching of surface features but also takes into account the understanding of deep semantics, thus achieving more accurate content matching.
[0069] Moreover, the user feedback satisfaction increased from 3.2 points of the traditional method to 4.7 points (with a full score of 5 points) of the method of the present invention, an increase of 1.5 points, with a promotion rate of 46.9%. The 7-day retention rate of the method of the present invention reached 75.8%, an increase of 30.2 percentage points compared to the 45.6% of the traditional method, with a promotion rate of 66.2%. The relatively high retention rate indicates that the method of the present invention can continuously provide valuable content push services.
[0070] Specifically, in terms of system performance, the calculation time consumption of the method of the present invention was 595 ms, a reduction of 255 ms compared to the 850 ms of the traditional method, with an optimization rate of 30%. This shows that although the present invention adopts a more complex algorithm strategy, through the optimized implementation method, the calculation efficiency is improved. The push conversion rate of the method of the present invention reached 20.1%, an increase of 11.7 percentage points compared to the 8.4% of the traditional method, with a promotion rate of 139.3%. This indicator comprehensively reflects the ultimate business value of the push effect, and the improvement fully proves the superiority of the method of the present invention in practical applications.
[0071] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for browsing and pushing graphic information based on user preference settings, characterized in that: include, Obtain the user's interaction data when browsing graphic and text information, and establish a user behavior feature matrix; The user's interest in the current graphic information is calculated by using the feature matrix and the preset weight vector, and the graphic information in the user's historical browsing records is clustered to generate a user preference category set; Based on the user preference category set, a dual matching algorithm combining cosine similarity and deep semantic matching is used to screen candidate image and text information with a matching score exceeding a second threshold from the image and text information library to form a push candidate set; The image and text information of the push candidate set are sorted in ascending order according to the matching degree with the user preference category set, and the image and text information are selected to construct a push sequence, and the push sequence is pushed to the user terminal.
2. The method for browsing and pushing graphic and text information based on user preference settings according to claim 1, characterized in that: The image and text information of the push candidate set is sorted from high to low according to the matching degree with the user preference category set, and the image and text information is selected to construct a push sequence, and the push sequence is pushed to the user terminal, including: Obtain the comprehensive matching score of the image and text information in the push candidate set and the corresponding category information; Based on the interest threshold of each category in the user preference category set, the comprehensive matching score is subjected to category weighting adjustment to obtain a ranking score; Sort the graphic and text information in the push candidate set in ascending order according to the sorting scores to generate an initial push sequence; Selecting graphic information from the initial push sequence according to a preset time window size to construct a push sequence, wherein category identifiers of adjacent graphic information in the push sequence are not repeated; The push sequence is encapsulated as a push message, and the push message is sent to the user terminal through the push interface of the user terminal, wherein the push message contains the title, summary, thumbnail and category identification of the graphic information.
3. The method for browsing and pushing graphic and text information based on user preference settings according to claim 2, characterized in that: The method for forming the push candidate set is: Based on the user preference category set, the text feature vector is extracted using the TF-IDF algorithm, and the image feature vector is obtained using the pre-trained image feature extraction model; The text feature vector and the image feature vector are concatenated to form a mixed feature vector of the image and text information, and the cosine similarity between the mixed feature vector and the feature vectors of each category in the user preference category set is calculated to obtain a surface matching score; At the same time, a deep semantic matching model based on the Transformer architecture is used to calculate the semantic similarity between the image and text information content and the sample content of each category in the user's preference category set to obtain a deep matching score. The deep semantic matching model captures the contextual association of the content through the attention mechanism. Arrange the image and text information in descending order according to the deep matching score, and select the top N image and text information to form a push candidate set; Performing weighted fusion on the surface matching score and the deep matching score to obtain a matching score; When the matching score between a certain piece of graphic information and any category in the user preference category set exceeds a second threshold, the graphic information is added to the push candidate set.
4. The method for browsing and pushing graphic and text information based on user preference settings according to claim 3, characterized in that: The specific formula of the deep semantic matching model is as follows: Among them, S deep (Q, D) is the matching score, Q is the query text feature vector, D is the target text feature vector, n is the feature dimension, w i is the weight coefficient of the i-th dimension feature, α is the fusion parameter, H(·) is the semantic encoding function, and σ is the Gaussian kernel parameter.
5. The method for browsing and pushing graphic and text information based on user preference settings according to claim 3, characterized in that: The method for generating the user preference category set is: Normalizing the feature matrix and setting a preset weight vector, wherein the components in the preset weight vector W correspond to the weight coefficients of the interaction data; Performing a dot product operation on the interaction feature vector of the normalized feature matrix and the preset weight vector to obtain the user's interest in the graphic information; Based on the interest level, an improved K-means algorithm is used to perform topic clustering on the graphic information in the user's historical browsing records, and a topic feature vector of the graphic information is extracted through a text topic model; Clustering is performed through the cosine similarity of the topic feature vectors to determine the topic categories, and the topic features of each category are extracted to form a category feature vector; Calculating the mapping relationship between the category feature vector and the predefined subject system, counting the number and interest of the graphic information associated with the semantic category, and generating the category weight; The semantic categories are sorted according to the category weights, and categories whose category weights exceed a first threshold form a user preference category set.
6. The method for browsing and pushing graphic and text information based on user preference settings according to claim 5, characterized in that: The method for establishing the feature matrix is: The interaction data is normalized and combined into a behavior feature vector, and arranged in order according to the identification ID of the graphic information to form a user behavior feature matrix, wherein the row vector of the feature matrix represents the user interaction data of the graphic information, and the column vector represents the interaction feature dimension.
7. The method for browsing and pushing graphic and text information based on user preference settings according to claim 6, characterized in that: The interaction data includes dwell time, click frequency, favorite mark, sharing behavior and comment content semantic score; the dwell time is to record the start time point and end time point of the user's stay on the graphic information page, and the dwell time is obtained by calculating the time difference between the start time point and the end time point; The click frequency is to count the number of times a user clicks on an interactive element in a graphic information within a specified time window, and to calculate the click frequency per unit time in combination with the duration of the time window; the collection mark is to record the user's collection operation on the graphic information, and the collection mark is set as a binary feature, where the collected state is recorded as 1 and the uncollected state is recorded as 0; the sharing behavior is to detect the user's behavior of sharing the graphic information to different social platforms, record the sharing target platform type and the number of shares, and calculate the comprehensive sharing index according to the weight coefficients of different social platforms; The comment content semantic score is obtained by using a natural language processing model to perform sentiment analysis on the comments posted by users, extracting the sentiment polarity and sentiment intensity of the comments, and calculating the semantic comprehensive score of the comment content in combination with the number of words in the comments and the comment time.
8. A system for browsing and pushing graphic and text information based on user preference settings, based on the method for browsing and pushing graphic and text information based on user preference settings according to any one of claims 1 to 7, characterized in that: include, The acquisition module is used to obtain the interaction data of the user when browsing the graphic information and establish the user behavior feature matrix; A cluster analysis module, used to calculate the user's interest in the current graphic information through the feature matrix and the preset weight vector, and to perform cluster analysis on the graphic information in the user's historical browsing records to generate a user preference category set; A screening module, based on the user preference category set, uses a dual matching algorithm combining cosine similarity and deep semantic matching to screen candidate image and text information with a matching score exceeding a second threshold from the image and text information library to form a push candidate set; The push module is used to sort the image and text information of the push candidate set in ascending order according to the matching degree with the user preference category set, select the image and text information to construct a push sequence, and push the push sequence to the user terminal.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for browsing and pushing graphic and text information based on user preference settings are implemented as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for browsing and pushing graphic and text information based on user preference settings according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Song menu recommendation method and apparatus
CN108021568A
Information recommendation method and device, computer equipment and storage medium
CN113590935A
Plasticizing product recommendation method and system based on user demands
CN119128177A
Question recommendation method, device and system, electronic device, and readable storage medium
US20220198300A1
Cited By
E-commerce commodity image-text generation optimization system based on generative AI
CN120472050A
E-commerce product image and text generation optimization system based on generative AI
CN120472050B
Marketing copywriting iteration generation method and device based on potential semantic optimization and medium
CN120805857A