Information retrieval method and device, storage medium and electronic equipment
By analyzing the topic categories and sentiment attributes of search requests, user interest profiles are constructed and ranking weights are dynamically adjusted. This addresses the shortcomings of existing information retrieval methods in terms of timeliness and scenario adaptability, and achieves more efficient personalized information recommendation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE INTERNET CO LTD
- Filing Date
- 2025-12-08
- Publication Date
- 2026-04-17
AI Technical Summary
Existing information retrieval methods struggle to effectively capture the timeliness requirements of information and cannot flexibly adjust result ranking strategies according to different retrieval scenarios, resulting in poor recommendation performance in differentiated scenarios.
By acquiring search request information, analyzing its topic categories and sentiment attributes, constructing user interest profiles, dynamically adjusting the search ranking weights of each topic, and sorting candidate search results according to the adjusted weights.
It improves the relevance, timeliness, and personalization of search results, overcoming the rigidity of ranking caused by the lack of scenario understanding and timeliness modeling in traditional methods.
Smart Images

Figure CN121880428A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to an information retrieval method, device, storage medium, and electronic device. Background Technology
[0002] With the rapid development of artificial intelligence and natural language processing technologies, intelligent assistants have been widely used in information retrieval, content recommendation, and human-computer interaction. Users have placed higher demands on the personalization, real-time performance, and accuracy of information services, prompting intelligent systems to gradually shift from traditional keyword matching to deep personalized services based on user behavior understanding and preference modeling.
[0003] Currently, personalized information retrieval methods typically construct static user profiles based on users' historical behavior data and combine them with collaborative filtering, content similarity calculation, or deep learning models for recommendations. However, this information retrieval method struggles to effectively capture the timeliness requirements of information and cannot flexibly adjust result ranking strategies according to different retrieval scenarios, resulting in poor recommendation performance in some differentiated scenarios. Summary of the Invention
[0004] In view of this, this application provides an information retrieval method, apparatus, storage medium and electronic device, the main purpose of which is to improve the current technical problem that it is difficult to effectively capture the timeliness requirements of information and cannot flexibly adjust the result ranking strategy according to different retrieval scenarios, resulting in poor recommendation effect in some differentiated scenarios.
[0005] Firstly, this application provides an information retrieval method, including: Obtain search request information; The search request information is analyzed to identify the topic category and sentiment attribute corresponding to the information search request; Based on the topic categories and the sentiment attributes, construct user interest profiles; Based on the user interest profile, the search ranking weight of each topic is dynamically adjusted; The candidate search results obtained from the initial search are sorted according to the adjusted search ranking weights to generate the target search results.
[0006] Secondly, this application provides an information retrieval device, comprising: The acquisition module is configured to acquire retrieval request information; The analysis module is configured to analyze the retrieval request information and identify the topic category and sentiment attribute corresponding to the information retrieval request. The construction module is configured to build a user interest profile based on the topic category and the sentiment attribute; The adjustment module is configured to dynamically adjust the search ranking weight of each topic based on the user interest profile. The generation module is configured to sort the candidate search results obtained from the initial search based on the adjusted search ranking weights, and generate the target search results.
[0007] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0008] Fourthly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect.
[0009] Fifthly, this application provides a computer program product having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method of the first aspect.
[0010] By employing the above technical solutions, this application provides an information retrieval method, apparatus, storage medium, and electronic device. First, it acquires retrieval request information; then, it analyzes the retrieval request information to identify the topic category and sentiment attribute corresponding to the information retrieval request; next, it constructs a user interest profile based on the topic category and sentiment attribute; then, it dynamically adjusts the retrieval ranking weights of each topic according to the user interest profile; and finally, it ranks the candidate retrieval results obtained from the initial retrieval based on the adjusted retrieval ranking weights to generate the target retrieval results. Compared with existing technologies, this application dynamically associates the topic category and sentiment attribute of the retrieval request with the user interest profile, and introduces a perception and response mechanism for the timeliness requirements of information in the retrieval scenario. This makes the retrieval ranking strategy no longer limited to static interest matching or single relevance calculation, but can adaptively adjust the strength of different dimensional features in the ranking according to the context of the current retrieval, especially strengthening the identification and adaptation capabilities for time-sensitive scenarios. This effectively overcomes the rigid ranking problem caused by the lack of scenario understanding and timeliness modeling in traditional methods, significantly improving the relevance, timeliness, and personalization of the results when facing diverse retrieval intentions and differentiated usage scenarios.
[0011] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart illustrating an information retrieval method provided in an embodiment of this application is shown; Figure 2 A flowchart illustrating another information retrieval method provided in an embodiment of this application is shown; Figure 3 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 4 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 5 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 6 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 7 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 8 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 9 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 10 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 11 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 12 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 13 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 14 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 15 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 16 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 17 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 18 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 19 A schematic diagram of the structure of an information retrieval device provided in an embodiment of this application is shown. Detailed Implementation
[0015] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0016] For example, user interests are multi-dimensional and evolve over time. Traditional static interest models lack the ability to perceive and respond to such dynamic changes, failing to capture shifts or mutations in user interests in a timely manner. This leads to significant discrepancies between the content recommended or retrieved by the system and the user's current actual needs. Meanwhile, although some existing solutions have made progress in improving retrieval accuracy, they often neglect effective control of computational complexity. With increasingly complex model structures and continuously increasing feature dimensions, the system struggles to achieve low-latency real-time responses with limited resources. This not only restricts retrieval efficiency but also directly weakens the user's smooth experience and satisfaction when using the intelligent assistant. Furthermore, as the core interactive entry point for users, the service quality of the intelligent assistant highly depends on the accuracy and response speed of the underlying information retrieval system. If the performance bottlenecks and experience shortcomings caused by the lag in interest modeling and excessive computational overhead cannot be effectively addressed, it will be difficult to truly achieve high-quality, highly relevant personalized information services, thereby affecting user stickiness and product competitiveness.
[0017] To address the current technical challenges of effectively capturing the timeliness requirements of information and flexibly adjusting result ranking strategies based on different retrieval scenarios, resulting in poor recommendation performance in some differentiated situations, this embodiment provides an information retrieval method. This method can be applied to electronic devices or server platforms with real-time retrieval and user modeling capabilities, such as intelligent assistants, search engines, or personalized recommendation systems. Figure 1 As shown, the method includes: Step 101: Obtain search request information.
[0018] The search request information includes the search keywords entered by the user in the current search operation.
[0019] Step 102: Analyze the retrieval request information to identify the topic category and sentiment attribute corresponding to the information retrieval request.
[0020] For example, semantic vector representations of keywords can be extracted using a word vector model (such as Word2Vec) based on the search keywords, and a Naive Bayes classifier can be used to classify the keywords contained in the search request into topics to identify their respective topic categories. At the same time, by combining a sentiment dictionary or a dictionary-based sentiment analysis method, the sentiment polarity of the keywords can be judged to determine the sentiment tendency attribute of the search request, including positive, negative, or neutral sentiment. In subsequent statistics and weight calculations, only positive and negative sentiments are usually included in the proportion calculation, while neutral sentiments are not included in the construction of the sentiment tendency proportion. The results of contextual semantic parsing can also be combined to improve the accuracy of topic and sentiment recognition.
[0021] Step 103: Construct user interest profiles based on topic categories and sentiment attributes.
[0022] For example, based on the topic category and sentiment attribute identified in the search request, the user's interest profile information can be obtained from a pre-established user interest profile database. This interest profile information includes the user ID, topic preference vector, sentiment ratio, and a list of frequently used information sources. The topic preference vector reflects the user's historical preference intensity in different topic categories (such as technology, entertainment, sports, etc.), the sentiment ratio records the distribution of positive and negative emotions shown by the user in past behaviors, and the list of frequently used information sources summarizes the types of information sources that the user frequently accesses or prefers. By comparing the topic category of the current search request with the user's topic preference vector and combining the consistency between the sentiment attribute and the user's historical sentiment ratio, a quantitative assessment of the matching degree between the user's current intention and long-term interests can be achieved.
[0023] Step 104: Dynamically adjust the search ranking weight of each topic based on user interest profiles.
[0024] For example, when performing personalized information retrieval, the search ranking weights of different topics are dynamically adjusted, and the information source types preferred by users are given priority. Specifically, for adjusting the search ranking weights of different topics, a preset number of days is set. Starting from the user's current date, the behavioral dataset of different topics searched by the user within the preset number of days is obtained, and the weights of the behavioral data are calculated and adjusted with the preset weight increment value. The weights of the behavioral data that exceed the preset number of days are calculated and adjusted with the preset weight decrement value.
[0025] In some examples, for information sources that users prefer to prioritize, a preset threshold is established. Information sources that users have accessed more than the threshold in the past are identified, and their content quality, update frequency, and relevance are scored. Priority recommendations are then made based on these scores. Furthermore, during personalized information retrieval, the scenario type of the user's immediate search request is determined. Based on the scenario type, the user's level of information timeliness requirement is determined, and the latest information is retrieved from high-level timeliness information sources. Personalized search results are then ranked in conjunction with a comprehensive user preference representation. For each level of timeliness requirement, a preset threshold is established. For scenarios with timeliness requirements exceeding the threshold, the ranking weight of immediate information is increased; for scenarios with timeliness requirements below the threshold, the weight of immediate information is decreased, thus matching the differentiated information timeliness requirements of users in different scenarios.
[0026] Step 105: Sort the candidate search results obtained from the preliminary search according to the adjusted search ranking weights to generate the target search results.
[0027] In some examples, personalized information retrieval strategies are generated based on the ranking of personalized search results. The intelligent assistant then uses these user information retrieval strategies learned from these preferences to make personalized recommendations. Furthermore, features from different dimensions, such as topic category distribution, sentiment tendency ratio, and information source authority, are weighted and fused. By collecting historical data of users in their respective dimensions, the influence of each dimension on user behavior is analyzed, key features are extracted, and a preset initial weight is assigned to each dimension. By weighted and fused features from different dimensions, a comprehensive user preference representation is generated to support the aforementioned ranking and recommendation process.
[0028] Compared with existing technologies, the technical solution of this embodiment first obtains retrieval request information; analyzes the retrieval request information to identify the topic category and sentiment attribute corresponding to the information retrieval request; then, based on the topic category and sentiment attribute, constructs a user interest profile; next, dynamically adjusts the retrieval ranking weight of each topic according to the user interest profile; and sorts the candidate retrieval results obtained from the preliminary retrieval according to the adjusted retrieval ranking weight to generate the target retrieval result. Compared with existing technologies, this application dynamically associates the topic category, sentiment attribute of the retrieval request with the user interest profile, and on this basis introduces a perception and response mechanism for the timeliness requirements of information in the retrieval scenario. This makes the retrieval ranking strategy no longer limited to static interest matching or single relevance calculation, but can adaptively adjust the strength of the role of different dimensions of features in the ranking according to the differences in the current retrieval context. In particular, it strengthens the identification and adaptation capabilities for timeliness-sensitive scenarios, thereby effectively overcoming the rigid ranking problem caused by the lack of scenario understanding and timeliness modeling in traditional methods. When facing diverse retrieval intentions and differentiated usage scenarios, it significantly improves the relevance, timeliness, and personalization of the results.
[0029] To further illustrate the specific implementation process of the method in this embodiment, this embodiment provides a method such as... Figure 2 The method shown includes: Step 201: Obtain search request information.
[0030] In some examples, such as Figure 3 As shown, the system obtains retrieval request information, extracts keywords using the Word2Vec word vector model based on the retrieval request information, and uses a Naive Bayes classifier to identify the topic categories involved in the retrieval request; it also obtains user interest profile information from a pre-established user interest profile database, which includes user ID, topic preference vector, sentiment tendency ratio, and a list of commonly used information sources.
[0031] Step 202: Analyze the retrieval request information to identify the topic category and sentiment attribute corresponding to the information retrieval request.
[0032] Optionally, step 202 may specifically include: cleaning and standardizing the retrieval request information to generate user behavior time-series data; extracting each retrieval keyword within a preset time window based on the user behavior time-series data; determining the topic category mapped by each retrieval keyword and its corresponding sentiment tendency; calculating the positive and negative ratio of sentiment tendency based on the frequency and / or weight of the occurrence of topic categories and sentiment tendencies within each time window; and determining the sentiment tendency attribute based on the positive and negative ratio of sentiment tendency.
[0033] For example, the original search text first needs to be cleaned and standardized, including preprocessing operations such as removing special characters, unifying capitalization, word segmentation, and stop word filtering, in order to generate standardized input that can be used for subsequent analysis. On this basis, the system can combine the user's historical behavior sequence within a preset time window to construct a lightweight user behavior time-series data context to assist in understanding the intent of the current request. Topic category identification typically employs supervised text classification methods, such as Naive Bayes, Support Vector Machines, or fine-tuned classifiers based on pre-trained language models. This semantic classification is achieved by mapping keywords or phrases to predefined topic systems (e.g., news, technology, finance, health, etc.). Sentiment attribute determination can utilize external sentiment dictionaries for rule matching or employ sentiment analysis models to calculate overall sentiment polarity scores. Furthermore, the frequency or weighted proportion of positive and negative occurrences of each keyword within the window is statistically analyzed, with weights dynamically assigned based on TF-IDF values, click feedback, or dwell time. By setting thresholds or proportion ranges (e.g., a positive percentage exceeding 60% is considered a positive tendency), the continuous sentiment distribution is discretized into explicit sentiment attributes, providing structured and computable semantic feature inputs for subsequent interest matching and ranking optimization.
[0034] Step 203: Based on the authority rating of the information source clicked by the user under the topic category and the corresponding page browsing time, generate the comprehensive score of the information source corresponding to the topic.
[0035] In some examples, based on user behavior time series datasets, the topic categories and sentiment tendencies involved in keywords within different time windows can be extracted, the distribution probability of topic categories and the positive and negative proportions of sentiment tendencies can be statistically analyzed, and a multi-dimensional user interest profile time series can be constructed by combining the information source type, authority and browsing time duration of user clicks.
[0036] For example, such as Figure 4As shown, search keywords within different time windows are obtained from a user behavior time-series dataset. A Naive Bayes classifier is used to classify the search keywords into topics, obtaining the topic categories involved in each time window. Based on the topic categories, a sentiment dictionary is used to determine the sentiment polarity of the search keywords, thus identifying their sentiment tendency. For each time window, the type of information source clicked by the user, the corresponding authority score, and the page browsing time are obtained. A weighted average method is used to calculate the comprehensive score of information source authority and browsing time. Combining the topic category distribution probability, the positive and negative ratio of sentiment tendency, and the average comprehensive score of information sources for each topic category, a multi-dimensional user interest feature vector is constructed. These values are then concatenated into a single vector in sequence. A feature vector is generated for each time window and arranged chronologically to form a user interest profile time series. The Euclidean distance between user interest feature vectors in adjacent time windows is calculated to obtain the rate of change of user interest. The difference in the rate of change is further calculated to determine the acceleration of user interest change. If the rate of change or acceleration of interest exceeds a preset threshold, it is marked as an interest abrupt change point to identify drastic changes in user interest.
[0037] For example, the features of different dimensions such as topic category distribution, sentiment tendency ratio, and information source authority are weighted and fused. By collecting users' historical data on their respective dimensions, the influence of each dimension on user behavior is analyzed, key features are extracted, and a preset initial weight is assigned to each dimension. By weighted and fused features of different dimensions, a comprehensive user preference representation is generated.
[0038] In some examples, such as Figure 5 As shown, the process involves obtaining the user's most recent historical data from the user behavior database. This historical data includes dimensions such as topic categories, sentiment tendencies, and information source selection. Based on the historical data, the distribution and trends of each dimension's features are calculated to obtain feature distribution data. Quantitative analysis is performed on the feature distribution data to determine the correlation between each dimension's features and user behavior. If the correlation exceeds a preset threshold, the corresponding feature is identified as a key feature. A logistic regression model is used to process the key features, obtaining feature importance data. Based on the feature importance data, the analytic hierarchy process (AHP) is used to assign preset initial weights to dimensions such as topic category distribution, sentiment tendency ratio, and information source authority. A weight vector is calculated using a judgment matrix. The key features are normalized to obtain a normalized feature vector. A weighted average method is used to fuse the normalized feature vector and the weight vector to obtain a comprehensive user preference representation vector, which contains user interest and preference information across multiple dimensions.
[0039] For example, historical data on users over the past 90 days in terms of topic categories, sentiment tendencies, and information source selection are extracted from a user behavior database. Independent statistical analysis is performed on the data for each dimension, calculating the distribution and trends of each feature, including mean, variance, and percentiles. A mutual information algorithm is used to quantify the correlation between each feature and user behavior. For discrete features, frequency counting is used to calculate the joint probability distribution; for continuous features, kernel density estimation is used to calculate the probability density function. The influence of each dimension is ranked according to the calculated mutual information values, and features with higher mutual information values are selected as key features. A logistic regression model is used to calculate feature importance, further validating the influence of key features. Historical data on 1000 users over the past 90 days in terms of topic categories, sentiment tendencies, and information source selection are extracted from the user behavior database. Topic categories include technology product reviews, technology industry news, new technology product releases, and technology trivia. Statistics show that among 1000 users, 350 followed tech product reviews, 280 followed tech industry news, 220 followed new tech product releases, and 150 followed tech trivia. Sentiment levels were categorized as positive, negative, and neutral. 420 records showed positive sentiment, 180 showed negative sentiment, and 400 showed neutral sentiment. Information sources were selected from professional tech media, social media platforms, and official manufacturer websites. 380 users chose professional tech media, 460 chose social media platforms, and 160 chose official manufacturer websites. Assuming each topic category is assigned a numerical value (e.g., tech product reviews = 1, tech industry news = 2, new tech product releases = 3, tech trivia = 4), the mean is 2.2. The variance is approximately 1.1, indicating a degree of dispersion in user topic category selection. The 25th percentile corresponds to the topic category of new technology product launches; the 50th percentile (median) corresponds to technology industry news; and the 75th percentile corresponds to technology product reviews, reflecting the general distribution of user-focused topic categories. If positive is 1, negative is -1, and neutral is 0, the mean is 0.24, indicating a slightly positive overall sentiment. The variance is approximately 0.96, demonstrating significant fluctuations in sentiment. The 25th percentile corresponds to negative sentiment; the 50th percentile to neutral sentiment; and the 75th percentile to positive sentiment.The mean of the information source selection dimension, with professional technology media set as 1, social media platforms as 2, and official manufacturer websites as 3, is approximately 1.8, indicating that users prefer social media platforms and professional technology media. The variance is 0.64, indicating that the dispersion of information source selection is relatively small. The 25th percentile is professional technology media, the 50th percentile is social media platforms, and the 75th percentile is social media platforms, further illustrating users' information source preferences.
[0040] For example, a mutual information algorithm is used to quantitatively analyze the correlation between various features and user behavior (assuming this is the sum of user interactions, such as likes, comments, and shares). The joint probability distribution is calculated using frequency counting. For instance, regarding topic categories and user interaction frequency, statistics show that 200 users follow tech product reviews and have a high interaction frequency (more than 10 times). After a series of frequency counting and calculations, the mutual information value between topic categories and user behavior is 0.38. Similarly, the mutual information value between sentiment and user behavior is 0.26, and the mutual information value between information source selection and user behavior is 0.32. For continuous features (assuming user dwell time on the platform is a continuous feature), the probability density function is calculated using kernel density estimation. The calculated mutual information value between user dwell time and user behavior is 0.42. The influence of each dimension was ranked based on the calculated mutual information values: user dwell time (mutual information value 0.42), topic category (mutual information value 0.38), information source selection (mutual information value 0.32), and sentiment tendency (mutual information value 0.26). User dwell time, topic category, and information source selection, with higher mutual information values, were selected as key features. A logistic regression model was used to calculate feature importance, further validating the influence of the key features. A logistic regression model was established using high user interaction behavior (defined as more than 10 interactions) as the dependent variable and the selected key features as independent variables. The model calculations showed that the regression coefficients for user dwell time, topic category, and information source selection were 0.8, 0.6, and 0.5, indicating that these key features have a significant impact on user behavior, further validating the rationality of the previous key feature selection. Based on the importance of the key features, the analytic hierarchy process (AHP) was used to assign preset initial weights to dimensions such as topic category distribution, sentiment tendency ratio, and information source authority. An automated method was used to construct a judgment matrix, and the feature vectors were calculated to obtain the weight vectors. The feature vectors of each dimension are normalized to ensure that features of different dimensions are comparable.
[0041] For example, a weighted average method is used to fuse the normalized feature vector with the corresponding weight vector to obtain a comprehensive user preference representation vector, which contains user interest preference information across multiple dimensions. Assuming user ID is 12345, their behavioral data for the past 90 days is extracted, including browsed topic categories (technology 40%, entertainment 30%, sports 20%, news 10%), sentiment tendencies (positive 60%, negative 25%, neutral 15%), and frequently used information sources (website A 45%, website B 30%, website C 25%). The mean, variance, and 75th percentile of each feature dimension are calculated. Mutual information is used to analyze the correlation between features and user click behavior. For the discrete feature of topic categories, the frequency counting method is used to obtain the joint probability distribution; for the continuous feature of information source authority, the kernel density estimation method is used to calculate the probability density function. The results show that the mutual information value for topic categories is 0.8, for sentiment tendencies it is 0.6, and for information source authority it is 0.7. Verification using a logistic regression model confirms that the feature importance ranking is consistent with the mutual information values. A judgment matrix was constructed using the analytic hierarchy process (AHP). The scoring results were: topic category 9 points, sentiment 5 points, and information source authority 7 points. The weight vector was calculated to be [0.45, 0.25, 0.30]. The feature vectors were normalized, such as the topic category vector being normalized to [0.4, 0.3, 0.2, 0.1]. The feature vectors and weight vectors were then fused using a weighted average method to obtain the comprehensive user preference representation vector [0.18, 0.135, 0.09, 0.045, 0.15, 0.0625, 0.0375, 0.135, 0.09, 0.075].
[0042] Step 204: Based on the comprehensive score of information sources and the user's behavioral characteristics, a fusion model of topic categories and sentiment attributes is performed to obtain the user interest model.
[0043] Optionally, step 204 may specifically include: obtaining historical behavioral data of users in multiple dimensions based on the comprehensive score of information sources and the user's behavioral characteristics, including topic category dimension, sentiment tendency dimension, and information source authority dimension; determining the initial weight parameters of each dimension by analyzing the influence of historical behavioral data in multiple dimensions on user decision-making behavior; normalizing the feature values of each dimension to generate normalized feature vectors; weighting and fusing the normalized feature vectors with the corresponding initial weight parameters to generate a comprehensive user preference representation vector; and constructing a user interest model based on the comprehensive user preference representation vector.
[0044] For example, such as Figure 6 As shown, user search behavior data is obtained in different time windows, including search keywords, types of information sources clicked, authority of information sources, and browsing duration. Text and numerical data of different types are cleaned and standardized to obtain a structured user behavior time series dataset.
[0045] In some examples, such as Figure 7 As shown, the process involves acquiring raw data of user search behavior within multiple preset time windows. This raw data includes search keywords, clicked information source types, information source authority scores, and page browsing duration. Based on this raw data, the user behavior dataset is sorted by user ID and timestamp. A sliding window method is used to divide the dataset into samples of fixed time lengths. TF-IDF features are extracted from the search keywords within these fixed-time-length samples using the TF-IDF Vectorizer tool. The information source type field is one-hot encoded to obtain user behavior feature vectors. KMeans is used to cluster these feature vectors to identify the behavior patterns of different user types. Based on these patterns, a prediction model based on the Prophet time series prediction tool is constructed. This model receives historical user behavior sequence data and outputs the search keywords, clicked information source types, and browsing duration for the next time window.
[0046] For example, the jieba word segmentation tool is used to segment Chinese words and remove stop words for text fields, and the min-max scaling method is used for normalization of numerical fields, resulting in a preliminarily cleaned user behavior dataset. The cleaned user behavior dataset is then sorted by user ID and timestamp, and a sliding window method is used to divide the data into samples of fixed time lengths. This captures user behavior patterns across different time periods. TF-IDF feature extraction is performed on the search keywords within each sample using the TF-IDF Vectorizer tool from the sklearn library. One-hot encoding is used for the information source type field, and equal-frequency binning is used to divide the data into five equal-frequency intervals for information source authority ratings and browsing duration, generating behavioral feature vectors for users in different time windows.
[0047] In some examples, based on the generated user behavior feature vectors, statistical indicators such as the cosine similarity of search topics, information source preference distribution, and average page dwell time are calculated for users across different time windows. The KMeans tool from the sklearn library is then used to cluster user groups, revealing the behavioral patterns of different user types. Feature centers are calculated for each cluster category to identify typical user behavior patterns, including high-frequency search terms, commonly used information source types, and average browsing time. Hierarchical clustering algorithms or K-means can be used to group users. Hierarchical clustering constructs a hierarchy of similarity for user browsing features using Euclidean distance, suitable for discovering fine-grained user groups; K-means requires determining the optimal number of clusters using the elbow method or silhouette coefficient; TF-IDF is used to extract the weights of user search keywords, and the weighted average of keywords in each class is taken; the frequency distribution of information source types visited by each user class (such as news, product pages, forums) is statistically analyzed to form a probability vector; the mean and median time spent on the page by each user class are calculated; the PrefixSpan algorithm is used to mine frequent behavior sequences of each user class (such as "search → click on product → add to cart"); ANOVA is used to analyze the significant differences between different cluster groups (such as the influence of gender and cognitive style on behavior patterns), and silhouette coefficient is used to evaluate cluster quality to ensure high intra-class similarity and significant inter-class differences.
[0048] For example, based on cluster centers, a typical user type is defined: low search term diversity but high frequency, preference for vertical information sources (such as e-commerce product pages), average browsing time exceeding the 90th percentile of similar users, and frequent access to in-depth content (such as long articles and videos). Behavioral sequences are fragmented, with low continuity shown in the page transition probability matrix. The following indicators are statistically analyzed by time window (e.g., hour / day): search term frequency change rate (time-series fluctuations calculated using TF-IDF via a sliding window), entropy value of information source type distribution (reflecting user interest concentration), and periodicity of browsing time (daily / weekly patterns detected using Fourier transform).
[0049] For example, based on the identified different types of user behavior patterns, a prediction model based on the Prophet time series forecasting tool is constructed. Inputting historical user behavior sequence data, the model outputs the user's possible search keywords, clicked information source types, and browsing duration for the next time window, thus forming a complete user behavior analysis and prediction dataset. The user's historical behavior is converted into a time series format, and the frequency of the Top-N high-frequency words is statistically analyzed daily to form a multivariate sequence. One-hot encoding is performed on the access frequency of each type of information source, generating multiple independent sequences. Browsing duration is directly used as a continuous variable, aligned by timestamps. The time axis is extended using Prophet's `make_future_dataframe` function, and the prediction frequency is set (e.g., `freq='6H'`). A non-linear saturation growth trend is enabled for browsing duration, and a piecewise linear trend is used for search term frequency. Daily / weekly / monthly seasonality is explicitly added (e.g., weekend peak visits for e-commerce users). A "holiday effect" is introduced for information source type prediction (e.g., a surge in visits to news pages during promotional periods). The Prophet model is trained independently for each type of behavior (search term, information source, duration), and the optimal parameter combination is selected through cross-validation. For discrete variables (such as information source type), the predicted values are mapped to a probability distribution, and the term with the highest probability is taken as the final output. The prediction results are then merged into a structured dataset with fields including: user ID, prediction time window, Top-K search terms and their confidence levels, most likely information source type, and browsing duration interval.
[0050] In some examples, Mean Average Precision (MAP) can be used to measure the Top-K hit rate. The confusion matrix and F1-score are used. Calculate the MAE and RMSE of the browsing duration. When obtaining the original data of user retrieval behavior, set three time windows of 24 hours, 7 days, and 30 days, and record the user ID, retrieval timestamp, retrieval keywords, click information source type (such as news, blog, video), information source authority score (0 - 100), and page browsing duration (seconds). Use the jieba word segmentation tool to segment the retrieval keywords, for example, "artificial intelligence application" is segmented into "artificial intelligence" and "application", and stop words such as "de" and "le" are removed. Perform min-max scaling normalization on the information source authority score and browsing duration, and map the values to the 0 - 1 interval. Adopt a sliding window of 60 minutes, sliding every 10 minutes, to generate user behavior samples. For the retrieval keywords within each sample, use TfidfVectorizer to calculate the TF-IDF value, for example, the TF-IDF value of "artificial intelligence" is 0.8. The information source type adopts one-hot encoding, for example, [1, 0, 0] represents the news type. Divide the information source authority score and browsing duration into 5 intervals using the equal-frequency binning method, such as 0 - 20, 21 - 40, 41 - 60, 61 - 80, 81 - 100. Calculate the cosine similarity of the retrieval topics of the user within different time windows, for example, the similarity within 24 hours is 0.75, and within 7 days is 0.6. Statistically analyze the information source preference distribution, for example, news accounts for 60%, blog accounts for 30%, and video accounts for 10%. Calculate the average page stay time, for example, 30 seconds. Use the KMeans tool to cluster users into 3 categories, and calculate the feature center of each category. For example, the high-frequency retrieval words of the first category of users are "technology", the common information source is news, and the average browsing duration is 45 seconds. Based on the Prophet tool, construct a time series prediction model, input the user behavior data of the past 7 days, and predict the keywords that the user may search, the information source type to be clicked, and the expected browsing duration within the next 24 hours. For example, predict that the user will search for "5G technology", click on the blog type information source, and the browsing duration is about 40 seconds.
[0051] Step 205, construct a user interest portrait based on the user interest model.
[0052] Optionally, step 205 may specifically include: converting the topic distribution probability, sentiment tendency ratio, and information source comprehensive score of each topic corresponding to each time window into a structured feature vector; arranging the structured feature vectors in chronological order to form a user interest portrait time series, which is used to characterize the dynamic evolution process of the user interest; constructing a user interest portrait based on the user interest portrait time series.
[0053] In some examples, quantitative indicators such as the probability distance of topic category distribution and the difference between positive and negative sentiment ratios can be designed for different dimensions of the user interest profile time series. By calculating the magnitude of change of these indicators in different time windows, it can be determined whether the user's interest preferences in different topics and sentiment dimensions have changed significantly, and the significance level of the change can be determined based on the magnitude of the change.
[0054] For example, Jensen-Shannon divergence is used to calculate the probability distribution distance between adjacent time windows. The Jensen-Shannon divergence value is compared with a preset threshold. If the divergence value is greater than the preset threshold, a significant change in the topic category distribution is considered. The threshold can be determined based on the specific application scenario and data characteristics. The Manhattan distance between adjacent time windows is calculated to obtain a quantitative indicator of sentiment change. Manhattan distance is the total distance two points move along coordinate axes in a multidimensional space. For two probability distributions P and Q, the Manhattan distance is defined as:
[0055] If the Manhattan distance is greater than a preset sentiment change threshold, a significant change in sentiment is determined. A weighted average comprehensive change index is constructed by combining the probability distance of topic category distributions and the quantitative index of the difference between positive and negative sentiment proportions. This involves combining the JS divergence and Manhattan distance to create a weighted average comprehensive change index. Assuming the weights are w1 and w2, the comprehensive change index C can be expressed as:
[0056] Optionally, the method in this embodiment may further include: obtaining the magnitude of change of user interest preferences in each dimension based on the feature values of each dimension in the user interest profile time series; and dynamically adjusting the parameters associated with the corresponding dimension in the user interest model when the magnitude of change in any dimension exceeds a preset significance level threshold.
[0057] Optionally, the above-mentioned dynamic adjustment of parameters associated with the corresponding dimension in the user interest model may specifically include: obtaining the target feature vector corresponding to the target dimension whose change exceeds the significance level threshold; retraining the user interest model parameters corresponding to the target dimension based on the target feature vector using a preset machine learning model; and weighted fusion of the updated model parameters obtained from the training with the historical model parameters.
[0058] For example, the choice of weights can be adjusted according to actual needs and data characteristics. For instance, if more attention is paid to changes in the distribution of topic categories, the value of w1 can be increased; if more attention is paid to changes in sentiment, the value of w2 can be increased. For the comprehensive change index, time series processing is performed, assuming a window size of n, and the moving average MA of the time window is calculated:
[0059] Then, based on the moving average, calculate the mean and standard deviation:
[0060]
[0061] For example, such as Figure 8 As shown, the Z-score standardization method can be used to convert the mean and standard deviation into scores under a standard normal distribution. Based on the properties of the standard normal distribution, the Z-score is mapped to a significance level to determine the significance of changes in user interests and preferences. If the JS divergence or Manhattan distance is greater than a preset threshold, it is considered that the distribution of topic categories or sentiment tendencies has changed significantly. If the Z-score of the comprehensive change index is greater than the significance level (e.g., 1.96), it is considered that user interests and preferences have changed significantly.
[0062] For example, regarding the probability dimension of topic category distribution in the time series of user interest profiles, three fixed time windows—24 hours, 7 days, and 30 days—can be set. Jensen-Shannon divergence is used to calculate the probability distribution distance between adjacent time windows. Example: The divergence value for the 24-hour and 7-day windows is 0.08, less than the preset threshold of 0.1, indicating no significant change; the divergence value for the 7-day and 30-day windows is 0.15, greater than the threshold, indicating a significant change. The calculated Jensen-Shannon divergence value is compared with the preset threshold of 0.1. If the divergence value is greater than the threshold, a significant change in topic category distribution is determined. For the positive / negative proportion dimension of sentiment, the Manhattan distance between adjacent time windows is calculated to obtain a quantitative indicator of sentiment change. A sentiment change threshold of 0.2 is set. When the Manhattan distance is greater than the threshold, a significant change in sentiment is determined. Example: The distance between the 24-hour and 7-day windows is 0.1, less than the threshold of 0.2, indicating no significant change; the distance between the 7-day and 30-day windows is 0.3, greater than the threshold, indicating a significant change. A weighted average comprehensive change index is constructed by combining the probability distance of topic category distribution and the quantitative indicators of the positive and negative proportions of sentiment. The analytic hierarchy process (AHP) is used to determine the weights, with topic category weight set at 0.6 and sentiment weight at 0.4. The comprehensive change index is compared with a preset threshold of 0.15 to determine whether user interests and preferences have changed significantly. The obtained comprehensive change index undergoes time series processing, calculating a moving average over five time windows to eliminate the impact of short-term fluctuations. Based on the processed comprehensive change index series, the mean and standard deviation are calculated, and the change amplitude is converted into a score under a standard normal distribution using the Z-score standardization method. The specific calculation process is Z = (X - μ) / σ, where X is the original change amplitude, μ is the mean, and σ is the standard deviation. According to the properties of the standard normal distribution, the Z-score is mapped to a significance level: an absolute Z-score greater than 1.96 corresponds to a significance level of 0.05, and greater than 2.58 corresponds to a significance level of 0.01, thus determining the significance of changes in user interests and preferences. Three fixed time windows—24 hours, 7 days, and 30 days—are set for the user interest profile time series. For the probability distribution of topic categories, such as technology 0.3, entertainment 0.2, and sports 0.5, calculate the Jensen-Shannon divergence between adjacent windows. Assuming the divergence value for the 24-hour and 7-day windows is 0.08, which is less than the preset threshold of 0.1, the judgment is considered to have no significant change; the divergence value for the 7-day and 30-day windows is 0.15, which is greater than the threshold, the judgment is considered to have changed significantly. For the positive-negative ratio of sentiment, such as positive 0.6 and negative 0.4, calculate the Manhattan distance between adjacent windows. If the distance between the 24-hour and 7-day windows is 0.1, which is less than the threshold of 0.2, the judgment is considered to have no significant change; if the distance between the 7-day and 30-day windows is 0.3, which is greater than the threshold, the judgment is considered to have changed significantly.The weights were determined using the analytic hierarchy process (AHP). The weight of topic category was set at 0.6, and the weight of sentiment tendency at 0.4. A weighted average comprehensive change index was calculated. For example, if the topic change was 0.15 and the sentiment change was 0.3, the comprehensive index would be 0.15 * 0.6 + 0.3 * 0.4 = 0.21, which is greater than the preset threshold of 0.15, indicating a significant change in user interests and preferences. A moving average of five time windows was applied to the comprehensive change index. For example, the original sequence [0.21, 0.18, 0.23, 0.20, 0.22] would have a moving average of 0.208. The mean (0.208) and standard deviation (0.02) were calculated. The original change magnitude of 0.21 was standardized, and Z = (0.21 - 0.208) / 0.02 = 0.1. This Z-score is less than 1.96, corresponding to a significance level greater than 0.05, indicating that the change is not significant enough.
[0063] In some examples, when a significant change in a certain dimension of user interest preference is detected and exceeds a preset significance level threshold, the user interest preference update mechanism for that dimension is triggered, and the user interest model parameters for that dimension are adjusted.
[0064] For example, such as Figure 9 As shown, the chi-square test can be used to calculate the statistics for each dimension of user interest preferences, and these statistics are compared with a preset significance level threshold. If the p-value corresponding to the statistic is less than the significance level threshold, the user interest preference update mechanism for that dimension is triggered. Once the update mechanism is triggered, the feature vector for that dimension needs to be extracted from the user's most recent behavioral data. Feature vector extraction and preprocessing are important steps in training machine learning models. Feature vectors can be generated from user behavioral data (such as browsing, clicking, purchasing, etc.).
[0065] For example, a matrix R can be used to represent a user's preference vector, where rows of matrix R represent users, columns represent feature terms, and each element represents the degree of user preference for a feature term. Data cleaning is the first step in preprocessing, including operations such as removing missing values, noise, and outliers. This step ensures data quality and improves model training performance. Normalization scales the feature vector to a uniform range, typically between 0 and 1. This step reduces the influence of dimensions between features and improves model stability. Based on the update mechanism, feature vectors for that dimension are extracted from the user's most recent behavioral data, and preprocessed, including data cleaning and normalization. Depending on the type of dimension triggering the update, an appropriate update algorithm is selected, and the random forest algorithm is used to process the preprocessed feature vectors. Model performance is evaluated using methods such as cross-validation to ensure the model's accuracy and stability. The random forest algorithm is used to process the feature vectors to obtain updated model parameters. The updated model parameters are applied to the user interest model, and an exponential moving average method is used to smooth the old and new parameters. This completes the update of user interest preferences for a dimension while recording the update timestamp and the updated dimension information, where α is the smoothing coefficient, calculated using the following formula: New parameter = α Update parameter + (1-α) Old parameters For example, significance testing is performed on each dimension of user interests and preferences. The chi-square test is used to calculate the statistic for each dimension, and the statistic is compared with a preset significance level threshold of 0.05 to determine if the dimension has changed significantly. Assuming the chi-square statistic for the "technology" dimension is 10.5, the corresponding p-value is 0.003, which is less than the preset significance level threshold of 0.05, triggering an update mechanism. If the p-value corresponding to the statistic of a certain dimension is less than the significance level threshold, the user interest preference update mechanism for that dimension is triggered. The feature vector for that dimension is extracted from the user's behavioral data over the past 7 days and used as input data for updating the model. The extracted feature vectors are preprocessed, including data cleaning and normalization, removing outliers, and scaling numerical features to the 0-1 range. Based on the type of dimension triggering the update, an appropriate update algorithm is selected. A random forest algorithm is used to handle mixed-type features, and the parameters of the random forest model are retrained using the preprocessed feature vectors. The updated model parameters are applied to the user interest model, and an exponential moving average method is used to smooth the old and new parameters to avoid abrupt model changes. The specific calculation formula is: new parameter = α. Update parameter + (1-α) The old parameter, where α is the smoothing coefficient, is set to 0.3. Example: If the old value of a parameter is 0.6, the updated value is 0.8, and the smoothing coefficient α is set to 0.3, then the new parameter value is 0.3. 0.8 + 0.7 0.6 = 0.66. Complete the update of user interest preferences for this dimension, recording the update timestamp and the updated dimension information for subsequent user interest evolution analysis. Perform significance testing on the dimensions of user interest preferences such as "technology," "entertainment," and "sports," using the chi-square test to calculate the statistic for each dimension. Assuming the chi-square statistic for the "technology" dimension is 10.5, the corresponding p-value is 0.003, which is less than the preset significance level threshold of 0.05, indicating a significant change in this dimension. Trigger the user interest preference update mechanism for the "technology" dimension, extracting feature vectors for this dimension from the user's behavioral data over the past 7 days, such as pageview count, dwell time, and click-through rate. Preprocess the extracted feature vectors, removing outliers, such as records with more than 1000 pageviews, and normalizing numerical features such as dwell time to the range of 0-1. Use the random forest algorithm to process mixed-type features, setting 100 decision trees with a maximum depth of 10, and retrain the random forest model using the preprocessed feature vectors. After obtaining the new model parameters, an exponential moving average method is used to smooth the old and new parameters, with a smoothing coefficient α set to 0.3. For example, if the old value of a parameter is 0.6 and the updated value is 0.8, then the new parameter value is 0.3. 0.8 + 0.7 0.6 = 0.66. After the update is complete, record the current timestamp and the update information for the "Technology" dimension, and store them in the user interest evolution log for subsequent analysis of the dynamic changes in user interests.
[0066] In some examples, by dynamically constructing and updating the time series of user interest profiles by combining contextual information, significant changes in interest preferences are monitored using quantitative indicators such as the probability distance of topic category distribution and the difference in the proportion of sentiment tendency. Once a change in a certain dimension is detected to exceed the preset significance level threshold, the model parameter update process for that dimension is triggered, thereby more accurately predicting the user's immediate needs and providing personalized search results that are highly consistent with the current intent.
[0067] Step 206: Dynamically adjust the search ranking weight of each topic based on user interest profiles.
[0068] Optionally, step 206 may specifically include: adjusting the intensity of interest in each topic in the user's interest profile based on a preset period; using the intensity of interest in each topic after the time decay adjustment as the basic weight of the corresponding topic in the search ranking; and adjusting the basic weight in combination with the similarity between the user's current search request and the interest profile to obtain the search ranking weight.
[0069] In some examples, it is also possible to adjust the search ranking weights of different topics, preset the number of days, start from the user's current date, obtain the behavioral dataset of different topics searched by the user within the preset number of days, calculate and adjust the weight of the behavioral data with the preset weight increment value, and calculate and adjust the weight value of the behavioral data outside the preset number of days with the preset weight decrement value.
[0070] For example, such as Figure 10 As shown, the process involves retrieving the user ID, search timestamp, and search topic fields from the user behavior database. The dataset of user search behaviors that meet the time criteria is then filtered based on the user's current date and a preset number of days. Specifically, according to user requirements, the user ID, search timestamp, and search topic fields must be extracted from the user behavior database, and the dataset of search behaviors meeting the criteria is filtered based on the current date and a preset number of days. The de-identification method may include using symmetric encryption technology to reversibly de-identify the user ID using an encryption key and algorithm, ensuring that the ciphertext format is consistent with the original data logic rules, while also supporting data recovery requirements after authorization. An offset rounding scheme is applied, randomly shifting the timestamp value to maintain the authenticity of the time range while avoiding the exposure of the specific time point. A random value replacement method is used to replace the original search topic content with fictitious text that conforms to business characteristics, while preserving field format and semantic relevance. For each search behavior data in the user search behavior dataset, an exponential decay function is used to calculate the time decay factor to obtain the initial weight. If the search behavior falls within a preset number of days, the first weighted average method is used to determine the adjusted weight; if the search behavior exceeds the preset number of days, the second weighted average method is used to determine the adjusted weight. For each search topic category, the cumulative weight of the search behavior is calculated and divided by the total number of search behaviors under that topic to obtain the average weight. If there are no search behaviors under a search topic category, a default minimum weight value is assigned. The average weight or the default minimum weight value is used as the final weight value of the search topic in the search ranking.
[0071] For example, fields such as user ID, search timestamp, and search topic are extracted from a user behavior database. This database contains fields such as user_id (integer), search_time (timestamp), and search_topic (string). Based on the user's current date and a preset 30-day range, a dataset of user search behaviors meeting the time criteria is selected and categorized according to search topic. For each search behavior within the preset 30-day range, its time decay factor is calculated using the exponential decay function f(t) = e^(-λt), where t is the number of days since the search behavior, and λ is the decay rate parameter with a value of 0.1. The calculated time decay factor is used as the initial weight for that search behavior. For search behaviors within the preset 30-day range, the initial weight is multiplied by a preset weight increment of 1.2, and a weighted average method is used to obtain the adjusted weight value, which is 0.7. Initial weight +0.3 1.2. For search queries exceeding the preset number of days, multiply their initial weight by a preset weight reduction value of 0.8, and then apply a similar weighted average method. The adjusted weight is 0.7. Initial weight +0.3 0.8. The weights of all search behaviors under each topic category are summed and divided by the total number of search behaviors under that topic to obtain the average weight of that topic. If there are no search behaviors under a topic, a default minimum weight value of 0.1 is assigned. The calculated average weight is used as the final weight value of that topic in the search ranking. Assume user ID is 12345 and the current date is April 15, 2023. Extract the user's search records for the most recent 60 days from the user behavior database, obtaining 100 records. Filter according to a preset 30-day range, obtaining 80 records that meet the criteria. After classifying these records by topic, we find 30 records for the "Technology" topic, 40 for the "Entertainment" topic, and 10 for the "Sports" topic. For a search record from 5 days ago in the "Technology" topic, its time decay factor is calculated as e^(-0.15) = 0.61. Combining this initial weight of 0.61 with the preset increment value of 1.2, we obtain the adjusted weight: 0.7 / 0.61 + 0.3 / 1.2 = 0.79. For a search record on the topic "Entertainment" from 35 days ago, its initial weight is e^(-0.135) = 0.03. Combined with the preset reduction value of 0.8, the adjusted weight is 0.7 * 0.03 + 0.3 * 0.8 = 0.26. A similar calculation is performed on 30 records on the topic "Technology," and the average weight is calculated to be 0.65. The average weights for "Entertainment" and "Sports" are 0.58 and 0.72, respectively. Since the topic "News" did not appear in the user's search records, it is assigned the default minimum weight of 0.1. The final weights for each topic in the search ranking are: Technology 0.65, Entertainment 0.58, Sports 0.72, and News 0.1.
[0072] In some examples, to address the issue of excessive computational complexity impacting real-time response, a series of optimization measures can be introduced, including but not limited to employing lightweight machine learning models, using efficient representation methods such as TF-IDF vectorization and one-hot encoding during the feature engineering stage, optimizing the inverted index structure to accelerate candidate result recall, implementing an incremental update strategy to avoid full retraining, and caching and asynchronously processing key modules such as time decay weights and interest mutation detection. These measures aim to significantly reduce the CPU and memory resource consumption during system operation, achieve millisecond-level response while ensuring retrieval accuracy, and especially improve availability and smoothness in resource-constrained environments such as mobile devices.
[0073] Optionally, the method in this embodiment may further include: filtering high-frequency access information sources based on the user's historical access frequency to each information source; evaluating the high-frequency access information sources in multiple dimensions, including content quality, update frequency, and relevance to the user's interest profile, to generate a comprehensive score; constructing a priority recommendation information source sequence based on the comprehensive score, and increasing the ranking priority of result items belonging to the priority recommendation information source sequence in the candidate search results.
[0074] For example, when conducting personalized information retrieval, the topic category and sentiment attributes in the user's search request are obtained, the user's latest interest profile is matched, and the search ranking weight of different topics is dynamically adjusted, and the information source types preferred by the user are given priority recommendation.
[0075] In some examples, the cosine similarity between the search request topic and the user's interest topics is calculated. If the cosine similarity is greater than a preset threshold, the weight parameters for ranking the search results are adjusted based on the cosine similarity. Based on the adjusted weight parameters, combined with topic relevance, sentiment matching, and information source preference, a preliminary personalized search result ranking list is generated. Specifically, cosine similarity can be used as part of the weight parameters to adjust the ranking of search results. For example, cosine similarity can be combined with other factors such as topic relevance, sentiment matching, and information source preference to generate a preliminary personalized search result ranking list. After adjusting the weight parameters, combining topic relevance, sentiment matching, and information source preference, a preliminary personalized search result ranking list can be generated. These factors can be implemented in the following ways: Topic relevance: assessed by calculating the relevance of documents or content to the user's interest topics. Sentiment matching: assessed by analyzing the user's sentiment tendency and the sentiment characteristics of the content. Information Source Preferences: The quality and reliability of information sources are evaluated based on users' historical behavior and preferences. The initial personalized search result ranking list is adjusted using the maximum marginal relevance algorithm, selecting items with high relevance and significant differences from the already selected results to obtain the final personalized search result ranking list. To further optimize the personalized search result ranking list, the maximum marginal relevance algorithm can be used. This algorithm aims to select items with high relevance and significant differences from the already selected results. The specific steps are as follows: For each candidate item, calculate its marginal relevance to the currently selected items; select the item with the highest marginal relevance; ensure that the selected items have a high degree of content difference from the already selected items. Through these steps, the final personalized search result ranking list can be obtained. This list not only considers the similarity between user interests and search requests but also incorporates factors such as topic relevance, sentiment matching, and information source preferences, ensuring the accuracy and diversity of the recommendation results.
[0076] Optionally, the method in this embodiment may further include: performing semantic parsing and feature extraction on the retrieval request information to identify the scenario type corresponding to the retrieval request information, wherein the scenario type is used to characterize the user's sensitivity to the timeliness of the retrieval results; determining the timeliness requirement level corresponding to the current scenario type based on the mapping relationship between the scenario type and the information timeliness requirement level; when the timeliness requirement level meets the preset high timeliness condition, acquiring the real-time information provided by the real-time information source and increasing the retrieval ranking weight of the real-time information; when the timeliness requirement level meets the preset low timeliness condition, decreasing the retrieval ranking weight of the real-time information.
[0077] For example, semantic parsing is performed on the user's input search request. Keywords are extracted from the request using the Word2Vec word vector model, and the topic categories involved in the request are identified using a Naive Bayes classifier. A dictionary-based method is then used to determine the sentiment tendency of the request. The latest interest profile of the user is obtained from the user interest profile database, including fields such as user ID, topic preference vector, sentiment tendency ratio, and a list of frequently used information sources. The cosine similarity between the search request topic and the user's interest topics is calculated. Based on the similarity calculation results, the weight parameters of the search ranking are dynamically adjusted, assigning higher weights to topics highly relevant to the user's interests. The specific adjustment formula is: Adjusted weight = Original weight. The similarity is calculated as (1 + cosine similarity), while also considering sentiment matching. When the sentiment of the search results matches the user's preferences, the weight is multiplied by 1.2. Example: The cosine similarity between the search query topic "technology" and the user's interest topic is 0.95; the original weight of 0.5 is adjusted to 0.5. (1 + 0.95) = 0.975. In the search results, information source types preferred by the user are prioritized, and the content of these information sources is weighted to improve the score. The weighting formula is: Final Score = 0.5 Topic relevance +0.3 Emotional compatibility +0.2 Information source preference. A preliminary personalized search result ranking list is generated by combining topic relevance, sentiment matching, and information source preference. To ensure the diversity of search results, the maximum marginal relevance algorithm is used to adjust the ranking list, selecting items with high relevance and significant differences from the selected results to form the final personalized search result ranking list. The user inputs the search request "latest technology news," which is converted into word vectors [0.2, 0.5, -0.1, 0.3] using the Word2Vec model. The Naive Bayes classifier identifies the topic category as "technology" with a probability of 0.85. The sentiment tendency is determined to be neutral based on a dictionary method. The latest interest profile of user ID 12345 is obtained from the user interest profile database, with topic preference vectors of [0.3, 0.6, 0.1] (technology, entertainment, sports), sentiment tendency ratios of [0.6, 0.3, 0.1] (positive, negative, neutral), and frequently used information sources of ["Science and Technology Daily", " The search query topic "technology" was calculated to have a cosine similarity of 0.95 with the user's interest topic. The search ranking weight was adjusted based on this similarity, from an original weight of 0.5 to 0.5. (1 + 0.95) = 0.975. The sentiment of the search results is consistent with user preferences, and the weight is further adjusted to 0.975. 1.2 = 1.17. The content of the user's preferred information source, "Science and Technology Daily," is weighted. Assuming a news item has a topic relevance of 0.8, a sentiment match of 1, and an information source preference of 1, the final score is 0.5. 0.8 + 0.3 1+0.2 1=0.9. Using the maximum marginal relevance algorithm, with a diversity parameter λ=0.3, the 5 news items with high relevance and large differences are selected from the Top-10 results to form the final personalized search results ranking list.
[0078] Optionally, the method in this embodiment may further include: generating a user feedback signal based on the user's click behavior and page dwell time corresponding to the target retrieval result; and using a reinforcement learning model to iteratively optimize the user's historical retrieval strategy based on the user feedback signal to generate an optimized user information retrieval strategy, which is used to dynamically adjust the display ratio of retrieval results for different topic categories.
[0079] In some examples, a preset threshold is set for the information sources that users prefer to be recommended. Information sources that users have visited more than the threshold in the past are identified, and the content quality, update frequency and relevance of the identified information sources are scored. Priority recommendations are made based on the scores.
[0080] For example, such as Figure 11 As shown, the system retrieves the user ID, information source URL, and access count fields from the user's historical access record database. Based on these fields, it statistically calculates the access count for each information source. The statistical calculation results are used to determine a list of frequently accessed information sources. For the list of frequently accessed information sources, the system retrieves the latest content from each source. This latest content is analyzed using the NLTK natural language processing tool to obtain content keywords and topic features. Based on the content keywords and topic features, the system retrieves the historical update records of the information sources. These historical update records are used to calculate the average update time interval, resulting in an information source update frequency score. The system uses a cosine similarity algorithm to calculate the relevance score between the information source content and the user's interest profile. The user interest profile is constructed by analyzing the user's historical browsing content and search keywords. The content quality score, update frequency score, and relevance score are weighted and summed. If the weighted sum is greater than a preset threshold, the information source is determined to be a priority recommended information source. Priority recommended information sources are used to generate a list of recommended information sources.
[0081] For example, the system extracts fields such as user ID, accessed information source URL, and number of visits from the user's historical access record database. It then statistically analyzes the number of visits for each information source, calculating the average and standard deviation of visits over the past 30 days. The average plus the standard deviation is used as a preset threshold to filter out information sources with more than 20 visits. For instance, calculating the average and standard deviation of visits for all information sources over the past 30 days, assuming an average of 15 visits and a standard deviation of 5 visits, the preset threshold is 15 + 5 = 20 visits. After filtering, a list of information sources with more than 20 visits is obtained. For these high-frequency access sources, the Scrapy web crawler is used to obtain their latest content. The NLTK natural language processing tool is then used to analyze the content, extracting keywords, themes, and other features. The Flesch-Kincaid readability index is calculated as a basis for scoring content quality. Historical update records of information sources are obtained through RSS subscriptions or periodic crawling. The average update interval is calculated, and the update frequency is converted into a numerical score. Simultaneously, a cosine similarity algorithm is used to calculate the relevance score between the information source content and the user's interest profile. The user interest profile is constructed by analyzing the user's historical browsing content and search keywords. Historical update records of information sources are obtained through RSS subscriptions or periodic crawling (e.g., crawling once a day). Assuming that update records from the past 60 days are collected, and a total of 12 updates are found, then the average update interval is 60 / 12 = 5 days. The update frequency can be converted into a numerical score according to certain rules, such as an update interval of less than 3 days scoring 3 points, 3-7 days scoring 2 points, and more than 7 days scoring 1 point. Therefore, the update frequency score of this information source is 2 points.
[0082] In addition, user interest profiles are constructed by analyzing users' historical browsing content and search keywords. Assuming users' interest keywords are concentrated in "technological innovation" and "smart living," a cosine similarity algorithm is used to calculate the relevance score between information source content and the user's interest profile. Assuming that text content and user interests are represented by vectors, a relevance score of 0.8 is calculated, indicating a high relevance between the information source content and the user's interests. A weighted sum of 40% for content quality, 30% for update frequency, and 30% for relevance score is obtained to obtain a comprehensive score for each frequently accessed information source. Information sources are then sorted in descending order based on their comprehensive scores to generate a priority recommendation list. Assuming user ID is 12345, historical browsing records show that this user accessed 50 different information sources in the past 30 days, with an average of 10 visits and a standard deviation of 3. A threshold of 13 visits (10+3) is set to filter out 5 frequently accessed information sources. Scrapy is used to crawl the latest content from these 5 information sources, crawling 10 articles from each source. NLTK was used for text analysis to extract keywords and themes, such as "technology" accounting for 60% and "finance" for 40%. The Flesch-Kincaid readability index was calculated, ranging from 0 to 100. The average scores of the five sources were assumed to be 75, 80, 70, 85, and 65. Update frequency was obtained through RSS subscriptions, revealing average update intervals of 6 hours, 12 hours, 24 hours, 8 hours, and 48 hours, translating to ratings of 83, 67, 50, 75, and 33 out of 100. Based on user interest profiles ("technology" interest score 0.7, "finance" interest score 0.3), content relevance scores were calculated, yielding 90, 85, 70, 80, and 60. Finally, a weighted sum was performed, resulting in a total score of 75 for the first source. 0.4+83 0.3+90 0.3 = 81.9. Similar to calculating the scores of other sources, the final priority recommendation order is 81.9, 79.1, 64.0, 80.5, 54.4.
[0083] For example, when performing personalized information retrieval, the scenario type of the user's immediate search request is obtained, the user's demand level for information timeliness is determined according to the scenario type, the latest information is obtained from high-level timeliness information sources, and personalized search results are sorted in combination with comprehensive user preference representation.
[0084] In some examples, such as Figure 12As shown, the BERT model is used for natural language processing on the retrieval request to extract keywords and semantic features. Based on the keywords and semantic features, the random forest algorithm is used to identify the scenario type of the retrieval request, and a pre-established scenario-timeliness mapping table is queried to determine the timeliness requirement level of the retrieval request. According to the timeliness requirement level, the latest information is obtained from the corresponding information source, where high-level timeliness requirements are obtained from real-time updated news sources, medium-level requirements are obtained from daily updated websites, and low-level requirements are obtained from periodically updated knowledge bases. Combining the latest information and the pre-obtained user preference representation vector, the relevance between the latest information and the retrieval request is calculated, and the cosine similarity algorithm is used to calculate the matching degree between the latest information and user preferences. A weighted linear combination model is used to integrate timeliness, relevance, and user preference matching degree to generate a preliminary personalized search result ranking list. The maximum marginal relevance algorithm is used to re-rank the preliminary ranking list to form the final personalized search result ranking list.
[0085] For example, the BERT model is used for natural language processing on user-input search requests to extract keywords and semantic features. A random forest algorithm is used as a scene classifier to identify the scene type of the request, such as emergency events, daily queries, or learning and research. Based on the identified scene type, a scene-timeliness mapping table constructed based on expert experience and historical data analysis is queried to determine the user's level of information timeliness needs, categorizing them into high, medium, and low levels. For different timeliness needs, the latest information is obtained from corresponding information sources: high-level timeliness needs are met by real-time updated news sources and social media platforms; medium-level needs by daily updated websites; and low-level needs by periodically updated academic and knowledge bases. The difference between the information's publication time and the current time is used to determine the information's timeliness, filtering out information that meets the timeliness requirements. Combining the latest information and the user's comprehensive preference representation vector, the TF-IDF algorithm is used to calculate the relevance between the information and the search request, and the cosine similarity algorithm is used to calculate the degree of matching between the information and the user's preferences. A weighted linear combination model is used to comprehensively consider timeliness, relevance, and user preference matching to generate a preliminary personalized search result ranking list. The ranking list is then reordered using the maximum marginal relevance algorithm to improve result diversity, forming the final personalized search result ranking list. Assuming a user inputs the search request "latest typhoon path data," the BERT model is used to extract keywords ["typhoon," "path," "data"] and semantic feature vectors [-0.2, 0.5, 0.3]. A random forest classifier with 100 decision trees and a maximum depth of 10 identifies the request as an "emergency event" scenario with a confidence level of 0.85. A query scenario-timeliness mapping table is used to determine the timeliness requirement level as "high." The latest information is obtained from real-time updated news sources and social media platforms, filtering out 10 relevant items published within the past 2 hours. The TF-IDF algorithm is used to calculate the relevance of the information to the search request, with a score ranging from 0 to 1. Assuming the user's overall preference representation vector is [0.3, 0.2, 0.5] (topic, sentiment, authority), the cosine similarity algorithm is used to calculate the matching degree between information and user preferences, with a score range of 0-1. A weighted linear combination model is adopted, with weights allocated as follows: timeliness 0.4, relevance 0.3, and user preference 0.3, to calculate the overall score for each piece of information. For example, if a piece of information has a timeliness of 0.9, was published 30 minutes ago, has a relevance of 0.8, and a user preference matching degree of 0.7, then its overall score is 0.9 + 0.4 + 0.8 + 0.3 + 0.7. 0.3 = 0.81. The initial Top-10 results are re-ranked using the maximum marginal relevance algorithm (λ = 0.7) to obtain the final personalized search results ranking list.
[0086] In some examples, a preset threshold is set for the timeliness requirement level. For scenarios with timeliness requirements higher than the threshold level, the ranking weight of instant information is increased, and for scenarios with timeliness requirements lower than the threshold level, the weight of instant information is decreased, so as to match the differentiated timeliness requirements of users in different scenarios.
[0087] For example, such as Figure 13 As shown, the random forest algorithm can also be used to classify historical search records extracted from the user behavior database into scenarios. Scenario classification quantifies timeliness preferences through user click behavior and dwell time to obtain a scenario-timeliness demand mapping table. Different sensitivity levels of desensitization rules are defined through this mapping table. For example, core user behavior features are replaced with random values, while non-sensitive fields such as timestamps retain their original format to maintain data distribution characteristics. Based on the scenario-timeliness demand mapping table, a percentile method is used to set a timeliness demand level threshold, which divides the timeliness demand level into high, medium, and low levels. The BERT model is used to extract semantic features from the user's current search request. These semantic features are used for scenario identification and querying the scenario-timeliness demand mapping table to obtain the corresponding timeliness demand level. If the timeliness demand level is higher than the preset threshold, the sigmoid function is used to map the timeliness demand level to a weight adjustment coefficient, which is used to dynamically adjust the ranking weight of real-time information. By introducing historical user behavior data, the weight adjustment coefficient is fine-tuned to obtain personalized search result rankings adapted to different scenarios. These personalized search result rankings are used to generate the final search results.
[0088] For example, historical search records are extracted from a user behavior database, and the random forest algorithm is used to classify these records by scenario. Timeliness preferences are quantified based on user click behavior and dwell time, establishing a mapping table between scenarios and timeliness requirements. Based on this mapping table, a percentile method is used to set a threshold for timeliness requirements, dividing them into high, medium, and low levels, and determining the corresponding weight adjustment coefficient for each level. For the user's current search request, a BERT model is used to extract semantic features for scenario identification. The corresponding timeliness requirement level is obtained by querying the scenario-timeliness requirement mapping table, and the obtained level is compared with the preset level threshold. Based on the comparison result, the sigmoid function is used to map the timeliness requirement level to the weight adjustment coefficient, dynamically adjusting the ranking weight of immediate information. Scenarios above the threshold have a higher weight for immediate information, while scenarios below the threshold have a lower weight. Historical user behavior data is introduced to fine-tune the weight adjustment, generating personalized search result rankings adapted to different scenarios. The effect of the weight adjustment is evaluated through A / B testing, user feedback data is collected, and the scenario-timeliness requirement mapping table and weight adjustment parameters are updated regularly. Assuming user ID is 12345, 1000 data points are extracted from their historical search records. Using a random forest algorithm (100 decision trees, maximum depth 10), scenarios are classified into three categories: "Emergency Events," "Daily Queries," and "Study & Research." By analyzing the average click interval and page dwell time for each scenario, timeliness preferences are quantified. For example, in the "Emergency Events" scenario, the average click interval is 30 seconds and the dwell time is 2 minutes, corresponding to a high timeliness requirement. Using the 75th and 25th percentiles as thresholds, timeliness requirements are divided into high, medium, and low levels, with corresponding weight adjustment coefficients of 1.5, 1.0, and 0.5, respectively. When a user inputs the search request "latest earthquake news," the BERT model extracts the semantic feature vector [-0.2, 0.5, 0.3], identifying it as an "Emergency Events" scenario, and a high timeliness requirement level is obtained by looking up a table. The sigmoid function is used to map the timeliness requirement level 6 (out of 10) to a weight adjustment coefficient of 1.3. Considering that "emergency events" accounted for 40% of user queries in the past week, the weight adjustment coefficient was fine-tuned to 1.35. For real-time information in the search results, such as news published within 10 minutes, their original ranking score was multiplied by 1.35 to generate new ranking results. A / B testing was conducted to compare the click-through rate before and after the adjustment, revealing a 5% increase. This feedback data was used to update the scenario-timeliness requirement mapping table and optimize the weight adjustment parameters.
[0089] In some examples, considering that the performance and user experience of a personalized information retrieval system directly determine the service quality of an intelligent assistant, various collaborative improvement methods can be adopted: continuously optimize the weight allocation and ranking logic in the recommendation algorithm through explicit or implicit user feedback loops; add interpretability design to the result display layer, such as labeling the recommendation reasons because you frequently follow technology-related content to improve transparency and user trust; and establish an A / B testing and online experimentation framework to continuously verify and optimize different parameter configurations, ensuring that the system can provide high-quality personalized services that meet user expectations, while avoiding resource waste caused by over-recommendation or redundant computation, and retaining sensitivity to abnormal behavior or security-related signals in the overall architecture, so that the system can still support the security team's timely response to truly critical alerts while focusing on personalization.
[0090] Step 207: Sort the candidate search results obtained from the preliminary search according to the adjusted search ranking weights to generate the target search results.
[0091] In some examples, a personalized information retrieval strategy is generated based on the personalized search results ranking. The intelligent assistant then uses this user information retrieval strategy learned from the preferences to make personalized recommendations for search content.
[0092] For example, such as Figure 14 As shown, the system retrieves user search requests, click behavior, and browsing duration data from a database of historical user search records. Based on this user data, combined with content features and user behavior, an interest vector is generated. Random letters / numbers are then used to replace real content for search keywords, clicked URLs, and other behavioral data. For example, the search term "medical report" is replaced with "technology product," maintaining text format while eliminating sensitive semantics. Based on the user interest vector and personalized search result ranking, a personalized information retrieval strategy model is constructed using the random forest algorithm. This model is then input into the deep reinforcement learning framework of the intelligent assistant. The DQN algorithm iterates repeatedly and simulates user interactions to obtain an optimized recommendation strategy for the intelligent assistant. Based on the learned user information retrieval strategy and real-time user behavior data, the intelligent assistant dynamically adjusts the topic distribution, information source selection, and display order of recommended content using a multi-armed slot machine algorithm. If the user is new or there is new content, a content-based recommendation method is used to handle the cold start problem. Initial recommendation results are generated by extracting content features and user profile information. The workflow of the intelligent assistant is shown in the diagram below. Figure 15 As shown, the workflow diagram for intelligent assistant processing is as follows: Figure 16 As shown in the flowchart, the interaction pattern between the intelligent assistant and the user is as follows: Figure 17 As shown in the diagram, the learning process of the intelligent assistant is illustrated below. Figure 18 As shown.
[0093] For example, data such as user search requests, click behaviors, and browsing duration are extracted from the user's historical search record database. A matrix factorization algorithm is used to analyze user interests and preferences by combining content features and user behavior, generating user interest vectors. Based on these user interest vectors and personalized search result ranking, a random forest algorithm is used to construct a personalized information retrieval strategy model, which includes search keyword expansion, result filtering, and ranking rules. This generated information retrieval strategy model is input into the deep reinforcement learning framework of the intelligent assistant. The DQN algorithm is used to iterate and simulate user interactions to optimize the intelligent assistant's recommendation strategy, using user click-through rate and dwell time as reward signals to update the model. Based on the learned user information retrieval strategy and real-time user behavior data, the intelligent assistant uses a multi-armed slot machine algorithm to dynamically adjust the topic distribution, information source selection, and display order of recommended content, generating a personalized list of recommended search content. To address the cold start problem, a content-based recommendation method is used to handle new users or new content. Initial recommendation results are generated by extracting content features and user profile information. Assuming user ID is 12345, 1000 data points are extracted from their historical search records, including search keywords, clicked URLs, and browsing duration. Using a matrix factorization algorithm, the user-content interaction matrix is decomposed into 100-dimensional user latent vectors and content latent vectors. Combined with content features extracted using TF-IDF, a user interest vector [0.8, 0.5, 0.3, ...] is generated. Based on this vector, a retrieval strategy model is constructed using a random forest algorithm (100 decision trees, maximum depth 15), such as expanding "technology news" to "latest technology information" and "technology innovation updates," filtering out results published more than 7 days ago, and sorting by relevance and timeliness. This model is then input into the DQN algorithm, a 3-layer neural network with 128 neurons per layer. A discount factor γ = 0.9 is set, and an ε-greedy strategy is used, with ε linearly decaying from 0.9 to 0.1 during exploration. The reward is a weighted sum of user click-through rate and average dwell time (in seconds), such as 0.7 click-through rate + 0.3 log(dwell time). After 10,000 iterations, DQN converges to a stable strategy. Using a converged strategy, the intelligent assistant applies a multi-armed slot machine algorithm (UCB1 variant, confidence parameter α=1.5) to the real-time recommended content, selecting the top-3 from 10 topic categories for display and dynamically adjusting the display ratio of each topic. For new users, features such as age and occupation are extracted from their registration information, and an initial recommendation list is generated using a content-based recommendation method, such as recommending 70% technology, 20% finance, and 10% entertainment content to a 25-year-old IT professional.
[0094] Compared to existing technologies, this embodiment utilizes a structured user behavior time-series dataset and extracts the topic categories and sentiment tendencies associated with keywords using natural language processing techniques. It statistically analyzes the distribution probability and sentiment ratio of each topic category, and performs weighted calculations based on the information source type clicked by the user, authority rating, and browsing duration to form a dynamically updated multi-dimensional user interest profile time series. Specific quantitative indicators, such as the distance between topic category distribution probabilities and the difference between positive and negative sentiment ratios, are designed to measure the magnitude of changes in different dimensions of the user interest profile time series. A change detection mechanism is implemented to determine significant changes in user interests based on the magnitude of these quantitative indicators and to identify the significance level of the changes. An automatic triggering mechanism is developed to initiate a user interest preference update process when a significant change in user interests is detected. The update mechanism involves adjusting the user interest model parameters to ensure that the personalized information retrieval system can respond promptly to the latest changes in user interests. During actual retrieval, the system can match the user's latest interest profile and dynamically adjust the ranking weight of search results. Based on the information source type preferred by the user and the timeliness requirements of the scenario, the presentation order of search results is optimized to ensure that the most relevant content is displayed to the user first.
[0095] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides an information retrieval device, such as... Figure 19 As shown, the device includes: an acquisition module 31, an analysis module 32, a construction module 33, an adjustment module 34, and a generation module 35.
[0096] Module 31 is configured to acquire retrieval request information; Analysis module 32 is configured to analyze the retrieval request information and identify the topic category and sentiment attribute corresponding to the information retrieval request; Module 33 is configured to construct a user interest profile based on the topic category and the sentiment attribute; The adjustment module 34 is configured to dynamically adjust the search ranking weight of each topic based on the user interest profile. The generation module 35 is configured to sort the candidate search results obtained from the preliminary search according to the adjusted search ranking weights, and generate the target search results.
[0097] In some examples of this embodiment, the analysis module 32 is further configured to clean and standardize the retrieval request information to generate user behavior time-series data; extract each retrieval keyword within a preset time window based on the user behavior time-series data; determine the topic category mapped by each retrieval keyword and its corresponding sentiment tendency according to each retrieval keyword; calculate the positive and negative ratio of sentiment tendency based on the frequency and / or weight of the occurrence of topic category and sentiment tendency within each time window; and determine the sentiment tendency attribute based on the positive and negative ratio of sentiment tendency.
[0098] In some examples of this embodiment, the construction module 33 is further configured to generate a comprehensive score of information sources corresponding to the topic based on the authority rating of the information sources clicked by the user under the topic category and the corresponding page browsing time; to perform a fusion modeling of the topic category and the sentiment tendency attribute based on the comprehensive score of information sources and the user's behavioral characteristics to obtain a user interest model; and to construct the user interest profile based on the user interest model.
[0099] In some examples of this embodiment, the construction module 33 is further configured to acquire historical behavioral data of users in multiple dimensions based on the comprehensive score of information sources and the user's behavioral characteristics. The multiple dimensions include the topic category dimension, sentiment tendency dimension, and information source authority dimension. By analyzing the influence of historical behavioral data in multiple dimensions on user decision-making behavior, the initial weight parameters of each dimension are determined. The feature values of each dimension are normalized to generate a normalized feature vector. The normalized feature vector is weighted and fused with the corresponding initial weight parameters to generate a comprehensive user preference representation vector. Based on the comprehensive user preference representation vector, the user interest model is constructed.
[0100] In some examples of this embodiment, the construction module 33 is further configured to convert the topic distribution probability, sentiment tendency ratio and comprehensive information source score of each topic corresponding to each time window into a structured feature vector; arrange the structured feature vectors in chronological order to form a user interest profile time series, which is used to characterize the dynamic evolution process of user interests; and construct the user interest profile based on the user interest profile time series.
[0101] In some examples of this embodiment, the construction module 33 is further configured to obtain the magnitude of change of user interest preferences in each dimension based on the feature values of each dimension in the user interest profile time series; when the magnitude of change in any dimension exceeds a preset significance level threshold, the parameters associated with the corresponding dimension in the user interest model are dynamically adjusted.
[0102] In some examples of this embodiment, the construction module 33 is further configured to obtain the target feature vector corresponding to the target dimension whose change exceeds the significance level threshold; based on the target feature vector, retrain the user interest model parameters corresponding to the target dimension using a preset machine learning model; and perform weighted fusion of the updated model parameters obtained from the training with the historical model parameters.
[0103] In some examples of this embodiment, the adjustment module 34 is further configured to perform time decay adjustment on the interest intensity of each topic in the user interest profile based on a preset period; use the interest intensity of each topic after time decay adjustment as the basic weight of the corresponding topic in the search ranking; and adjust the basic weight in combination with the similarity between the user's current search request and the interest profile to obtain the search ranking weight.
[0104] In some examples of this embodiment, the adjustment module 34 is further configured to filter out high-frequency access information sources based on the user's historical access frequency to each information source; to perform multi-dimensional evaluation of the high-frequency access information sources in terms of content quality, update frequency, and relevance to the user's interest profile, and to generate a comprehensive score; to construct a priority recommendation information source sequence based on the comprehensive score, and to improve the ranking priority of the result items in the candidate search results that belong to the priority recommendation information source sequence.
[0105] In some examples of this embodiment, the adjustment module 34 is further configured to perform semantic parsing and feature extraction on the retrieval request information, identify the scenario type corresponding to the retrieval request information, wherein the scenario type is used to characterize the user's sensitivity to the timeliness of the retrieval results; determine the timeliness requirement level corresponding to the current scenario type based on the mapping relationship between the scenario type and the information timeliness requirement level; when the timeliness requirement level meets the preset high timeliness condition, obtain the instant information provided by the instant information source and increase the retrieval ranking weight of the instant information; when the timeliness requirement level meets the preset low timeliness condition, decrease the retrieval ranking weight of the instant information.
[0106] In some examples of this embodiment, the generation module 35 is further configured to generate a user feedback signal based on the user click behavior and page dwell time corresponding to the target retrieval result; based on the user feedback signal, the user's historical retrieval strategy is iteratively optimized using a reinforcement learning model to generate an optimized user information retrieval strategy, which is used to dynamically adjust the display ratio of retrieval results for different topic categories.
[0107] It should be noted that for other corresponding descriptions of the various functional units involved in the information retrieval device provided in this embodiment, please refer to... Figure 1 and Figure 2The corresponding descriptions in [the document] will not be repeated here.
[0108] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method shown.
[0109] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0110] Based on the above, Figure 1 and Figure 2 The method shown, and Figure 19 To achieve the above objectives, the present application also provides an electronic device, such as a personal computer, server, laptop computer, intelligent robot, or other intelligent terminal, as illustrated in the virtual device embodiment. This device includes a storage medium and a processor; the storage medium stores a computer program; and the processor executes the computer program to implement the above-described virtual device. Figure 1 and Figure 2 The method shown.
[0111] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0112] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0113] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0114] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms, or it can be implemented by hardware. By applying the solution of this embodiment, compared with the existing technology, this embodiment is more comprehensive and refined in acquiring and processing user behavior data. It not only collects multi-dimensional data such as search keywords, clicked information source types, information source authority scores, and page browsing time, but also ensures data quality and consistency through advanced data cleaning and standardization processing, forming a structured user behavior time series dataset. On this basis, it further extracts the topic categories and sentiment tendencies involved in keywords within different time windows, statistically analyzes the distribution probability of topic categories and the positive and negative proportions of sentiment tendencies, and combines them with the information source type clicked by the user, authority score, and browsing time for weighted fusion to construct a multi-dimensional user interest profile time series. The proposal introduces a more intelligent quantitative indicator design and a significant change detection mechanism. For different dimensions of the user interest profile time series, it designs quantitative indicators such as the probability distance of topic category distribution and the difference between the positive and negative proportions of sentiment tendencies. By calculating the magnitude of change of these indicators in different time windows, it determines whether the user's interest preferences in each dimension have changed significantly, and determines the significance level based on the magnitude of the change. When the change in interest in a certain dimension is detected to exceed the preset threshold, the system will automatically trigger the user interest preference update mechanism for that dimension and dynamically adjust the corresponding model parameters, thereby ensuring that the personalized model can respond to the evolution of user interests in a timely manner. During personalized information retrieval, the system acquires the topic category and sentiment attributes of the current search request, matches it with the user's latest interest profile, dynamically adjusts the search ranking weight of different topics, and prioritizes recommending information source types preferred by the user. Simultaneously, by identifying the scenario type of the search request, the system determines the user's level of information timeliness requirement, increasing the ranking weight of immediate information for scenarios with high timeliness requirements and decreasing it for scenarios with low timeliness requirements. The system retrieves the latest content from high-level timeliness information sources and integrates comprehensive user preference representations to rank search results, effectively meeting users' differentiated needs for information relevance and timeliness in different usage scenarios.
[0115] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0116] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. An information retrieval method, characterized in that, include: Obtain search request information; The search request information is analyzed to identify the topic category and sentiment attribute corresponding to the information search request; Based on the topic categories and the sentiment attributes, construct user interest profiles; Based on the user interest profile, the search ranking weight of each topic is dynamically adjusted; The candidate search results obtained from the initial search are sorted according to the adjusted search ranking weights to generate the target search results.
2. The method according to claim 1, characterized in that, The step of analyzing the retrieval request information and identifying the topic category and sentiment attribute corresponding to the information retrieval request includes: The retrieval request information is cleaned and standardized to generate user behavior time-series data; Based on the user behavior time-series data, each search keyword is extracted within a preset time window; Based on each search keyword, determine the topic category mapped by the keyword and its corresponding sentiment tendency; The positive and negative proportions of sentiment are calculated based on the frequency and / or weight of the occurrence of topic categories and sentiment tendencies within each time window. The emotional tendency attribute is determined based on the positive and negative proportions of the stated emotional tendency.
3. The method according to claim 1, characterized in that, The process of constructing a user interest profile based on the topic category and the sentiment attribute includes: Based on the authority rating of the information source clicked by the user under the topic category and the corresponding page browsing time, a comprehensive score for the information source corresponding to the topic is generated. Based on the comprehensive score of information sources and the user's behavioral characteristics, the topic category and the sentiment tendency attribute are fused and modeled to obtain the user interest model; Based on the user interest model, the user interest profile is constructed.
4. The method according to claim 3, characterized in that, The method of fusing and modeling the topic category and the sentiment attribute based on the comprehensive score of information sources and user behavioral characteristics to obtain a user interest model includes: Based on the comprehensive score of information sources and the user's behavioral characteristics, historical behavioral data of users in multiple dimensions are obtained, including topic category dimension, sentiment tendency dimension and information source authority dimension. By analyzing the impact of historical behavioral data from multiple dimensions on user decision-making behavior, the initial weight parameters for each dimension are determined. The feature values of each dimension are normalized to generate normalized feature vectors; The normalized feature vector is weighted and fused with the corresponding initial weight parameters to generate a comprehensive user preference representation vector. The user interest model is constructed based on the comprehensive user preference representation vector.
5. The method according to claim 3, characterized in that, The process of constructing the user interest profile based on the user interest model includes: The topic distribution probability, sentiment tendency ratio, and comprehensive information source score of each topic corresponding to each time window are transformed into structured feature vectors. The structured feature vectors are arranged in chronological order to form a time series of user interest profiles, which is used to characterize the dynamic evolution of user interests. The user interest profile is constructed based on the time series of the user interest profile.
6. The method according to claim 3, characterized in that, The method further includes: Based on the feature values of each dimension in the time series of the user interest profile, the magnitude of change of user interest preferences in each dimension is obtained; When the change in any dimension exceeds a preset significance level threshold, the parameters associated with the corresponding dimension in the user interest model are dynamically adjusted.
7. The method according to claim 6, characterized in that, The dynamic adjustment of parameters associated with the corresponding dimensions in the user interest model includes: Obtain the target feature vector corresponding to the target dimension whose change exceeds the significance level threshold; Based on the target feature vector, the user interest model parameters corresponding to the target dimension are retrained using a preset machine learning model; The updated model parameters obtained from training are weighted and fused with the historical model parameters.
8. The method according to claim 1, characterized in that, The step of dynamically adjusting the search ranking weights of each topic based on the user interest profile includes: Based on a preset period, the intensity of interest in each topic in the user interest profile is adjusted by time decay. The intensity of interest in each topic after time decay adjustment will be used as the basic weight of the corresponding topic in the search ranking. The search ranking weight is obtained by adjusting the basic weight based on the similarity between the user's current search request and the interest profile.
9. The method according to claim 1, characterized in that, The method further includes: High-frequency access information sources are selected based on the user's historical access frequency to each information source; The high-frequency access information sources are evaluated from multiple dimensions, including content quality, update frequency, and relevance to user interest profiles, to generate a comprehensive score. Based on the comprehensive score, a priority recommendation information source sequence is constructed, and the ranking priority of result items belonging to the priority recommendation information source sequence in the candidate search results is increased.
10. The method according to claim 1, characterized in that, The method further includes: The retrieval request information is semantically parsed and features are extracted to identify the scenario type corresponding to the retrieval request information. The scenario type is used to characterize the user's sensitivity to the timeliness of the retrieval results. Based on the mapping relationship between the scenario type and the information timeliness requirement level, the timeliness requirement level corresponding to the current scenario type is determined; When the timeliness requirement level meets the preset high timeliness condition, the instant information provided by the instant information source is obtained, and the retrieval ranking weight of the instant information is increased. When the timeliness requirement level meets the preset low timeliness condition, the retrieval ranking weight of the real-time information is reduced.
11. The method according to claim 1, characterized in that, The method further includes: Based on the user click behavior and page dwell time corresponding to the target search results, a user feedback signal is generated; Based on the user feedback signals, a reinforcement learning model is used to iteratively optimize the user's historical retrieval strategy to generate an optimized user information retrieval strategy. This user information retrieval strategy is used to dynamically adjust the display ratio of retrieval results for different topic categories.
12. An information retrieval device, characterized in that, include: The acquisition module is configured to acquire retrieval request information; The analysis module is configured to analyze the retrieval request information and identify the topic category and sentiment attribute corresponding to the information retrieval request. The construction module is configured to build a user interest profile based on the topic category and the sentiment attribute; The adjustment module is configured to dynamically adjust the search ranking weight of each topic based on the user interest profile. The generation module is configured to sort the candidate search results obtained from the initial search based on the adjusted search ranking weights, and generate the target search results.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 11.
14. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 11.
15. A computer program product having a computer program stored thereon, characterized in that, When the computer program product is executed by a processor, it implements the method of any one of claims 1 to 11.