Content recommendation method and device, electronic equipment and storage medium
By building user portraits and predictive scores based on multi-source data, the problem of insufficient diversity of content recommendations in the prior art is solved, and more accurate and diversified content recommendations are achieved.
Patent Information
- Application Number
- CN202510306340.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-06
AI Technical Summary
The content recommendation method based on user historical data in the prior art leads to insufficient diversity of recommendation results and the inability to accurately match user needs, resulting in low recommendation accuracy.
By obtaining the interactive data of the target user, building user portrait data, and determining the target user's prediction score for unknown content based on the user's historical content rating, similarity with other users, and other users' ratings for different content, the target user's prediction score for unknown content is thus made.
It achieves the exact matching of recommended content and user needs, improves the accuracy of recommendations, and increases the diversity of recommendation results.
Smart Images

Figure CN120106227A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a content recommendation method, device, electronic device and storage medium. Background Art
[0002] The existing content recommendation process is generally based on data obtained from users' historical interactions, such as historical behaviors or preferences.
[0003] Existing recommendation methods based on historical data often tend to recommend content that users already like or frequently visit, which may lead to insufficient diversity in recommendation results and the recommended content cannot match user needs, resulting in low recommendation accuracy. Summary of the invention
[0004] The present invention provides a content recommendation method, device, electronic device and storage medium, which are used to solve the defect in the prior art that the recommendation results are not diverse enough based on recommending content that users already like or frequently visit, achieve accurate matching of recommended content with user needs, and improve recommendation accuracy.
[0005] The present invention provides a content recommendation method, comprising the following steps: Acquire the target user's interaction data, and construct user portrait data of the target user based on the features extracted from the interaction data, wherein the interaction data includes audio data, video data, and user behavior data; Determining the predicted score of the target user for unknown content based on the historical content score of the target user, the similarity between the target user and other users, and the scores of the other users on different contents, wherein the similarity between the target user and the other users is determined based on the similarity between the user portrait data of the target user and the user portrait data of the other users; Based on the predicted score, content is recommended to the target user.
[0006] According to a content recommendation method provided by the present invention, the method of determining the predicted score of the target user for unknown content based on the historical content score of the target user, the similarity between the target user and other users, and the scores of the other users for different content includes: Determining a set of neighboring users similar to the target user based on the similarity between the user portrait data of the target user and the user portrait data of the other users; Based on the similarity between the users in the neighboring user set and the target user, the average rating of the target user on historical content, and the ratings of the users in the neighboring user set on different content, the predicted rating of the target user on unknown content is determined.
[0007] According to a content recommendation method provided by the present invention, the target user's predicted score for unknown content is: ; in, The target user For unknown content The prediction score of The target user Average rating for historical content, The target user With users The similarity between Is a user About content Ratings, Is a user The average rating of different content. Is with the target user A collection of similar nearby users.
[0008] According to a content recommendation method provided by the present invention, based on the features extracted from the interaction data, constructing user portrait data of the target user, comprising: Performing text extraction on the interaction data to obtain text data; Performing sentiment analysis on the text data to obtain sentiment analysis results, and performing context analysis on the text data to obtain context analysis results; Feature extraction is performed on the sentiment analysis result, the context analysis result, and the text data, and target user portrait data of the target user is constructed based on the extracted features.
[0009] According to a content recommendation method provided by the present invention, after acquiring the interaction data of the target user, the method further includes: Based on a noise detection algorithm, analyzing the background sound in the audio data segment by segment; Based on the analysis result of the background sound, it is determined that the target user is not listening to the target audio segment of the audio data, and the audio data corresponding to the target audio segment is deleted.
[0010] A content recommendation method provided by the present invention further includes: Acquire the real-time interaction data of the target user, perform feature extraction based on the real-time interaction data, and obtain the real-time interaction features of the target user; Based on the real-time interaction features, the user portrait data of the target user is updated.
[0011] The present invention also provides a content recommendation device, comprising the following modules: A data acquisition module is used to acquire the interaction data of the target user and construct user portrait data of the target user based on the features extracted from the interaction data, wherein the interaction data includes audio data, video data and user behavior data; A rating prediction module, used to determine the target user's predicted rating for unknown content based on the target user's historical content rating, the target user's similarity with other users, and the other users' ratings of different content, wherein the target user's similarity with other users is determined based on the similarity between the target user's user portrait data and the other users' user portrait data; A content recommendation module is used to recommend content to the target user based on the predicted score.
[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, any of the above-mentioned content recommendation methods is implemented.
[0013] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the content recommendation method described above is implemented.
[0014] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the content recommendation method described above is implemented.
[0015] The content recommendation method, device, electronic device and storage medium provided by the present invention integrate multi-source data such as the target user's own historical rating, similarity with other users, and ratings of other users. These data complement each other and provide a basis for predicting ratings from different angles. Compared with relying on a single data source, they can more comprehensively understand user needs and content characteristics and improve the accuracy of predictions. By analyzing the rating data of a large number of users, it is possible to discover potential similarities between users and associations between content. Even if the target user has not rated certain content, the ratings of similar users can be used to discover the target user's potential interest in these contents, so that the user can find new content that may be of interest but has not yet been exposed to, thereby improving the diversity of recommendation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0017] Figure 1 It is a flowchart of the content recommendation method provided by the present invention.
[0018] Figure 2 It is a schematic diagram of the structure of a device applying the content recommendation method provided by the present invention.
[0019] Figure 3 It is a structural schematic diagram of the content recommendation device provided by the present invention.
[0020] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] Figure 1 is a flow chart of the content recommendation method provided by the present invention, such as Figure 1 As shown, the method includes the following: Step 110, acquiring interaction data of a target user, and constructing user portrait data of the target user based on features extracted from the interaction data, wherein the interaction data includes audio data, video data, and user behavior data; Step 120, determining a predicted score of the target user for unknown content based on the historical content score of the target user, the similarity between the target user and other users, and the scores of the other users for different contents, wherein the similarity between the target user and the other users is determined based on the similarity between the user portrait data of the target user and the user portrait data of the other users; Step 130: Recommend content to the target user based on the predicted score.
[0023] The execution subject of the content recommendation method provided by the present invention may be an electronic device, a component in an electronic device, an integrated circuit, or a chip. The electronic device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a PDA, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), or a personal computer (PC), etc., which is not specifically limited by the present invention.
[0024] The following takes a computer executing the content recommendation method provided by the present invention as an example to describe the technical solution of the present invention in detail.
[0025] In step 110, the interaction data of the target user is obtained, and based on the features extracted from the interaction data, the user portrait data of the target user is constructed, wherein the interaction data includes audio data, video data and user behavior data.
[0026] Based on business needs, you can determine which data sources you need to obtain interaction data from, such as audio and video platforms, social media APIs, user devices, etc.
[0027] Among them, audio data and video data: including audio and video files, are used for subsequent content analysis, sentiment analysis, etc.
[0028] User behavior data: such as the target user’s browsing history, likes, comments, sharing and other behavior logs on the audio and video platform. Including text content such as posts, comments, tags, etc. on social media platforms.
[0029] Video data can be images extracted from video frames for image recognition or content analysis. Audio data can be the audio portion extracted from an audio file for speech recognition, sentiment analysis, etc.
[0030] The interactive data scenario can be an online video platform. When the target user watches a video on the platform, the emotional information and contextual information are extracted by analyzing the audio and image content in the video. The target user interacts with the system through voice commands. The system understands the user's needs through audio analysis and provides more accurate feedback in combination with image content (such as scenes in video frames). The system can also analyze the audio and image content uploaded by users to extract emotional tendencies and trend information for advertising and market analysis. The system can also provide students with personalized learning paths and feedback by analyzing the audio and image content in teaching videos. Use technical means such as API calls, database queries, and network crawling to establish a secure and stable connection with the data source. Extract the required audio and video files, user behavior logs, social media posts and other data from the data source.
[0031] During the data transmission process after obtaining the interactive data, encryption technologies such as SSL / TLS can be used to ensure the security of the data during transmission, including but not limited to using the HTTPS protocol for API calls and using encrypted connection strings for database connections. Check whether there is invalid data such as duplicate data, missing data, and abnormal data in the data set. Delete the identified invalid data to ensure data quality. For missing data, adopt appropriate filling strategies such as mean filling, interpolation, model prediction, etc. to maintain data integrity. In the process of data extraction and cleaning, strictly adhere to the principle of data minimization, only collect and process data directly related to business goals, and avoid collecting irrelevant or redundant data.
[0032] After acquiring the interaction data, convert the data in different formats into a unified format for subsequent processing and analysis. This includes encoding conversion of text or numerical data to ensure data consistency and readability. For data containing personally identifiable information, anonymization is performed, such as using hash functions, tokenization and other technologies to ensure that the processed data cannot be associated with the original data, thereby protecting user privacy. For high-dimensional data, such as text content in social media data, text mining, feature extraction and other technologies are used to reduce dimensionality, extract key information, and reduce the complexity of data processing. For data of different dimensions or ranges, normalization is performed to eliminate dimensional differences between data and improve the accuracy and efficiency of data processing. For data that needs to be aggregated, such as user behavior data, aggregation is performed according to dimensions such as time and user ID for subsequent analysis.
[0033] The pre-processed data can be first subjected to text recognition to identify text from audio data, video data, and user behavior data to obtain text data. Specifically, the text and voice content in the audio and video can be identified through a deep learning algorithm to obtain text data.
[0034] Based on feature extraction of the acquired text data, user portrait data of the target user can be constructed.
[0035] The features corresponding to the speech data extracted from the text data and the features corresponding to the image data can be fused using methods such as splicing, weighted summation, and attention mechanism. Splicing: directly splice the speech features and image features together to form a longer feature vector. Weighted summation: perform weighted summation of speech features and image features according to the weights to obtain a comprehensive feature vector. Weighted summation formula: F = w1×F_speech + w2×F_image, where F is the comprehensive feature vector, F_speech is the speech feature vector, F_image is the image feature vector, and w1 and w2 are weights. Attention mechanism: by calculating the attention weight, the speech features and image features are weightedly fused.
[0036] Adaptive learning and encoding processing: Use deep learning models to identify the encoding format of audio and video data and adaptively adjust the processing strategy. Encoding recognition can use a classification model based on a convolutional neural network (CNN). Adjust the parameters of the speech recognition model or select the corresponding language model based on the recognized language characteristics.
[0037] Speech-to-text and subtitle extraction: Use deep learning models (such as recurrent neural networks (RNNs), Transformers, etc.) to convert speech signals into text. Speech recognition models are usually trained using the CTC (Connectionist Temporal Classification) loss function. Use optical character recognition (OCR) technology to extract text information (such as subtitles, tags, etc.) from video frames.
[0038] Contextual understanding and sentiment analysis use natural language processing (NLP) technology to understand the context of text content. Word embedding: Convert words in the text into vector representations. Methods such as Word2Vec and GloVe can be used. Syntactic analysis: Analyze the syntactic structure of the text. A syntactic analyzer based on deep learning can be used. Semantic role labeling: Label the semantic roles in the text to understand the semantic structure of the text. Use sentiment analysis models to analyze the sentiment color of the text content. Sentiment analysis models can use sentiment classifiers based on LSTM or Transformer. The output of a sentiment classifier is usually a probability distribution, which indicates the probability that the text belongs to different sentiment categories.
[0039] Model optimization and training: Use strategies such as data enhancement, transfer learning, and self-supervised learning to improve the generalization ability and training efficiency of the model. Tune the model's hyperparameters to achieve optimal performance. Hyperparameter tuning can use methods such as grid search and random search. After processing, the analysis results (including recognized text, voice content, context understanding, and sentiment analysis results) are used as an important basis for subsequent modules to perform tasks such as content recommendation, sentiment analysis, and user behavior prediction.
[0040] In step 120, the predicted score of the target user for unknown content is determined based on the historical content score of the target user, the similarity between the target user and other users, and the scores of the other users for different contents. The similarity between the target user and the other users is determined based on the similarity between the user portrait data of the target user and the user portrait data of the other users.
[0041] It should be noted that other users refer to all other users in the system except the target user. The target user refers to the specific user for whom content is currently being recommended, while other users are concepts relative to the target user. Their relevant data (such as user portrait data, content ratings, etc.) are used to assist in analyzing and predicting the target user's preference for unknown content.
[0042] Specifically, a rating matrix may be constructed based on the user portrait data of other users and the ratings of other users on different contents. The matrix includes the user portrait data of multiple other users and the ratings of each other user on different contents.
[0043] Based on the similarity calculation between the user profile of the target user and the user profile data of other users in the rating matrix, similar users with similar user profiles are determined. Several users with high similarity to the target user can be specifically called "neighbor users". They can be sorted according to the similarity, and the top k users with the highest similarity ranking are selected as neighbors.
[0044] It should be noted that users with similar interests often have similar evaluations on the same content. In other words, if user A and user B show similar rating tendencies on multiple rated content, then when user A has not rated a certain content, but user B has rated it, the possible rating of user A on the content can be predicted based on the similarity between user A and user B and user B's rating of the content. In this way, the potential relationship between users can be mined to recommend content that the target users may be interested in but have not yet come into contact with.
[0045] When predicting the target user's rating of unknown content, the target user's historical content rating and the ratings of similar users on the content are comprehensively considered. Through certain algorithms, such as the weighted average method, different weights are assigned to the ratings of similar users according to the similarity between the target user and similar users and the target user's own historical ratings, and the final predicted rating is calculated.
[0046] In step 130, content is recommended to the target user based on the predicted score.
[0047] After calculating the predicted ratings of all unrated items, these items can be sorted from high to low according to the predicted ratings. The specific steps are as follows: Generate a predicted rating list: For the target user uT, generate a list L containing all unrated items and their predicted ratings.
[0048] L={(i1, r^uT_i1), (i2, r^uT_i2),…, (iN, r^uT_iN)}; Where i1, i2, ..., iN are unrated items, r^uT_i1, r^uT_i2, ..., r^uT_iN are the corresponding predicted ratings. Sorting is to sort the list L in descending order according to the predicted rating values.
[0049] Lsorted=sort(L, key=r^uT_i); Based on the sorted predicted ratings, the system generates a personalized recommendation list for the target user uT. The specific steps are as follows: Select top items: Select the top K items with the highest scores from the sorted list Lsorted as the recommendation list. Usually K is a predefined number of recommendations, such as K=10 or K=20.
[0050] r^uT ={i1,i2,…,iK}; Where i1, i2, …, iK are the K items with the highest predicted scores.
[0051] Filter rated items: If the target user uT has rated some items, these items usually do not appear in the recommendation list. Therefore, it is necessary to filter out the items that the user has rated from Lsorted.
[0052] Consider diversity: To improve the diversity of recommendations, the system may introduce some strategies when generating recommendation lists, such as: Category diversity: Make sure your recommendation list contains items from different categories.
[0053] Novelty: Prioritize new items that users have not come into contact with.
[0054] Based on business needs and user feedback, the system may dynamically adjust the recommendation strategy, for example: Popular content recommendations: Add some popular items to the recommendation list to increase the coverage of the recommendation.
[0055] New content recommendation: Prioritize newly released items to attract users to explore new content.
[0056] User feedback adjustment: Dynamically adjust the recommendation strategy based on user feedback on recommended content (such as click-through rate, viewing time, likes, etc.).
[0057] Optionally, for the content recommendation process, the specific implementation process of content recommendation can be assigned to the edge server closest to the data source for processing. Once the task is assigned, the relevant data will be unloaded from the central server or data source to the designated edge server. The data transmission process will adopt efficient transmission protocols and compression algorithms to reduce transmission time and bandwidth usage. The edge server has strong parallel processing capabilities and can handle multiple tasks at the same time. By utilizing technologies such as multi-core processors and GPU acceleration, the edge server can significantly improve the speed and efficiency of data processing. The edge server will perform real-time analysis on the received data, including data cleaning, format conversion, feature extraction, etc. Based on the analysis results, the edge server can make real-time decisions, such as whether the data needs to be further transmitted to the central server for in-depth analysis or processing. After the edge server completes the processing, it will transmit the processing results back to the central server or the designated receiving module. The returned data will be compressed and encrypted to ensure the security and integrity of the data. The central server will integrate and update the data received from the edge server. The integrated data will be used for subsequent analysis, recommendation or decision-making processes.
[0058] Continuously monitor the resource usage of edge servers, including CPU usage, memory usage, network bandwidth, etc. Dynamically schedule and allocate resources based on resource usage to ensure smooth execution of tasks and system stability. When the amount of tasks increases, automatically increase the number of edge servers or improve server performance to meet processing needs. Conversely, when the amount of tasks decreases, reduce the number of edge servers or reduce server performance accordingly to save resource costs. Regularly check the operating status and health of edge servers. Once a fault or abnormality is found, immediately trigger the alarm mechanism and notify relevant personnel to handle it. For detected faults or abnormalities, try to automatically recover or restart the edge server. If automatic recovery fails, pass the fault information to relevant personnel for further investigation and repair.
[0059] Optionally, after determining the recommended content, a copyright analysis can be performed on the recommended content, which can be specifically implemented based on the copyright detection and management module. Specifically, the copyright detection and management module receives data analysis results. These analysis results may contain information from multiple data sources, such as social media comments, blog posts, news reports, etc., which are associated with the audio and video content after comprehensive analysis. The received data is cleaned and formatted to ensure the consistency and accuracy of the data. Key information, such as text content, keywords, timestamps, etc., is extracted to prepare for subsequent analysis.
[0060] Using machine learning algorithms, extract copyright-related features from preprocessed data. These features may include audio fingerprints, video frame features, text similarity, etc. Match the extracted features with known copyright databases. Copyright databases may contain metadata, audio fingerprints, video frames, etc. of protected works. By comparing the similarity between features, determine whether there are protected copyright elements in the audio and video content.
[0061] For the matched copyright elements, calculate their similarity with the protected works in the database. Similarity calculation may involve a variety of algorithms and technologies, such as audio fingerprint comparison, video frame similarity calculation, text similarity algorithm, etc. Based on the similarity calculation results, determine whether the audio and video content constitutes infringement. A certain threshold may be set, and when the similarity exceeds the threshold, it is determined to be infringement. Once it is determined to be infringement, the copyright detection and management module sends an infringement notice to the copyright owner. The notice may contain detailed information on the infringing content, the time of infringement, the degree of infringement, etc.
[0062] The content recommendation method provided by the present invention integrates multi-source data such as the target user's own historical ratings, similarities with other users, and ratings of other users. These data complement each other and provide a basis for predicting ratings from different angles. Compared with relying on a single data source, it can more comprehensively understand user needs and content characteristics and improve the accuracy of prediction. By analyzing the rating data of a large number of users, it is possible to discover potential similarities between users and associations between content. Even if the target user has not rated certain content, the ratings of similar users can also be used to dig out the target user's potential interest in these contents, so that the user can find new content that may be of interest but has not yet been exposed to, thereby improving the diversity of recommendation results.
[0063] In one embodiment, based on the historical content score of the target user, the similarity between the target user and other users, and the scores of the other users for different contents, the predicted score of the target user for unknown content is determined, including: based on the similarity between the user portrait data of the target user and the user portrait data of the other users, determining a set of neighboring users similar to the target user; based on the similarity between users in the neighboring user set and the target user, the average score of the target user for historical content, and the scores of users in the neighboring user set for different contents, determining the predicted score of the target user for unknown content.
[0064] The similarity between the user portrait data of the target user and the user portrait data of the other users may be determined based on cosine similarity, by calculating the cosine value of the angle between the two user portrait vectors to measure the similarity.
[0065] After calculating the similarity between the target user and all other users, they are sorted according to the similarity, and several users with higher similarity are selected to form a set of neighboring users.
[0066] Optionally, the calculation process of user similarity can use methods such as Pearson correlation coefficient. The calculation formula of Pearson correlation coefficient is given below: ; Among them, sim(ua, ub) represents the similarity between user ua and user ub; ra_i and rb_i represent the ratings of user ua and ub on item i respectively; rä and rb represent the average ratings of user ua and ub on all items respectively; ∑ represents the sum of all items with common ratings. In order to improve the similarity calculation, the weight of item relevance can be introduced. Let rel(i, iT) represent the relevance between item i and item iT, then the improved user similarity calculation formula is: ; The target user's rating of historical content reflects his or her overall preference and evaluation criteria for different content. The target user's average rating is obtained by calculating the average rating of all historical content that the target user has rated.
[0067] For unknown content, the ratings of neighboring users are an important basis for predicting the ratings of target users. At the same time, the similarity between neighboring users and target users determines the weight of their ratings.
[0068] Specifically, the target user's predicted score for unknown content is: ; in, The target user For unknown content The prediction score of The target user Average rating for historical content, The target user With users The similarity between Is a user About content Rating, Is a user The average rating of different content. Is with the target user A collection of similar nearby users.
[0069] This formula takes into account the historical rating level of the target user, the ratings of neighboring users, and the similarity between users. It can more accurately predict the target user's rating for unknown content and provide strong support for subsequent content recommendations.
[0070] In one embodiment, based on the features extracted from the interaction data, user portrait data of the target user is constructed, including: performing text extraction on the interaction data to obtain text data; performing sentiment analysis on the text data to obtain sentiment analysis results, and performing context analysis on the text data to obtain context analysis results; performing feature extraction on the sentiment analysis results, the context analysis results, and the text data, and constructing target user portrait data of the target user based on the extracted features.
[0071] Automatic speech recognition (ASR) technology is used to convert the collected audio data into text format. ASR can adapt to speech in different accents, speech speeds, and language environments, accurately recognize words in the speech, and restore it into coherent text based on grammatical rules and semantic understanding. For example, for a voice chat record mixed with dialects, ASR technology can recognize dialect vocabulary through a pre-trained dialect model and output standard Mandarin text. For the voice part of the video, ASR technology is also used for conversion; for the subtitles in the video, if it is in the external subtitle format, the text content can be directly extracted. If it is an embedded subtitle, it is necessary to separate the subtitles from the video screen by combining image recognition and text extraction technology. For example, when analyzing the video data of a foreign language movie, ASR technology is first used to identify the voice text of the character's dialogue, and then combined with the extracted embedded subtitle text, multiple versions are compared and integrated to improve the accuracy and completeness of the text.
[0072] The user's browsing history, click behavior, search history, etc. are structured and described according to certain rules and converted into text form. For example, the web page titles, product names, search keywords, etc. that the user has browsed are linked together to form a text description similar to "The user browsed [web page title 1], [web page title 2], searched [keyword 1], [keyword 2], and clicked [product name 1], [product name 2]". This text processing allows the originally scattered behavior data to be presented in a unified format, which is convenient for subsequent analysis operations.
[0073] Optionally, the text extraction process can be implemented based on the real-time multilingual translation module. The module first receives audio and video content. Next, the module uses machine translation technology to translate the received audio and video content in real time. This includes converting the text or voice content in the source language into the text or voice form of the target language. During the translation process, factors such as grammatical differences, vocabulary selection, and expression between languages are considered to ensure the accuracy and fluency of the translation. In addition to the basic translation function, cultural adaptation processing is also performed. This includes considering the cultural expression habits and sensitivities of the target language area and making appropriate adjustments and optimizations to the translated content. For example, some words or expressions with specific cultural connotations may be encountered during the translation process. The module will replace or interpret them according to the cultural background of the target language area to ensure that the translated content can be accurately understood and accepted by local users. After translation and cultural adaptation processing, the final translation results are obtained. These results may include translated text, voice content, and related metadata (such as translation time, translation quality score, etc.). The module also monitors and evaluates the translation quality and performance. This includes checking indicators such as accuracy, fluency, and user satisfaction of the translation results. Based on the monitoring results, the module may make necessary adjustments and optimizations to improve translation quality and performance. At the same time, it will also provide relevant feedback information to system designers and developers so that they can continuously improve and optimize the entire system architecture. Including BLEU score calculation: BLEU (Bilingual Evaluation Understudy) is a commonly used machine translation quality evaluation indicator. Its calculation process includes the following steps: Word segmentation: Both the machine-generated translation and the manual reference translation are word segmented. Matching: For each n-gram of the machine translation (usually 1-gram to 4-gram), check whether it appears in the reference translation. Counting: For each matching n-gram, count and record the total number of all n-grams in the machine translation. Precision: Calculate the precision of each n-gram, that is, the number of matching n-grams divided by the total number of n-grams in the machine translation. Penalty factor: The BLEU score also takes into account the short sentence penalty (Brevity Penalty, BP). If the length of the machine translation is shorter than the reference translation, the score will be penalized. Comprehensive score: Take the geometric average of the accuracy of all n-grams and multiply it by the short sentence penalty factor to get the final BLEU score. The formula is as follows: ; in, is the precision of n-gram, N is the maximum length of n-gram, BP is the short sentence penalty factor, defined as , c is the length of the machine translation, and r is the length of the closest reference text.
[0074] ROUGE score calculation: The ROUGE indicator is mainly used to evaluate text summarization tasks, but it can also be used for machine translation quality assessment. Its calculation process includes: Matching n-grams: Calculate the number of times the n-grams contained in the translation generated by the system appear in the reference translation. Recall rate: Divide the total number of matched n-grams by the total number of n-grams in the reference translation. ROUGE score: Calculate ROUGE-1, ROUGE-2, and ROUGE-L scores according to different n-gram sizes. According to the monitoring results, if it is necessary to improve the translation quality and performance, the optimization algorithm in machine learning can be used to adjust the model. Commonly used optimization algorithms include: Stochastic Gradient Descent (SGD): Use a batch of samples in each iteration to approximate the loss function and optimize the objective function. Adam algorithm: Integrates the adaptive learning rate and momentum term, and constructs two vectors through the gradient to construct the adaptive learning rate. Newton method: Use the second-order derivative of the function (Hessian matrix) to find the optimal solution. Taking the Adam algorithm as an example, its parameter update formula is: ; in, and are the first-order and second-order moment estimates of the gradient, β1 and β2 are algorithm parameters, η is the learning rate, and ϵ is a small constant used to prevent division by zero.
[0075] The extracted text data is input into the selected sentiment analysis model, and the model uses the learned knowledge to label each sentence, paragraph, or even the entire text with sentiment. The sentiment analysis results are statistically analyzed to calculate the proportion of different sentiment categories in the text data, and to observe the emotional fluctuations of users in different scenarios, so as to gain a deeper understanding of the user's emotional state and satisfaction.
[0076] Use semantic understanding technology in natural language processing (NLP) to conduct in-depth analysis of text data. By building lexical semantic networks, syntactic analysis trees and other tools, we can understand the semantic relationship between words in the text, sentence structure and paragraph logic.
[0077] Combine the results of sentiment analysis and context analysis to extract sentiment features and context features. Sentiment features can be the distribution of users’ sentiment tendencies in different scenarios, such as a high proportion of positive sentiment in shopping scenarios and a high proportion of negative sentiment in work scenarios; context features can be typical contexts that users often appear in. The extracted text features, sentiment features, context features, etc. are integrated to form a multi-dimensional user portrait data.
[0078] In one embodiment, after obtaining the interaction data of the target user, it also includes: based on a noise detection algorithm, analyzing the background sound in the audio data segment by segment; based on the analysis result of the background sound, determining that the target user is not listening to the target audio segment of the audio data, and deleting the audio data corresponding to the target audio segment.
[0079] The collected audio data of the target user is segmented according to a certain time window. The segment length can be set according to the overall duration of the audio data and the analysis accuracy requirements. For each audio segment, the selected noise detection algorithm is applied for analysis. The algorithm calculates the relevant indicators of the audio segment in the frequency domain or feature space to determine whether there is background noise and the type and intensity of the noise. For example, the spectrum analysis algorithm finds that a certain audio segment has a high-intensity continuous energy signal in the low-frequency band. Combined with prior knowledge, it is judged that it may be the roar of nearby construction machinery, which is background noise.
[0080] It should be noted that when users concentrate on listening to audio content, the background sound is usually relatively stable and at a low intensity, and occasional short and slight noise generally does not affect the user's attention. However, if there is a long-term, high-intensity specific background sound, such as a continuous sharp alarm, a high-decibel crowd noise, etc., it is likely to interfere with the user's listening, causing the user to be distracted or even stop listening.
[0081] Based on the research on the characteristics of common background sound interference, a mapping model between background sound characteristics and user listening status is established. For example, when a noise lasting more than 3 seconds and with a decibel value exceeding 80dB is detected in a certain audio segment, the model determines that the user is likely not concentrating on listening to the audio content during this period.
[0082] According to the above mapping relationship model, each audio segment that has been analyzed through noise detection is judged to filter out the target audio segment that meets the characteristics that the user has not listened to. After determining the target audio segment, the deletion operation of the target audio segment is performed. After the deletion operation is completed, the remaining audio data is integrated to ensure the continuity of the audio data.
[0083] In one embodiment, the method further includes: acquiring real-time interaction data of the target user, performing feature extraction based on the real-time interaction data to obtain real-time interaction features of the target user; and updating user portrait data of the target user based on the real-time interaction features.
[0084] Real-time acquisition of target users' real-time interaction data during the interaction process. Real-time interaction data may include real-time audio data, real-time video data, and real-time user behavior data.
[0085] Feature extraction is performed on real-time interaction data to obtain real-time interaction features, and user portrait data is updated based on the obtained real-time interaction features. Based on the incremental learning strategy, the entire user portrait model can be avoided from being retrained for each update, saving computing resources and time costs. Only the model part involved in the real-time interaction features is locally optimized, while ensuring the accuracy of the portrait, fast and continuous updates are achieved, so that the user portrait is always close to the user's current real status.
[0086] The present invention also provides a schematic diagram of a device structure for applying the content recommendation method provided by the present invention, such as Figure 2 As shown, the device specifically includes: Data collection and preprocessing module, identify data sources (audio and video platforms, social media APIs, user devices, etc.). Establish connections with data sources and extract required data (audio and video files, user behavior logs, social media posts, etc.). Identify and delete invalid data (duplicate, missing, abnormal data), and use appropriate strategies to fill missing data. Convert data in different formats into a unified format and perform necessary encoding conversions. Perform dimensionality reduction on high-dimensional data and normalize data of different dimensions. Aggregate data that needs to be aggregated. Transfer preprocessed data to subsequent modules and store it in appropriate storage media; The noise detection and intelligent context understanding module receives audio and video data from the data acquisition and preprocessing module. It uses the noise detection algorithm to identify and distinguish noise and valid voice content. It detects noise more accurately in noisy environments and assists in determining whether the user is listening to the audio content. It uses deep learning algorithms to identify text and voice content in audio and video, including voice-to-text and text information extracted from videos. It analyzes the contextual meaning and emotional color of the identified text and voice content. It transmits the analysis results to the adaptive content recommendation algorithm module; Real-time multilingual translation and cultural adaptation module: Receives audio and video content processed by the noise detection and intelligent context understanding module. Uses machine translation technology to perform real-time translation, taking into account grammatical differences, vocabulary selection, and expressions between languages. Performs cultural adaptation processing, taking into account the cultural expression habits and sensitivities of the target language region. Outputs the results of translation and cultural adaptation processing to the content distribution and compatibility processing module; Through the user privacy protection module, the data is encrypted during the data collection process. The collected data is anonymized. Ensure that the data processing process does not leak user privacy. The processed user data is transmitted to the adaptive content recommendation algorithm module using a secure transmission protocol. Set a strict data access control policy to ensure that the recommendation algorithm module can only access the required data; Through the distributed edge computing module, identify the data processing tasks from the data acquisition and preprocessing module, the noise detection and intelligent context understanding module, the adaptive content recommendation algorithm module, and the multimodal fusion analysis module. Intelligently assign tasks to the edge server closest to the data source for processing. The edge server performs real-time data analysis and makes real-time decisions. The processing results are transmitted back to the central server or the designated receiving module. Continuously monitor the resource usage of the edge server, dynamically schedule and allocate resources. Regularly check the operating status of the edge server, trigger the alarm mechanism and handle faults; Through the adaptive content recommendation algorithm module, the user behavior data processed by the user privacy protection module and the content analysis results of the noise detection and intelligent context understanding module are received. The data is integrated to form a comprehensive user portrait and content understanding framework. The user's preferences, interests and consumption habits are analyzed to identify the current emotional state. The recommendation algorithm is used to generate a personalized audio and video content recommendation list. The recommendation strategy is formulated and adjusted, and the recommendation results are transmitted to the content distribution module. The recommendation algorithm and strategy are optimized according to user feedback and recommendation effect evaluation results; The multimodal fusion analysis module obtains multiple data sources. Cleans the data, performs format conversion and preliminary processing. Extracts key information and features, and converts them into representations suitable for subsequent analysis. Aligns data from different data sources to ensure consistency. Fusion of features from different data sources to form a comprehensive feature representation. Performs sentiment analysis, trend prediction, and association mining. Outputs the analysis results to the copyright detection and management module driven by machine learning.
[0087] Machine learning-driven copyright detection and management module: Receives data analysis results from the multimodal fusion analysis module. Cleans and formats the data to extract key information. Extracts copyright-related features and matches them with the copyright database. Determines whether the audio and video content contains protected copyright elements and calculates similarity. Determines whether infringement occurs based on similarity and sends infringement notices to the content distribution and compatibility processing module and the copyright owner. Through the content distribution and compatibility processing module: Receive processed audio and video content. Obtain user device information and determine content compatibility. Perform format conversion and other adaptation processing on incompatible content. Develop distribution strategies and transmit content to user devices. Monitor transmission speed and stability and collect user feedback. Process and analyze user feedback and optimize distribution strategies.
[0088] Through cloud computing resource management and dynamic expansion modules: Real-time monitoring of resource usage of all connected processing modules. Predict future data processing needs. Dynamically allocate computing resources to each processing module. Trigger dynamic expansion operations to increase or decrease resource allocation. Monitor resource usage efficiency and optimize resource allocation strategies. Dynamically adjust task allocation according to actual load conditions to achieve load balancing.
[0089] The content recommendation device provided by the present invention is described below. The content recommendation device described below and the content recommendation method described above can be referred to each other.
[0090] like Figure 3 As shown, the device comprises: A data acquisition module 310 is used to acquire interaction data of a target user and construct user portrait data of the target user based on features extracted from the interaction data, wherein the interaction data includes audio data, video data, and user behavior data; A rating prediction module 320, configured to determine a predicted rating of the target user for unknown content based on the historical content ratings of the target user, the similarity between the target user and other users, and the ratings of the other users for different content, wherein the similarity between the target user and the other users is determined based on the similarity between the user portrait data of the target user and the user portrait data of the other users; The content recommendation module 330 is used to recommend content to the target user based on the predicted score.
[0091] The content recommendation device provided by the present invention integrates multi-source data such as the target user's own historical ratings, similarities with other users, and ratings of other users. These data complement each other and provide a basis for predicting ratings from different angles. Compared with relying on a single data source, it can more comprehensively understand user needs and content characteristics and improve the accuracy of prediction. By analyzing the rating data of a large number of users, it is possible to discover potential similarities between users and associations between content. Even if the target user has not rated certain content, the ratings of similar users can also be used to dig out the target user's potential interest in these contents, so that the user can find new content that may be of interest but has not yet been exposed to, thereby improving the diversity of recommendation results.
[0092] In one embodiment, the rating prediction module 320 is specifically used to: Determining a predicted score of the target user for unknown content based on the historical content score of the target user, the similarity between the target user and other users, and the scores of the other users for different content, includes: Determining a set of neighboring users similar to the target user based on the similarity between the user portrait data of the target user and the user portrait data of the other users; Based on the similarity between the users in the neighboring user set and the target user, the average rating of the target user on historical content, and the ratings of the users in the neighboring user set on different content, the predicted rating of the target user on unknown content is determined.
[0093] In one embodiment, the rating prediction module 320 is further specifically configured to: Determine the target user's predicted score for unknown content: ; in, The target user For unknown content The prediction score of The target user Average rating for historical content, The target user With users The similarity between Is a user About content Rating, Is a user The average rating of different content. Is with the target user A collection of similar nearby users.
[0094] In one embodiment, the data acquisition module 310 is specifically used for: Constructing user portrait data of the target user based on the features extracted from the interaction data, including: Performing text extraction on the interaction data to obtain text data; Performing sentiment analysis on the text data to obtain sentiment analysis results, and performing context analysis on the text data to obtain context analysis results; Feature extraction is performed on the sentiment analysis result, the context analysis result, and the text data, and target user portrait data of the target user is constructed based on the extracted features.
[0095] In one embodiment, the data acquisition module 310 is further specifically used for: After obtaining the target user's interaction data, it also includes: Based on a noise detection algorithm, analyzing the background sound in the audio data segment by segment; Based on the analysis result of the background sound, it is determined that the target user is not listening to the target audio segment of the audio data, and the audio data corresponding to the target audio segment is deleted.
[0096] In one embodiment, the content recommendation module 330 is specifically used to: Acquire the real-time interaction data of the target user, perform feature extraction based on the real-time interaction data, and obtain the real-time interaction features of the target user; Based on the real-time interaction features, the user portrait data of the target user is updated.
[0097] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communications interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the content recommendation method, which includes: obtaining the interaction data of the target user, and constructing the user portrait data of the target user based on the features extracted from the interaction data, wherein the interaction data includes audio data, video data and user behavior data; Determining the predicted score of the target user for unknown content based on the historical content score of the target user, the similarity between the target user and other users, and the scores of the other users on different contents, wherein the similarity between the target user and the other users is determined based on the similarity between the user portrait data of the target user and the user portrait data of the other users; Based on the predicted score, content is recommended to the target user.
[0098] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0099] On the other hand, the present invention further provides a computer program product, the computer program product includes a computer program, the computer program can be stored in a non-transitory computer-readable storage medium, when the computer program is executed by a processor, the computer can execute the content recommendation method provided by the above methods, the method includes: obtaining interaction data of a target user, and constructing user portrait data of the target user based on features extracted from the interaction data, the interaction data including audio data, video data and user behavior data; Determining the predicted score of the target user for unknown content based on the historical content score of the target user, the similarity between the target user and other users, and the scores of the other users on different contents, wherein the similarity between the target user and the other users is determined based on the similarity between the user portrait data of the target user and the user portrait data of the other users; Based on the predicted score, content is recommended to the target user.
[0100] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the content recommendation method provided by the above methods, the method comprising: obtaining interaction data of a target user, and constructing user portrait data of the target user based on features extracted from the interaction data, wherein the interaction data comprises audio data, video data, and user behavior data; Determining the predicted score of the target user for unknown content based on the historical content score of the target user, the similarity between the target user and other users, and the scores of the other users on different contents, wherein the similarity between the target user and the other users is determined based on the similarity between the user portrait data of the target user and the user portrait data of the other users; Based on the predicted score, content is recommended to the target user.
[0101] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0102] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A content recommendation method, characterized in that: include: Acquire the target user's interaction data, and construct user portrait data of the target user based on the features extracted from the interaction data, wherein the interaction data includes audio data, video data, and user behavior data; Determining the predicted score of the target user for unknown content based on the historical content score of the target user, the similarity between the target user and other users, and the scores of the other users on different contents, wherein the similarity between the target user and the other users is determined based on the similarity between the user portrait data of the target user and the user portrait data of the other users; Based on the predicted score, content is recommended to the target user.
2. The content recommendation method according to claim 1, characterized in that: The step of determining the predicted score of the target user for unknown content based on the historical content score of the target user, the similarity between the target user and other users, and the scores of the other users for different contents comprises: Determining a set of neighboring users similar to the target user based on the similarity between the user portrait data of the target user and the user portrait data of the other users; Based on the similarity between the users in the neighboring user set and the target user, the average rating of the target user on historical content, and the ratings of the users in the neighboring user set on different content, the predicted rating of the target user on unknown content is determined.
3. The content recommendation method according to claim 2, characterized in that: The target user's predicted score for unknown content is: ; in, The target user For unknown content The prediction score of The target user Average rating for historical content, The target user With users The similarity between Is a user About content Rating, Is a user The average rating of different content. Is with the target user A collection of similar nearby users.
4. The content recommendation method according to claim 1, characterized in that: The constructing user portrait data of the target user based on the features extracted from the interaction data includes: Performing text extraction on the interaction data to obtain text data; Performing sentiment analysis on the text data to obtain sentiment analysis results, and performing context analysis on the text data to obtain context analysis results; Feature extraction is performed on the sentiment analysis result, the context analysis result, and the text data, and target user portrait data of the target user is constructed based on the extracted features.
5. The content recommendation method according to claim 1, characterized in that: After acquiring the interaction data of the target user, the method further includes: Based on a noise detection algorithm, analyzing the background sound in the audio data segment by segment; Based on the analysis result of the background sound, it is determined that the target user is not listening to the target audio segment of the audio data, and the audio data corresponding to the target audio segment is deleted.
6. The content recommendation method according to claim 1, characterized in that: Also includes: Acquire the real-time interaction data of the target user, perform feature extraction based on the real-time interaction data, and obtain the real-time interaction features of the target user; Based on the real-time interaction features, the user portrait data of the target user is updated.
7. A content recommendation device, characterized in that: include: A data acquisition module is used to acquire the interaction data of the target user and construct user portrait data of the target user based on the features extracted from the interaction data, wherein the interaction data includes audio data, video data and user behavior data; A rating prediction module, used to determine the target user's predicted rating for unknown content based on the target user's historical content rating, the target user's similarity with other users, and the other users' ratings of different content, wherein the target user's similarity with other users is determined based on the similarity between the target user's user portrait data and the other users' user portrait data; A content recommendation module is used to recommend content to the target user based on the predicted score.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the content recommendation method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the content recommendation method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the content recommendation method according to any one of claims 1 to 6 is implemented.