A software information processing method based on big data
By comprehensively considering multiple dimensions of user behavior and content characteristics, and dynamically adjusting the correlation between users and content, the problem of inaccurate user behavior characteristics and inaccurate content preference prediction in existing technologies is solved, thus realizing personalized and diversified recommendation services.
Patent Information
- Application Number
- CN202511299376.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-09-12
AI Technical Summary
Existing big data-based software information processing methods in the field of social network analysis do not fully consider the comprehensive interaction intensity and multi-dimensional factors of users with different types of social content, resulting in inaccurate and incomplete user behavior characteristics. Furthermore, when predicting user content preferences, they ignore key information such as content popularity, the correlation between users and content types, and the freshness of content, leading to inaccurate prediction results.
The data collection module gathers user behavior data, the data processing module cleans and integrates the data, and the information processing module comprehensively considers various behavioral characteristics such as user likes, comments, shares and activity levels, and combines them with factors such as content popularity, the relevance between users and content types and content freshness. Through user behavior sub-modules, content preference sub-modules, recommendation analysis sub-modules and relevance sub-modules, the module conducts in-depth analysis and prediction, and dynamically adjusts the relevance between users and content.
It achieves comprehensive and accurate mining of user behavior characteristics, improves the accuracy of content preference prediction and the personalization of recommendations, enhances the accuracy and efficiency of the recommendation system, can dynamically reflect changes in user interests, and provide personalized and diversified recommendation services.
Smart Images

Figure CN120804430B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of software information processing technology, specifically to a software information processing method based on big data. Background Technology
[0002] With the rapid development of Internet technology and the widespread popularity of social media, social networks have become an indispensable information exchange platform in people's daily lives. On the platform, user behavior data, such as likes, comments and shares, are constantly generated and accumulated. This data contains valuable information such as user preferences, interests, social relationships and consumption habits.
[0003] Big data-based software information processing technology has been widely used in many fields, especially in the field of social network analysis. It can help enterprises to understand user behavior more deeply, predict user preferences, and thus provide more personalized and accurate services.
[0004] However, current big data-based software information processing methods, when applied to the field of social network analysis, may not fully consider the comprehensive interaction intensity of users with different types of social content. The comprehensive interaction types can include likes, comments, and shares, as well as the impact of multiple dimensions such as user activity on user behavior characteristics. This results in the discovery of user behavior characteristics that are not accurate and comprehensive enough. Furthermore, when predicting user content preferences, it may only consider the direct relationship between users and content, while ignoring key information such as content popularity, the correlation between users and content types, and the freshness of content. This leads to inaccurate prediction results and makes it difficult to meet users' personalized needs. Summary of the Invention
[0005] The purpose of this invention is to provide a software information processing method based on big data, which solves the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a software information processing method based on big data, comprising the following information processing methods:
[0007] Step 1: Collect user behavior logs from social media, video websites, and news applications using the data collection module to obtain the number of likes, comments, shares, and user online time.
[0008] Step II: The data processing module performs data cleaning, integration, and normalization on the number of likes, comments, shares, and online login time of users. The data processing module outputs the number of likes, comments, and shares of user i on content type j, and user i's activity level.
[0009] Step III: Input the number of likes, comments, and shares of user i on content type j into the information processing module. The information processing module outputs the behavioral feature value of user i on content type j, the extended preference prediction value of user i on content type j, and the recommendation score of user i on content type j.
[0010] Step IV: The analysis module analyzes the behavioral characteristic value of user i for content type j, the predicted value of user i's extended preference for content type j, and the recommendation score of user i for content type j to gain a deeper understanding of user needs and preferences and recommend content that best matches their preferences.
[0011] Optionally, the information processing module includes: a user behavior submodule, a content preference submodule, a recommendation analysis submodule, and a relevance submodule.
[0012] Optionally, the calculation formula for the user behavior submodule is as follows:
[0013] UPSA j,i = (LKA) j,i ×SA+CON j,i ×SB+SAE j,i ×SC) 2 +ACCTY i -QLT j -0.5 ;
[0014] in:
[0015] UPSA j,i This represents the behavioral characteristic value of user i for content type j;
[0016] j represents the content type, which refers to different types of content in the field of social network analysis. Content types include text, images, video, audio, and LKA. j,i This represents the number of times user i liked content of type j, CON j,i Indicates the number of comments user i made on content type j, SAE j,i表示 Number of times user i shares content type j, ACCTY i Indicates user i's activity level, QLT j SA represents the quality of content type j, SB represents the weighting factor for likes, SC represents the weighting factor for comments, and SC represents the weighting factor for shares.
[0017] The processing procedure of the user behavior submodule is as follows: The number of likes (LKA) given by user i to content type j is calculated. j,i CON is the number of comments made by user i on content type j. j,i Number of times user i shares content type j (SAE)j,i And user i's activity ACCTY i Input is sent to the user behavior submodule, and the output is the behavioral feature value UPSA of user i for content type j. j,i .
[0018] Optionally, the calculation formula for the content preference submodule is as follows:
[0019] CPAQ j,i = (UPSA) j,i ×PS j ) + (AFNY j,i ×FRHE 0.5 ) - (DTET) j,i / PLAT i );
[0020] in:
[0021] CPAQ j,i This represents the predicted value of user i's extended preference for content type j;
[0022] PS j Indicates the popularity of content type j, AFNY j,i The relevance between user i and content type j is represented by FRHE, which represents the freshness of content type j, and DTET is represented by DTET. j,i This indicates the degree of disinterest of user i in content type j, PLAT i Indicates the popularity of user i, UPSA j,i ×PS j This indicates the weighted effect of user behavior characteristics on popular content.
[0023] The processing procedure of the content preference submodule is as follows: The behavioral feature value UPSA of user i for content type j is... j,i Input is fed into the content preference submodule and based on the popularity PS of content type j. j The correlation between user i and content type j (AFNY) j,i Output the extended preference prediction (CPAQ) for user i on content type j. j,i .
[0024] Optionally, the calculation formula for the recommendation analysis submodule is as follows:
[0025] CRSQ j,i = (CPAQ) j,i ×RCN j ) + (UPSA j,i ×TA) + (MTC) 0.5 ×PPS);
[0026] in:
[0027] CRSQ j,i This represents the recommendation score for user i for content type j;
[0028] RCN j This indicates the timeliness of content type j; TA represents UPSA. j,i The weighting factor, MTC represents the matching degree between user i and content type j, and PPS represents the weighting factor of MTC;
[0029] The processing procedure of the recommendation analysis submodule is as follows: The predicted extended preference value (CPAQ) of user i for content type j is calculated. j,i And the behavioral feature value UPSA of user i for content type j j,i Input is fed into the recommendation analysis submodule and based on the timeliness RCN of content type j. j Output the CRSQ recommendation score for user i on content type j. j,i .
[0030] Optionally, the processing procedure of the correlation degree submodule is as follows:
[0031] Step 1: AFNY j,i,new =AFNY j,i,old +α×(CRSQ) j,i -TSD);
[0032] Step 2: Establish the loop termination condition;
[0033] Condition 1: |AFNY j,i,new -AFNY j,i,old | < 0.00493;
[0034] Condition 2: The number of iterations is 123;
[0035] in:
[0036] AFNY j,i,new This represents the degree of association between user i and content type j after iteration;
[0037] AFNY j,i,old The expression represents the correlation between user i and content type j before iteration; α represents the learning rate, used to control the update step size; and TSD represents the comparison threshold, used to determine the CRSQ. j,i Whether the expected level has been achieved;
[0038] CRSQ based on user i's recommendation score for content type j j,i The difference between the comparison threshold TSD and the value of AFNY is calculated, and the correlation between user i and content type j before the iteration is controlled by the learning rate α. j,i,old The degree of change, outputting the correlation between user i and content type j after iteration. j,i,newThe correlation between user i and content type j after iteration is used. j,i,new Replace the correlation between user i and content type j in the content preference submodule with AFNY. j,i To continuously optimize CPAQ j,i and CRSQ j,i And by using conditions one and two, the loop terminates and converges.
[0039] Optionally, the data collection module can obtain user behavior data by crawling web technology, by using log analysis software, or by calling tools through API interfaces.
[0040] Optionally, the data cleaning in the data processing module involves removing duplicate data and correcting erroneous data; data integration involves integrating data from different sources and standardizing the data to ensure consistency in data format; and normalization involves converting the data into a format suitable for processing and then normalizing the data to ensure that the data is on the same scale.
[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0042] I. This invention outputs the behavioral feature value UPSA of user i on content type j through the user behavior submodule. j,i The behavior submodule comprehensively considers various user behavioral characteristics, such as number of views, number of clicks, dwell time, and number of shares, thus comprehensively reflecting user interests and preferences. Multi-dimensional analysis makes user profiles more three-dimensional and accurate. UPSA j,i The calculations provide the foundational data for subsequent personalized recommendations. By deeply analyzing user behavioral characteristics, this submodule can identify users' unique interests and needs, providing strong support for personalized recommendations. j,i It not only considers the user's basic attributes but also deeply analyzes the user's behavioral characteristics to make the user profile more comprehensive and accurate. This sub-module can also dynamically reflect changes in user interests, improving the timeliness and accuracy of recommendations.
[0043] II. This invention outputs the extended preference prediction value (CPAQ) of user i for content type j through the content preference submodule. j,i This submodule comprehensively considers multiple factors, including user behavior characteristics, content popularity, and the correlation between users and content types, to accurately predict user preferences for specific content. This submodule can recommend content that better matches users' interests and needs while improving the diversity of recommendations. (CPAQ) j,i Computation can provide the basis for optimization strategies in recommender systems. By analyzing user preference prediction results, recommender systems can adjust their recommendation strategies to improve the accuracy and efficiency of recommendations. CPAQ j,iThe calculation takes into account more factors, including user behavior characteristics and content popularity, making the recommendation results more comprehensive and accurate. In addition, this sub-module can also delve into the potential interests behind user behavior, improving the diversity and innovation of recommendations.
[0044] Third, this invention outputs the recommendation score (CRSQ) of user i for content type j through the recommendation analysis submodule. j,i CPAQ j,i Based on user preference predictions, this module can further calculate a recommendation score for specific content. This score intuitively reflects the user's level of interest in the content, providing a precise scoring basis for the recommendation system. This submodule provides a more accurate scoring basis for the recommendation system, making the recommendation results more in line with user expectations and needs. (CPAQ) j,i The calculation provides a basis for the optimization and improvement of the recommendation system. By analyzing the distribution of recommendation scores, the recommendation system can discover potential problems and directions for improvement, thereby improving the accuracy of recommendations. This submodule considers more factors, including user behavior characteristics and content preference prediction, making the scoring results more comprehensive and accurate. In addition, this submodule can also comprehensively consider the relationship between users and content in multiple dimensions, improving the objectivity and fairness of the scoring.
[0045] IV. This invention uses a correlation submodule to iterate the correlation between user i and content type j. j,i This iterative approach dynamically adjusts the correlation between users and content types based on changes in user behavior and content. This dynamism allows the recommendation system to more accurately reflect users' current interests and preferences, thus achieving more personalized recommendations. Through continuous iteration and optimization, AFNY... j,i Recommendation systems can gradually reduce the discrepancy between recommendations and users' actual interests, thereby improving the accuracy of recommendations. The iterative process reflects the complexity and dynamism of software information when processing and analyzing large-scale data. Through continuous iteration and optimization of model parameters, the software can more accurately capture and reflect the information and patterns in the data, thus providing users with better services. Attached Figure Description
[0046] Figure 1 This is a flowchart illustrating the steps of a software information processing method based on big data.
[0047] Figure 2 This is a schematic diagram of the overall structure of the software information processing method based on big data.
[0048] Figure 3 This is a schematic diagram of the information processing module in this big data-based software information processing method.
[0049] Figure 4 This is a flowchart illustrating the steps of the information processing module in a software information processing method based on big data. Detailed Implementation
[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0051] This big data-based software information processing method differs from existing methods. Existing methods in the field of social network analysis often only collect, store, and perform preliminary analysis of user behavior data, lacking in-depth mining and comprehensive understanding of user behavior characteristics. At the same time, when predicting user content preferences, existing technologies have failed to fully consider multi-dimensional factors such as content popularity, the correlation between users and content types, and the freshness of content, resulting in inaccurate prediction results.
[0052] This software information processing method based on big data comprehensively considers multiple dimensions of user interaction data, including interaction frequency, interaction type weight, user activity, and content quality. This invention can more accurately mine user behavior characteristics, providing a solid foundation for subsequent content preference prediction and recommendation score calculation. By integrating key information such as content popularity, the correlation between users and content types, and content freshness, it can achieve accurate prediction of user content preferences. When calculating content recommendation scores, this invention not only considers conventional factors such as the timeliness of the content and the matching degree between users and content types.
[0053] Example 1: Please refer to Figures 1 to 4 This implementation provides a software information processing method based on big data, including the following information processing methods:
[0054] Step 1: Collect user behavior logs from social media, video websites, and news applications using the data collection module to obtain the number of likes, comments, shares, and user online time.
[0055] Step II: The data processing module performs data cleaning, integration, and normalization on the number of likes, comments, shares, and online login time of users. The data processing module outputs the number of likes, comments, and shares of user i on content type j, and user i's activity level.
[0056] Step III: Input the number of likes, comments, and shares of user i on content type j into the information processing module. The information processing module outputs the behavioral feature value of user i on content type j, the extended preference prediction value of user i on content type j, and the recommendation score of user i on content type j.
[0057] Step IV: Analyze the behavioral feature values of user i for content type j, the predicted value of user i's extended preference for content type j, and the recommendation score of user i for content type j through the analysis module, in order to gain a deeper understanding of user needs and preferences and recommend content that best matches their preferences.
[0058] The information processing module includes: user behavior submodule, content preference submodule, recommendation analysis submodule, and relevance submodule.
[0059] In this embodiment, the interconnectedness and synergy of multiple sub-modules within the information processing module constitute the core of the software information processing method based on big data. Their collaborative effect enables the recommendation system to more accurately understand user needs, improving the accuracy and efficiency of recommendations. This method provides personalized recommendation services, which not only enhances user satisfaction and loyalty but also strengthens user trust and reliance on the recommendation system. The combined use of multiple sub-modules also provides a basis for the optimization and improvement of the recommendation system. By analyzing the calculation results and recommendation effects of the sub-modules, the recommendation system can identify potential problems and improvement directions, thereby continuously improving the accuracy and efficiency of recommendations. In the era of big data, recommendation systems need to process massive amounts of user and content data. The combined use of multiple sets of sub-modules fully reflects the demand for big data processing capabilities. They can effectively utilize this data to extract valuable information for recommendation services. The combined use of multiple sets of sub-modules also reflects the technical means of software information processing. Through algorithms and models, in-depth data mining and analysis are achieved, improving the accuracy and efficiency of recommendations. This technical means enables the recommendation system to respond to user needs more quickly and provide higher-quality recommendation services.
[0060] Please see Figures 1 to 4 The user behavior submodule processing procedure is as follows:
[0061] UPSA j,i = (LKA) j,i ×SA+CON j,i ×SB+SAE j,i ×SC) 2 +ACCTY i -QLT j -0.5 ;
[0062] in:
[0063] UPSA j,i This represents the behavioral characteristic value of user i for content type j;
[0064] j represents the content type, which refers to different types of content in the field of social network analysis. Content types include text, images, video, audio, and LKA. j,i This represents the number of times user i liked content of type j, CON j,i Indicates the number of comments user i made on content type j, SAE j,i表示 The number of times user i shares content type j;
[0065] ACCTY i This indicates the activity level of user i. Activity level is calculated based on metrics such as the total number of interactions, online time, and access frequency on the platform. Users with high activity levels usually have a higher interest in and engagement with the platform's content.
[0066] QLT j This represents the quality of content type j. Content quality is evaluated in various ways, based on user interaction, such as likes, comments, and shares, as well as factors like the professionalism, originality, and attractiveness of the content. High-quality content usually attracts more user attention and interaction.
[0067] SA represents the weighting factor for likes, SB represents the weighting factor for comments, and SC represents the weighting factor for shares.
[0068] The user behavior submodule processes the following: It calculates the number of likes user i has for content type j using the LKA method. j,i CON is the number of comments made by user i on content type j. j,i Number of times user i shares content type j (SAE) j,i And user i's activity ACCTY i Input is sent to the user behavior submodule, and the output is the behavioral feature value UPSA of user i for content type j. j,i .
[0069] In this embodiment: This submodule can comprehensively consider various user behavioral characteristics, such as browsing frequency, click frequency, dwell time, sharing frequency, and comment content, thereby more comprehensively reflecting the user's interests and preferences. This multi-dimensional analysis makes the user profile more three-dimensional and accurate. Furthermore, this submodule can dynamically update as user behavior changes, capturing subtle changes in user interests. This dynamism allows the recommendation system to adjust its recommendation strategy in a timely manner, improving the timeliness and accuracy of recommendations. j,iThe calculations provide the foundational data for subsequent personalized recommendations. By deeply analyzing user behavior characteristics, this sub-module can identify users' unique interests and needs, providing strong support for personalized recommendations. Traditional user profiling technologies based on big data software information processing often only focus on basic user attributes, such as age, gender, and occupation, while ignoring user behavior characteristics. This limitation makes user profiles insufficiently comprehensive and accurate, making it difficult to meet the needs of personalized recommendations. Compared with existing technologies, UPSA... j,i Not only does it consider basic user attributes, but it also delves into user behavioral characteristics, making user profiles more comprehensive and accurate. Furthermore, this submodule can dynamically reflect changes in user interests, improving the timeliness and accuracy of recommendations. In the era of big data, user behavior data is massive and complex; UBFEij can effectively process and utilize this data, extracting valuable information and providing strong support for software information processing. j,i It is a crucial part of software information processing, transforming user behavior data into analyzable metrics to provide a basis for subsequent content preference prediction and recommendation. This transformation process enables the recommendation system to more accurately understand user needs, improving recommendation satisfaction and conversion rates. The subscripts i and j represent the user and content type, respectively. This representation allows this submodule to clearly express the behavioral characteristic relationship between users and specific content types. The parameter subscripts i and j help to accurately locate specific users and content types. This location method improves the accuracy and efficiency of processing, enabling the recommendation system to respond to user needs more quickly.
[0070] Please see Figures 1 to 4 The content preference submodule processes the content as follows:
[0071] CPAQ j,i = (UPSA) j,i ×PS j ) + (AFNY j,i ×FRHE 0.5 ) - (DTET) j,i / PLAT i );
[0072] in:
[0073] CPAQ j,i This represents the predicted value of user i's extended preference for content type j;
[0074] PS j PS represents the popularity of content type j. j This parameter is calculated based on all users' interaction data with the content and reflects the content's general popularity. It is obtained through direct data collection, acquiring relevant data from content management systems or data analysis platforms via API interfaces.
[0075] AFNY j,i This represents the degree of association between user i and content type j. This value is calculated based on the similarity between the user's historical behavior and the content topic, reflecting the user's degree of interest in a specific content type. It is calculated through a comprehensive analysis of the user's historical behavior and content characteristics.
[0076] FRHE represents the freshness of content type j. This value is calculated based on the difference between the content's publication time and the current time, reflecting the content's timeliness and novelty. It can be obtained by querying the timestamp information in the content management system.
[0077] DTET j,i This value represents the degree of disinterest of user i in content type j. It is calculated based on the estimated value of the user's historical behavior and reflects the degree of user aversion to a specific content type. It is calculated through a comprehensive analysis of the user's historical behavior.
[0078] PLAT i This indicates the popularity of user i. This value is calculated based on the total number of interactions or followers of the user on the social network, reflecting the user's activity and influence on the social network.
[0079] UPSA j,i ×PS j This indicates the weighted effect of user behavior characteristics on popular content.
[0080] The content preference submodule processes the following: It assigns the behavioral feature value UPSA of user i to content type j. j,i Input is fed into the content preference submodule and based on the popularity PS of content type j. j The correlation between user i and content type j (AFNY) j,i Output the extended preference prediction (CPAQ) for user i on content type j. j,i .
[0081] In this embodiment: This submodule can comprehensively consider multiple factors such as user behavior characteristics, content popularity, and the correlation between users and content types, thereby more accurately predicting users' preferences for specific content. This submodule can recommend content that better matches users' interests and needs, while improving the diversity of recommendations. This diversity enables the recommendation system to provide users with richer choices, improving user satisfaction and loyalty. CPAQ j,iThe computational power of CPAQ provides a basis for optimizing recommendation strategies. By analyzing the results of user preference predictions, the recommendation system can adjust its strategies to improve accuracy and efficiency. Traditional computational processing methods often focus only on the direct relationship between users and content, ignoring the underlying interests behind user behavior; while content-based recommendation algorithms focus only on the features of the content itself, ignoring user interaction with the content. Both algorithms have limitations and struggle to meet the needs of personalized recommendations. Compared with existing technologies, CPAQ... j,i The calculations consider more factors, including user behavior characteristics and content popularity, making the recommendation results more comprehensive and accurate. Furthermore, this submodule can delve deeper into the potential interests behind user behavior, improving the diversity and innovation of recommendations. In the era of big data, content data is equally massive and complex. CPAQ... j,i CPAQ's calculations can effectively process and utilize this data to recommend valuable content to users. j,i It is one of the key steps in software information processing. It combines user and content data to achieve personalized content recommendations. This combination allows the recommendation system to understand user needs more accurately, improve recommendation satisfaction and conversion rate. The subscripts i and j also represent user and content type. This representation allows submodules to clearly express the preference relationship between users and specific content types.
[0082] Please see Figures 1 to 4 The recommended analysis submodule processing procedure is as follows:
[0083] CRSQ j,i = (CPAQ) j,i ×RCN j ) + (UPSA j,i ×TA) + (MTC) 0.5 ×PPS);
[0084] in:
[0085] CRSQ j,i This represents the recommendation score for user i on content type j, used to evaluate the recommendation value of content to the user;
[0086] RCN j This indicates the timeliness of content type j. This value is calculated based on the difference between the content's publication time and the current time, reflecting the latest status of the content.
[0087] TA stands for UPSA j,i Weighting factors;
[0088] MTC represents the matching degree between user i and content type j. This value is calculated based on the similarity between user interests and content topics, reflecting the user's adaptability and interest in the content.
[0089] PPS represents the weighting factor of MTC;
[0090] The processing procedure of the recommendation analysis submodule is as follows: The predicted extended preference value (CPAQ) of user i for content type j is calculated. j,i And the behavioral feature value UPSA of user i for content type j j,i Input is fed into the recommendation analysis submodule and based on the timeliness RCN of content type j. j Output the CRSQ recommendation score for user i on content type j. j,i .
[0091] In this embodiment: CPAQ j,i Based on user preference predictions, this module can further calculate a recommendation score for specific content. This score directly reflects the user's level of interest in the content, providing a more accurate scoring basis for the recommendation system. This submodule enables the recommendation system to provide more precise scoring criteria, making the recommendation results more aligned with user expectations and needs. This optimization improves recommendation satisfaction and conversion rates, enhances user trust and reliance on the recommendation system, and CPAQ... j,i The calculations can provide a basis for the optimization and improvement of recommendation systems. By analyzing the distribution of recommendation scores, recommendation systems can discover potential problems and directions for improvement, thereby improving the accuracy and efficiency of recommendations. However, traditional rating systems often only focus on the direct rating of content, ignoring the multi-dimensional relationship between users and content. This limitation makes the rating results incomplete and inaccurate, making it difficult to meet the needs of personalized recommendations. Compared with existing technologies, this submodule considers more factors, including user behavior characteristics and content preference prediction, making the rating results more comprehensive and accurate. In addition, this submodule can also comprehensively consider the multi-dimensional relationship between users and content, improving the objectivity and fairness of the rating. In the era of big data, recommendation systems need to process large amounts of user and content data. CPAQ j,i The calculations can effectively utilize this data to provide users with personalized recommendation services, as well as the calculated CPAQ. j,i As a crucial component of software information processing, it transforms user preference predictions into specific recommendation scores, providing a basis for optimizing and improving the recommendation system. This transformation process enables the recommendation system to more accurately understand user needs, improving the accuracy and efficiency of recommendations. The subscripts i and j used in this method emphasize the relationship between users and content types. This representation allows this submodule to clearly express the user's recommendation score for specific content. In subsequent analysis and calculation, the parameter subscripts i and j help to accurately assess the user's level of interest in specific content. This assessment method improves the accuracy and efficiency of recommendations, enabling the recommendation system to respond to user needs more quickly.
[0092] Please see Figures 1 to 4 The correlation degree submodule processing procedure is as follows:
[0093] Step 1: AFNY j,i,new =AFNY j,i,old +α×(CRSQ) j,i -TSD);
[0094] Step 2: Establish the loop termination condition;
[0095] Condition 1: |AFNY j,i,new -AFNY j,i,old | < 0.00493;
[0096] Condition 2: The number of iterations is 123;
[0097] in:
[0098] AFNY j,i,new This represents the degree of association between user i and content type j after iteration;
[0099] AFNY j,i,old The expression represents the correlation between user i and content type j before iteration; α represents the learning rate, used to control the update step size; and TSD represents the comparison threshold, used to determine the CRSQ. j,i Whether the expected level has been achieved;
[0100] CRSQ based on user i's recommendation score for content type j j,i The difference between the comparison threshold TSD and the value of AFNY is calculated, and the correlation between user i and content type j before the iteration is controlled by the learning rate α. j,i,old The degree of change, outputting the correlation between user i and content type j after iteration. j,i,new The correlation between user i and content type j after iteration is used. j,i,new Replace the correlation between user i and content type j in the content preference submodule with AFNY. j,i To continuously optimize CPAQ j,i and CRSQ j,i And by using conditions one and two, the loop terminates and converges.
[0101] In this embodiment: This iterative approach can dynamically adjust the correlation between users and content types based on changes in user behavior and content. This dynamism allows the recommendation system to more accurately reflect users' current interests and preferences, thereby achieving more personalized recommendations. Through continuous iteration and optimization, AFNY... j,iRecommendation systems can gradually reduce the discrepancy between recommendations and users' actual interests, thereby improving recommendation accuracy. This increased accuracy helps enhance user experience and satisfaction. Furthermore, in a big data environment, this iterative approach requires processing and analyzing massive amounts of user behavior and content feature data. This demands stronger data processing capabilities from the software information processing system, including data cleaning, preprocessing, storage, and analysis. The iterative approach reflects the complexity and dynamism of software information processing and analyzing large-scale data. By continuously iterating and optimizing model parameters, the software can more accurately capture and reflect information and patterns in the data, thus providing users with better services. Compared to traditional recommendation systems based on big data software information processing, this iterative approach places greater emphasis on user behavior... To accommodate the dynamism and diversity of content features, it no longer relies solely on static user profiles and content characteristics. Instead, it continuously learns and iterates to optimize recommendation performance. This iterative learning approach makes the recommendation system more intelligent and flexible. In general, this iterative approach enables users to more easily find content they are interested in through more accurate recommendations, thereby improving user experience and satisfaction. Accurate recommendations also help distribute content more efficiently to interested users, increasing content exposure and click-through rates. A better recommendation system can help content creators attract more attention, thus promoting their creation and development. Through intelligent recommendations, unnecessary resource waste can be reduced, such as avoiding recommending content to uninterested users, thereby optimizing resource utilization efficiency.
[0102] In the specific implementation process, the system architecture is constructed using various sub-modules from this method, and the number of likes given by user i to content type j is LKA. j,i CON is the number of comments made by user i on content type j. j,i Number of times user i shares content type j (SAE) j,i And user i's activity ACCTY i Input is sent to the user behavior submodule, and the output is the behavioral feature value UPSA of user i for content type j. j,i This submodule comprehensively considers various user behavioral characteristics, such as number of views, number of clicks, dwell time, number of shares, and comment content, thereby more comprehensively reflecting user interests and preferences. This multi-dimensional analysis makes user profiles more three-dimensional and accurate. j,i The calculations provide the foundational data for subsequent personalized recommendations. By deeply analyzing user behavioral characteristics, this submodule can identify users' unique interests and needs, providing strong support for personalized recommendations. j,i It not only considers the user's basic attributes but also analyzes the user's behavioral characteristics in depth, making the user profile more comprehensive and accurate. This sub-module can also dynamically reflect changes in user interests, improving the timeliness and accuracy of recommendations.
[0103] UPSA is the behavioral feature value of user i to content type j. j,i Input is fed into the content preference submodule and based on the popularity PS of content type j. j The correlation between user i and content type j (AFNY) j,i Output the extended preference prediction (CPAQ) for user i on content type j. j,i This submodule comprehensively considers multiple factors, including user behavior characteristics, content popularity, and the correlation between users and content types, to more accurately predict user preferences for specific content. It recommends content that better matches users' interests and needs, while also increasing the diversity of recommendations. (CPAQ) j,i Computation can provide a basis for optimization strategies in recommender systems. By analyzing user preference prediction results, recommender systems can adjust their recommendation strategies to improve the accuracy and efficiency of recommendations. (CPAQ) j,i The calculation takes into account more factors, including user behavior characteristics and content popularity, making the recommendation results more comprehensive and accurate. In addition, this submodule can also delve into the potential interests behind user behavior, improving the diversity and innovation of recommendations.
[0104] The CPAQ prediction value of user i for content type j j,i And the behavioral feature value UPSA of user i for content type j j,i Input is fed into the recommendation analysis submodule and based on the timeliness RCN of content type j. j Output the CRSQ recommendation score for user i on content type j. j,i CPAQ j,i Based on user preference predictions, this module can further calculate a recommendation score for specific content. This score intuitively reflects the user's level of interest in the content, providing a precise scoring basis for the recommendation system. This submodule can provide the recommendation system with more accurate scoring criteria, making the recommendation results more in line with user expectations and needs. (CPAQ) j,i The calculation can provide a basis for the optimization and improvement of the recommendation system. By analyzing the distribution of recommendation scores, the recommendation system can discover potential problems and improvement directions, thereby improving the accuracy of recommendations. This submodule considers more factors, including user behavior characteristics and content preference prediction, making the scoring results more comprehensive and accurate. In addition, this submodule can also comprehensively consider the relationship between users and content in multiple dimensions, improving the objectivity and fairness of the scoring.
[0105] Through CRSQ j,i The difference between the comparison threshold TSD is used to iterate the correlation between user i and content type j (AFNY). j,iThis iterative approach dynamically adjusts the correlation between users and content types based on changes in user behavior and content. This dynamism allows the recommendation system to more accurately reflect users' current interests and preferences, thereby achieving more personalized recommendations. Through continuous iteration and optimization, AFNY... j,i Recommendation systems can gradually reduce the discrepancy between recommendations and users' actual interests, thereby improving the accuracy of recommendations. The iterative process reflects the complexity and dynamism of software information when processing and analyzing large-scale data. Through continuous iteration and optimization of model parameters, the software can more accurately capture and reflect the information and patterns in the data, thus providing users with better services.
[0106] This allows the various sub-modules to cooperate and calculate in pairs, and also enables overall looping and iteration, giving the overall system an automated optimization and update effect, thus improving its adaptability.
[0107] Example 2: Please refer to Figure 1 , Figure 2 and Figure 3 The data collection module uses web crawler technology to capture user behavior data, log analysis software to obtain user data, and API interface to call tools to obtain user behavior data. The data processing module includes data cleaning to remove duplicate data and correct erroneous data, data integration to integrate data from different sources and standardize the data to ensure data format consistency, and normalization to convert the data into a format suitable for processing and then normalize the data to ensure that the data is on the same scale.
[0108] In this embodiment, the web crawler software is BeautifulSoup, and the API interface calling tool is Postman. Through web crawling technology and API interface calling tool, user behavior and related data can be collected, so as to realize the concept of software information processing based on big data. The collected data is further processed by the data processing module so that the data can be successfully input into subsequent modules.
[0109] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A software information processing method based on big data, characterized in that: Including the following information processing methods: Step 1: Collect user behavior logs from social media, video websites, and news applications using the data collection module to obtain the number of likes, comments, shares, and user online time. Step II: The data processing module performs data cleaning, integration, and normalization on the number of likes, comments, shares, and online login time of users. The data processing module outputs the number of likes, comments, and shares of user i on content type j, and user i's activity level. Step III: Input the number of likes, comments, and shares of user i on content type j into the information processing module. The information processing module outputs the behavioral feature value of user i on content type j, the extended preference prediction value of user i on content type j, and the recommendation score of user i on content type j. Step IV: Analyze the behavioral feature values of user i for content type j, the predicted value of user i's extended preference for content type j, and the recommendation score of user i for content type j through the analysis module, in order to gain a deeper understanding of user needs and preferences and recommend content that best matches their preferences. The information processing module includes: a user behavior submodule, a content preference submodule, a recommendation analysis submodule, and a relevance submodule; The calculation formula for the user behavior submodule is as follows: UPSA j,i =(LKA j,i ×SA+CON j,i ×SB+SAE j,i ×SC) 2 +ACCTY i -QLT j -0.5 ; in: UPSA j,i This represents the behavioral characteristic value of user i for content type j; j represents the content type, which refers to different types of content in the field of social network analysis. Content types include text, images, video, audio, and LKA. j,i This represents the number of times user i liked content of type j, CON j,i Indicates the number of comments user i made on content type j, SAE j,i表示 Number of times user i shares content type j, ACCTY i Indicates user i's activity level, QLT j SA represents the quality of content type j, SB represents the weighting factor for likes, SC represents the weighting factor for comments, and SC represents the weighting factor for shares. The processing procedure of the user behavior submodule is as follows: The number of likes (LKA) given by user i to content type j is calculated. j,i CON is the number of comments made by user i on content type j. j,i Number of times user i shares content type j (SAE) j,i And user i's activity ACCTY i Input is sent to the user behavior submodule, and the output is the behavioral feature value UPSA of user i for content type j. j,i ; The calculation formula for the content preference submodule is as follows: CPAQ j,i =(UPSA j,i ×PS j )+(AFNY j,i ×FRHE 0.5 )-(DTET j,i / PLAT i ); in: CPAQ j,i This represents the predicted value of user i's extended preference for content type j; PS j Indicates the popularity of content type j, AFNY j,i The relevance between user i and content type j is represented by FRHE, which represents the freshness of content type j, and DTET is represented by DTET. j,i This indicates the degree of disinterest of user i in content type j, PLAT i Indicates the popularity of user i, UPSA j,i ×PS j This indicates the weighted effect of user behavior characteristics on popular content. The processing procedure of the content preference submodule is as follows: The behavioral feature value UPSA of user i for content type j is... j,i Input is fed into the content preference submodule and based on the popularity PS of content type j. j The correlation between user i and content type j (AFNY) j,i Output the extended preference prediction (CPAQ) for user i on content type j. j,i ; The calculation formula for the recommendation analysis submodule is as follows: CRSQ j,i =(CPAQ j,i ×RCN j )+(UPSA j,i ×TA) + (MTC 0.5 ×PPS); in: CRSQ j,i This represents the recommendation score for user i for content type j; RCN j This indicates the timeliness of content type j; TA represents UPSA. j,i The weighting factor, MTC represents the matching degree between user i and content type j, and PPS represents the weighting factor of MTC; The processing procedure of the recommendation analysis submodule is as follows: The predicted extended preference value (CPAQ) of user i for content type j is calculated. j,i And the behavioral feature value UPSA of user i for content type j j,i Input is fed into the recommendation analysis submodule and based on the timeliness RCN of content type j. j Output the CRSQ recommendation score for user i on content type j. j,i ; The processing procedure of the correlation degree submodule is as follows: Step 1: AFNY j,i,new =AFNY j,i,old +α×(CRSQ) j,i -TSD); Step 2: Establish the loop termination condition; Condition 1: |AFNY j,i,new -AFNY j,i,old | < 0.00493; Condition 2: The number of iterations is 123; in: AFNY j,i,new This represents the degree of association between user i and content type j after iteration; AFNY j,i,old The expression represents the correlation between user i and content type j before iteration; α represents the learning rate, used to control the update step size; and TSD represents the comparison threshold, used to determine the CRSQ. j,i Whether the expected level has been achieved; CRSQ based on user i's recommendation score for content type j j,i The difference between the comparison threshold TSD and the value of AFNY is calculated, and the correlation between user i and content type j before the iteration is controlled by the learning rate α. j,i,old The degree of change, outputting the correlation between user i and content type j after iteration. j,i,new The correlation between user i and content type j after iteration is used. j,i,new Replace the correlation between user i and content type j in the content preference submodule with AFNY. j,i To continuously optimize CPAQ j,i and CRSQ j,i And by using conditions one and two, the loop terminates and converges.
2. The software information processing method based on big data according to claim 1, characterized in that: The data collection module can collect user behavior data through web crawler technology, log analysis software, and API interface calls.
3. The software information processing method based on big data according to claim 1, characterized in that: The data processing module includes data cleaning (removing duplicate data and correcting erroneous data), data integration (integrating data from different sources and standardizing the data to ensure consistency in data format), and normalization (converting the data into a format suitable for processing and then normalizing the data to ensure that the data is on the same scale).
Citation Information
Patent Citations
Method, device and system for recommending real-time information
CN106503014A
News recommendation method based on image-text combination
CN120011634A
Method and device for carrying out recommendation prediction by utilizing recommendation model
CN120105329A
Big data analysis method and system based on image processing
CN120256733A