Big data analysis method and system based on image processing
Through the data acquisition and analysis method based on image processing, the CNN model is used to extract image features and combine multiple data to calculate the comprehensive interest index, the problem of incomplete user interest analysis in the existing technology is solved, personalized content push is realized, and user experience is improved.
Patent Information
- Application Number
- CN202510411939.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art relies too much on text data in user interest analysis, ignores the value of image information, cannot comprehensively and accurately grasp user interests, and ignores user interaction data, search data and platform popularity data.
Through the data acquisition and analysis method based on image processing, image features are extracted using pre-trained CNN models, combined with user image behavior data, likes, comments, forwarding data, searches, platform content releases and user interactions, the comprehensive interest index is calculated, and a periodic interest evaluation model is constructed to realize personalized content push.
It realizes comprehensive and accurate analysis of user interests, improves the intelligence of data analysis, and improves the matching degree and user experience of content push.
Smart Images

Figure CN120256733A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis, and particularly to a big data analysis method and system based on image processing. Background Art
[0002] In today's digital age, social media and various online platforms have flourished, and users have shared a vast amount of image information on platforms such as Weibo, Twitter, and Flickr. These images contain rich features of users' interests and hobbies.
[0003] However, the big data analysis methods in the existing technology still have the following deficiencies when applied to the field of user interest analysis:
[0004] On the one hand, most current user interest analysis technologies focus on text data because text data is relatively mature in format and analysis tools and is easy to process. However, this analysis method that overly relies on text data greatly ignores the huge value of image information, resulting in an incomplete and in-depth analysis of user interests.
[0005] On the other hand, when traditional image analysis methods are used to process the images of users on social media platforms to determine the users' interest preferences, they ignore the users' interaction data, search data, and platform popularity data, and cannot comprehensively and accurately grasp users' interests.
[0006] Therefore, a big data analysis method and system based on image processing are introduced. Summary of the Invention
[0007] The purpose of the present invention is to solve the problems pointed out in the background art, and to propose a big data analysis method and system based on image processing.
[0008] The purpose of the present invention can be achieved through the following technical solutions: A big data analysis method based on image processing, including:
[0009] Data acquisition: Based on a preset user evaluation time period; capturing the image behavior data of the user within the current evaluation time period; the image behavior data includes the image data published by the user and the image data of likes, comments, and forwards.
[0010] User analysis: Inputting the image behavior data of the user within the current evaluation time period into a pre-trained CNN model for classification of interest categories, and conducting a comprehensive evaluation after classification to determine the user's periodic interest evaluation model.
[0011] Result Push: According to the periodic interest evaluation model of the user within the current evaluation time period, the proportion of each interest category within the periodic interest evaluation model is used as the proportion of the push content of each interest category pushed to the user by the platform within the next evaluation time period.
[0012] As a preferred implementation manner of the present invention, the image behavior data of the user within the current evaluation time period is input into a pre-trained CNN model for classification of interest categories. Specifically:
[0013] For the image behavior data of the user within the current evaluation time period; using the convolutional layer and pooling layer of the trained CNN model to extract features from the image to be classified, obtaining the feature vector of the image; inputting the extracted feature vector into the fully connected layer and Softmax layer of the model, and the CNN model outputs the probability distribution of the image belonging to each preset interest category, and selects the category with the largest probability value as the interest category to which the image belongs.
[0014] As a preferred implementation manner of the present invention, comprehensive evaluation is performed after classification. Specifically:
[0015] Different interest categories are represented by the number i, i = 1, 2,..., t, where t is the total number of interest categories; the image data published by the user is classified according to the interest category to which it belongs, obtaining the number of images corresponding to each interest category of the user within the current evaluation time period, and taking the number of images corresponding to each interest category of the user as the preliminary interest value of each interest category of the user within the current evaluation time period.
[0016] As a preferred implementation manner of the present invention, comprehensive evaluation after classification further includes:
[0017] Extract the like, comment, and forward image data of the user within the current evaluation time period, and after classifying according to the interest category to which it belongs, determine the number of likes, comments, and forwards of the user corresponding to different interest categories, and analyze to obtain the interaction interest value of different interest categories of the user within the current evaluation time period;
[0018] Obtain the search times of each interest category of the user within the current evaluation time period, and analyze to obtain the search interest score of each interest category of the user within the current evaluation time period;
[0019] Obtain the content release volume and user interaction volume of each interest category of the platform within the current evaluation time period, and analyze to obtain the popularity interest score of each interest category of the platform within the current evaluation time period.
[0020] As a preferred implementation manner of the present invention, determining the periodic interest evaluation model of the user specifically:
[0021] Extract the interaction interest values, search interest scores, and popularity interest scores corresponding to each interest category of the user within the current evaluation time period, multiply them by the set weight coefficients respectively, and then sum them to obtain the interest preference index of each interest category of the user within the current evaluation time period;
[0022] Extract the preliminary interest values and interest preference indices of each interest category of the user within the current evaluation time period, denoted as ps1 and ps2; calculate the average values of each group of preliminary interests and interest preference indices of the user within the past X evaluation time periods respectively, as the reference preliminary interest value and reference interest preference index corresponding to the current evaluation time period, denoted as pf1 and pf2;
[0023] According to the formula Calculate the interest comprehensive index pm of each interest category of the user within the current evaluation time period; where b1 and b2 are the influence weight factors of the preliminary interest value and the interest preference index respectively;
[0024] Sum the interest comprehensive indices pm of each interest category to obtain the index sum, calculate the proportion of the interest comprehensive index pm corresponding to each interest category in the index sum, as the interest division ratio of each interest category within the current evaluation time period;
[0025] Draw a pie chart and divide it into t equal parts according to the interest division ratio of each interest category of the user within the current evaluation time period. After division, mark the numbers of the corresponding interest categories in each equal part. After marking, it serves as the periodic interest evaluation model of the user within the current evaluation time period.
[0026] As a preferred embodiment of the present invention, the obtaining of the interaction interest values of different interest categories of the user within the current evaluation time period is specifically as follows:
[0027] Preset the intervals where each group of times corresponding to the number of likes, comments, and forwards are located. Each interval of the number of likes corresponds to a like basic score; each interval of the number of comments corresponds to a comment basic score; each interval of the number of forwards corresponds to a forward basic score;
[0028] Match the number of likes, comments, and forwards of different interest categories of the user within the current evaluation time period with the corresponding intervals of the number of times respectively, and then sum the obtained like basic scores, comment basic scores, and forward basic scores. After summing, divide by the integer three to obtain the interaction interest values of different interest categories of the user within the current evaluation time period.
[0029] As a preferred embodiment of the present invention, the obtaining of the search interest scores of each interest category of the user within the current evaluation time period is specifically as follows:
[0030] Obtain the number of searches of the user within the current evaluation time period, extract the search terms input by the user for each search, match the search terms with the pre-constructed matching term libraries of each interest category, check each search term in the search term list one by one to determine whether it contains a term in a certain matching term library. If it contains, associate the search term with the interest category corresponding to the matching term library, and after association, increment the search count of the corresponding interest category of the user within the current evaluation time period by one;
[0031] If it does not contain, use the edit distance algorithm to calculate the similarity between the search term and the terms in each matching term library. When the similarity with a certain term exceeds the set threshold, associate the search term with the interest category corresponding to the term, and after association, increment the search count of the corresponding interest category of the user within the current evaluation time period by one;
[0032] After determining the search counts of each interest category of the user within the current evaluation time period, match them with the preset groups of search count ranges. Each group of search count ranges corresponds to a search interest score; obtain the search interest scores of each interest category of the user within the current evaluation time period.
[0033] As a preferred embodiment of the present invention, obtaining the heat interest scores of each interest category of the platform within the current evaluation time period specifically includes:
[0034] Obtain the content release volume and user interaction volume of each interest category of the platform within the current evaluation time period, denoted as gt1 and gt2, and calculate the average value of the content release volume and user interaction volume of each group of each interest category within the past X evaluation time periods respectively as the reference content release volume and reference user interaction volume corresponding to each interest category within the current evaluation time period, denoted as gh1 and gh2; where X > 3 and is a positive integer;
[0035] According to the formula Perform weighted calculation on the content release volume and user interaction volume of each interest category of the platform within the current evaluation time period to determine the heat evaluation value gw of each interest category of the platform within the current evaluation time period; where a1 and a2 are the influence weight factors of the content release volume and user interaction volume respectively;
[0036] Match the heat evaluation value gw of each interest category of the platform within the current evaluation time period with the preset groups of evaluation value ranges. Each group of evaluation value ranges corresponds to a heat interest score; obtain the heat interest scores of each interest category of the platform within the current evaluation time period.
[0037] A big data analysis system based on image processing, comprising:
[0038] The image analysis module is used to capture the image behavior data of the user within the current evaluation time period; the image behavior data includes the image data published by the user and the liked, commented, and forwarded image data;
[0039] The model construction module is used to input the image behavior data of the user within the current evaluation time period into a pre-trained CNN model for classification of interest categories, and conduct a comprehensive evaluation after classification to determine the user's periodic interest evaluation model;
[0040] The push adjustment module is used to, according to the periodic interest evaluation model of the user within the current evaluation time period, take the proportion of each interest category in the periodic interest evaluation model as the proportion of the push content of each interest category pushed to the user by the platform within the next evaluation time period.
[0041] Compared with the prior art, the beneficial effects of the present invention are:
[0042] The present invention determines the user's interest by comprehensively considering various data. When conducting a comprehensive evaluation after classification, not only is the preliminary interest value determined based on the number of images published by the user, but also the interaction interest value is calculated by extracting the liked, commented, and forwarded data, the search interest score is analyzed by obtaining the search times, and the popularity interest score is obtained by combining the platform content release volume and the user interaction volume. By multiplying these different-dimensional data with the set weight coefficients respectively and summing them to obtain the interest deviation index, and then combining the preliminary interest value to calculate the interest comprehensive index, the user's interest is comprehensively and accurately grasped, and the intelligence level of data analysis is improved;
[0043] The present invention realizes personalized push by, according to the constructed periodic interest evaluation model, taking the proportion of each interest category in the model as the proportion of the content pushed to the user by the platform within the next evaluation time period, improves the matching degree between the pushed content and the user's interest, and enhances the user experience;
[0044] The present invention captures the user's image behavior data, including the published, liked, commented, and forwarded image data, also records the image metadata, processes the images using a pre-trained CNN model, extracts feature vectors from the images to determine the interest categories to which the images belong, makes full use of the image data, makes up for the defect of the traditional technology's over-reliance on text data, and realizes a more comprehensive analysis of the user's interest. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the accompanying drawings.
[0046] Figure 1 is the flow chart of the present invention;
[0047] Figure 2 is the system diagram of the present invention. Detailed implementation mode
[0048] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.
[0049] Embodiment 1
[0050] Please refer to Figure 1 As shown, a big data analysis method based on image processing includes:
[0051] Data collection: Based on a preset user evaluation time period; for example, an evaluation is performed every 2 days; the image behavior data of the user within the current evaluation time period is captured; the image behavior data includes the image data published by the user and the liked, commented, and forwarded image data;
[0052] It should be noted that a crawler program specifically developed for mainstream social media platforms such as Weibo, Twitter, and the photo-sharing website Flickr is used; according to the data interface specifications and anti-crawler mechanisms of each platform, technologies such as dynamic IP proxy, random request headers, and randomized request intervals are adopted to stably and efficiently capture the image data published by users. At the same time, the metadata of the images, such as the publication time, user ID, image resolution, image format, etc., is recorded;
[0053] Preprocessing operations are performed on the filtered images, including removing duplicate or low-quality images, image enhancement, denoising, and size normalization. Methods such as histogram equalization and adaptive filtering are used to enhance the contrast and clarity of the images, remove interference such as salt-and-pepper noise and Gaussian noise, and uniformly adjust all images to a fixed size;
[0054] User analysis: The image behavior data of the user within the current evaluation time period is input into a pre-trained CNN model for classification of interest categories, and a comprehensive evaluation is performed after classification to determine the user's periodic interest evaluation model; the interest categories include but are not limited to food, travel, sports, art, technology, etc.;
[0055] Specifically:
[0056] S1: For the image behavior data of the user within the current evaluation time period, perform operations such as size normalization and image enhancement according to the preprocessing method during model training; use the convolutional layer and pooling layer of the trained CNN model to extract features from the image to be classified, obtaining the feature vector of the image; input the extracted feature vector into the fully connected layer and Softmax layer of the model, and the CNN model outputs the probability distribution of the image belonging to each preset interest category, and select the category with the largest probability value as the interest category to which the image belongs;
[0057] S2: Different interest categories are represented by the number i, where i = 1, 2,..., t, and t is the total number of interest categories; classify the image data published by the user according to the interest category to which it belongs, obtaining the number of images corresponding to each interest category of the user within the current evaluation time period, and use the number of images corresponding to each interest category of the user as the preliminary interest value of each interest category of the user within the current evaluation time period;
[0058] S3: Extract the like, comment, and forward image data of the user within the current evaluation time period, and after classifying them according to the interest category to which they belong, determine the number of likes, comments, and forwards of the user corresponding to different interest categories;
[0059] Preset the intervals where each group of times corresponding to the number of likes, comments, and forwards is located. Each interval of the number of likes corresponds to a like base score; each interval of the number of comments corresponds to a comment base score; each interval of the number of forwards corresponds to a forward base score; the ranges of the like base score, comment base score, and forward base score are all set between 1 and 10 and are positive integers, and the more the number of likes, comments, and forwards, the higher the corresponding base score obtained;
[0060] Match the number of likes, comments, and forwards of the user in different interest categories within the current evaluation time period with the corresponding intervals of times respectively, and then sum the like base score, comment base score, and forward base score obtained from the matching, and divide the sum by the integer three to obtain the interaction interest value of the user in different interest categories within the current evaluation time period;
[0061] It should be noted that by setting different intervals and corresponding different base scores for the number of likes, comments, and forwards, the participation degree of the user in different interest categories can be measured more carefully. The more the number of times, the higher the base score, so that the interest categories with frequent interactions can obtain higher scores, thereby more accurately evaluating the user's preference degree for various interests and avoiding the one-sidedness that may be brought by single-dimensional evaluation;
[0062] It not only considers the content classification of images, but also conducts a quantitative analysis of user interests from multiple dimensions through user behavior data such as likes, comments, and forwards, more comprehensively reflecting the true attention of users to different interest categories.
[0063] S4: Obtain the number of searches of the user within the current evaluation time period, extract the search terms input by the user for each search, match the search terms with the pre-constructed matching term libraries of each interest category, check each search term in the search term list one by one to determine whether it contains a term in a certain matching term library. If it contains, associate the search term with the interest category corresponding to the matching term library, and after association, increment the search count of the corresponding interest category of the user within the current evaluation time period by one;
[0064] It should be noted that for each interest category, a corresponding matching term library is pre-constructed, and the matching term library can be constructed through manual collation, data analysis, or machine learning methods. For example, for the "Food" category, the matching term library can include "pizza", "sushi", "hot pot", etc.; for the "Travel" category, the matching term library can include "Paris", "New York", "travel guide", etc.; regularly update the matching term library to adapt to newly emerging vocabulary and changes in user interests.
[0065] If it does not contain, use the edit distance algorithm (such as Levenshtein distance) to calculate the similarity between the search term and the terms in each matching term library. When the similarity with a certain term exceeds the set threshold, associate the search term with the interest category corresponding to the term, and after association, increment the search count of the corresponding interest category of the user within the current evaluation time period by one;
[0066] After determining the search counts of each interest category of the user within the current evaluation time period, match them with the preset groups of search count ranges. Each group of search count ranges corresponds to a search interest score; the range of the search bonus score is set from 1 to 10, and the more the search count, the higher the corresponding search bonus score obtained; obtain the search interest scores of each interest category of the user within the current evaluation time period;
[0067] It should be noted that by conducting a quantitative analysis of the search counts and obtaining the search bonus scores through matching with the preset search count ranges, it can further quantify the search behavior of users for different interest categories; this quantification method can more intuitively reflect the energy invested by users in each interest category, providing richer and more accurate data dimensions for constructing a user interest model, helping to more comprehensively evaluate user interests, and making up for the deficiencies of the image processing method in quantifying user search behavior.
[0068] S5: Obtain the content release volume and user interaction volume of each interest category within the current evaluation time period of the platform, denoted as gt1 and gt2. Calculate the average value of the content release volume and user interaction volume of each group of each interest category within the past X evaluation time periods respectively, as the reference content release volume and reference user interaction volume corresponding to each interest category within the current evaluation time period, denoted as gh1 and gh2; where X > 3 and X is a positive integer;
[0069] According to the formula Perform weighted calculation on the content release volume and user interaction volume of each interest category within the current evaluation time period of the platform, so as to determine the heat evaluation value gw of each interest category within the current evaluation time period of the platform; where a1 and a2 are the influence weight factors of the content release volume and user interaction volume respectively;
[0070] Match the heat evaluation value gw of each interest category within the current evaluation time period of the platform with the preset groups of evaluation value ranges. Each group of evaluation value ranges corresponds to a heat interest score; the higher the heat evaluation value gw, the higher the corresponding matched heat interest score, and the heat interest score range is set from 1 to 10; obtain the heat interest scores of each interest category within the current evaluation time period of the platform;
[0071] It should be noted that by comprehensively considering the content release volume and user interaction volume within the current evaluation time period and combining the average value of multiple past evaluation time periods as a reference, it can more comprehensively and accurately evaluate the heat of each interest category and more accurately reflect the true heat of the platform for different interest categories.
[0072] S6: Extract the interaction interest value, search interest score, and heat interest score corresponding to each interest category of the user within the current evaluation time period, multiply them by the set weight coefficients respectively, and then sum them to obtain the interest preference index of each interest category of the user within the current evaluation time period;
[0073] S7: Extract the preliminary interest value and interest preference index of each interest category of the user within the current evaluation time period, denoted as ps1 and ps2; calculate the average value of the groups of preliminary interests and interest preference indexes of the user within the past X evaluation time periods respectively, as the reference preliminary interest value and reference interest preference index corresponding to the current evaluation time period, denoted as pf1 and pf2;
[0074] According to the formula Calculate the interest comprehensive index pm of each interest category of the user within the current evaluation time period; where b1 and b2 are the influence weight factors of the preliminary interest value and interest preference index respectively;
[0075] Sum the comprehensive interest indices pm of each interest category to obtain the sum of indices, and calculate the proportion of the comprehensive interest index pm corresponding to each interest category in the sum of indices as the interest division ratio of each interest category within the current evaluation time period;
[0076] Draw a pie chart and divide it into t equal parts according to the interest division ratio of each interest category of the user within the current evaluation time period. After division, label the numbers of the corresponding interest categories in each equal part. After labeling, it serves as the periodic interest evaluation model of the user within the current evaluation time period;
[0077] It should be noted that not only the preliminary interest values of each interest category of the user within the current evaluation time period (which may be obtained from multiple aspects such as image-related behaviors and search behaviors) are considered, but also the interest preference index is introduced. At the same time, the mean values of multiple past evaluation time periods are combined as a reference, and multi-dimensional data is comprehensively analyzed; this comprehensive evaluation method can more accurately capture the dynamic changes and true preferences of the user's interests, avoid the one-sidedness brought by single-dimensional evaluation, and make the evaluation results more in line with the actual interest situation of the user.
[0078] Result push: According to the periodic interest evaluation model of the user within the current evaluation time period, take the proportion of each interest category within the periodic interest evaluation model as the push content ratio of each interest category pushed to the user by the platform during the next evaluation time period;
[0079] Embodiment 2
[0080] Please refer to Figure 2 As shown, based on a big data analysis method based on image processing provided in Embodiment 1 of the present application, Embodiment 2 of the present application proposes a big data analysis system based on image processing. Embodiment 2 is merely a preferred manner of Embodiment 1, and the implementation of Embodiment 2 will not affect the independent implementation of Embodiment 1.
[0081] Specifically, the difference of a big data analysis system based on image processing provided in Embodiment 2 of the present application is that it includes an image analysis module, a model construction module, and a push adjustment module;
[0082] The image analysis module is used to capture the image behavior data of the user within the current evaluation time period; the image behavior data includes the image data published by the user and the liked, commented, and forwarded image data;
[0083] The model construction module is used to input the image behavior data of the user within the current evaluation time period into a pre-trained CNN model for interest category classification, and perform comprehensive evaluation after classification to determine the periodic interest evaluation model of the user;
[0084] The push adjustment module is used to, according to the periodic interest evaluation model of the user within the current evaluation time period, use the proportion of each interest category in the periodic interest evaluation model as the push content proportion of each interest category pushed to the user by the platform within the next evaluation time period;
[0085] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to only the specific embodiments. Obviously, many modifications and changes can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A big data analysis method based on image processing, characterized in that Including: Data collection: Based on a preset user evaluation time period; Scraping the user's image behavior data within the current evaluation time period; The image behavior data includes the image data published by the user and the liked, commented, and forwarded image data; User analysis: Inputting the user's image behavior data within the current evaluation time period into a pre-trained CNN model for classification of interest categories, and conducting a comprehensive evaluation after classification to determine the user's periodic interest evaluation model; Result push: According to the periodic interest evaluation model of the user within the current evaluation time period, taking the proportion of each interest category in the periodic interest evaluation model as the push content proportion of each interest category pushed to the user by the platform within the next evaluation time period.
2. The big data analysis method based on image processing according to claim 1, characterized in that Inputting the user's image behavior data within the current evaluation time period into a pre-trained CNN model for classification of interest categories, specifically: For the user's image behavior data within the current evaluation time period; Using the convolutional layer and pooling layer of the trained CNN model to extract features from the image to be classified, obtaining the feature vector of the image; Inputting the extracted feature vector into the fully connected layer and Softmax layer of the model, and the CNN model outputs the probability distribution of the image belonging to each preset interest category, and selecting the category with the largest probability value as the interest category to which the image belongs.
3. The big data analysis method based on image processing according to claim 2, characterized in that Conducting a comprehensive evaluation after classification, specifically: Using the number i to represent different interest categories, i = 1, 2,..., t, where t is the total number of interest categories; Classifying the image data published by the user according to the interest category to which it belongs, obtaining the number of images corresponding to each interest category of the user within the current evaluation time period, and taking the number of images corresponding to each interest category of the user as the preliminary interest value of each interest category of the user within the current evaluation time period.
4. The big data analysis method based on image processing according to claim 3, wherein Conducting a comprehensive evaluation after classification also includes: Extracting the liked, commented, and forwarded image data of the user within the current evaluation time period, classifying them according to the interest category to which they belong, determining the number of likes, comments, and forwards of the user corresponding to different interest categories, and analyzing to obtain the interactive interest value of different interest categories of the user within the current evaluation time period; Obtaining the search times of each interest category of the user within the current evaluation time period, and analyzing to obtain the search interest score of each interest category of the user within the current evaluation time period; Obtaining the content release volume and user interaction volume of each interest category of the platform within the current evaluation time period, and analyzing to obtain the popularity interest score of each interest category of the platform within the current evaluation time period.
5. A big data analysis method based on image processing according to claim 4, characterized in that Determining the user's periodic interest evaluation model, specifically: Extracting the interactive interest value, search interest score, and popularity interest score corresponding to each interest category of the user within the current evaluation time period, multiplying them by the set weight coefficients respectively, and then summing to obtain the interest bias index of each interest category of the user within the current evaluation time period. Extract the preliminary interest values and interest bias indices for each interest category of the user within the current evaluation time period, denoted as ps1 and ps2; calculate the mean of each group of preliminary interests and interest bias indices of the user within the past X evaluation time periods as the reference preliminary interest value and reference interest bias index corresponding to the current evaluation time period, denoted as pf1 and pf2; According to the formula calculate the comprehensive interest index pm of each interest category of the user within the current evaluation time period; where b1 and b2 are the influence weight factors of the preliminary interest value and the interest deviation index respectively; Sum up the interest comprehensive indices pm of each interest category to obtain the index sum, and calculate the proportion of the interest comprehensive index pm corresponding to each interest category in the index sum as the interest division ratio of each interest category within the current evaluation time period; Draw a pie chart and divide it into t equal parts according to the interest division ratio of each interest category of the user within the current evaluation time period. After division, label the numbers of the corresponding interest categories in each equal part. After labeling, it serves as the periodic interest evaluation model of the user within the current evaluation time period.
6. The big data analysis method based on image processing according to claim 5, characterized in that The obtaining of the interactive interest values of different interest categories of the user within the current evaluation time period is specifically as follows: Preset the intervals where each group of times corresponding to the number of likes, comments, and forwards are located. Each interval where each group of times of the number of likes is located corresponds to a like basic score; each interval where each group of times of the number of comments is located corresponds to a comment basic score; each interval where each group of times of the number of forwards is located corresponds to a forward basic score; Match the number of likes, comments, and forwards of different interest categories of the user within the current evaluation time period with the corresponding intervals where the times are located respectively, and then sum up the like basic score, comment basic score, and forward basic score obtained from the matching. After summing up, divide by the integer three to obtain the interactive interest values of different interest categories of the user within the current evaluation time period.
7. A big data analysis method based on image processing according to claim 6, characterized in that, The obtaining of the search interest scores of each interest category of the user within the current evaluation time period is specifically as follows: Obtain the number of searches of the user within the current evaluation time period, extract the search terms input by the user for each search, match the search terms with the pre-constructed matching word library of each interest category, check each search term in the search term list one by one to determine whether it contains a vocabulary in a certain matching word library. If it contains, associate the search term with the interest category corresponding to the matching word library, and after association, increment the number of searches of the corresponding interest category of the user within the current evaluation time period by one; If it does not contain, use the edit distance algorithm to calculate the similarity between the search term and the vocabulary in each matching word library. When the similarity with a certain vocabulary exceeds the set threshold, associate the search term with the interest category corresponding to the vocabulary, and after association, increment the number of searches of the corresponding interest category of the user within the current evaluation time period by one; After determining the number of searches of each interest category of the user within the current evaluation time period, match it with the preset ranges of each group of search times. Each range of search times corresponds to a search interest score; Obtain the search interest scores of each interest category of the user within the current evaluation time period.
8. A big data analysis method based on image processing according to claim 7, characterized in that, The obtaining of the popularity interest scores of each interest category of the platform within the current evaluation time period is specifically as follows: Obtain the content release volume and user interaction volume of each interest category within the current evaluation time period on the platform, denoted as gt1 and gt2. Calculate the average value of the content release volume and user interaction volume of each group of each interest category within the past X evaluation time periods respectively, and use it as the reference content release volume and reference user interaction volume of each interest category corresponding to the current evaluation time period, denoted as gh1 and gh2; where X > 3 and is a positive integer; According to the formula perform weighted calculations on the content release volume and user interaction volume of each interest category on the platform during the current evaluation time period, so as to determine the popularity evaluation value gw of each interest category on the platform during the current evaluation time period; where a1 and a2 are the influence weight factors of the content release volume and user interaction volume respectively; Match the heat evaluation value gw of each interest category on the platform within the current evaluation time period with the preset groups of evaluation value ranges, and each group of evaluation value ranges corresponds to a heat interest score; Obtain the heat interest scores of each interest category on the platform within the current evaluation time period.
9. A big data analysis system based on image processing, which is applied to a big data analysis method based on image processing as claimed in any one of claims 1-8 above, characterized in that, Including: The image analysis module is used to capture the image behavior data of users within the current evaluation time period; the image behavior data includes the image data published by users and the image data of likes, comments, and forwards; The model construction module is used to input the image behavior data of users within the current evaluation time period into a pre-trained CNN model for classification of interest categories, and conduct a comprehensive evaluation after classification to determine the periodic interest evaluation model of users; The push adjustment module is used to, according to the periodic interest evaluation model of users within the current evaluation time period, use the proportion of each interest category within the periodic interest evaluation model as the push content proportion of each interest category pushed to users by the platform within the next evaluation time period.
Citation Information
Cited By
Software information processing method based on big data
CN120804430A
A software information processing method based on big data
CN120804430B