Permeation shadow deduction method and system
By integrating multi-source user data and simulation control technology, combined with multi-device switching and proxy IP pool rotation, the problems of incomplete user portraits and poor real-time performance in the recommendation system are solved, high-precision user behavior modeling and data collection stability are achieved, and the real-time response and accuracy of the recommendation system are improved.
Patent Information
- Application Number
- CN202510846888.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN120744231A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer science and technology, and in particular to a method and system for infiltration deduction. Background Art
[0002] With the rapid development of short video platforms, platform recommendation algorithms have gradually become the main way for users to obtain information. Due to the black box characteristics of the recommendation system, content bias and information lag, some public opinion events have not been discovered in time or effectively monitored. Traditional public opinion monitoring methods rely on public data and manual analysis, which are inefficient and have large delays. Although existing technical means use crawlers to capture information, they are often limited by the platform's anti-crawler mechanism and lack the ability to actively simulate the behavior of the recommendation system. Therefore, they cannot efficiently obtain potential public opinion hotspots. This method usually relies on traditional data analysis tools, such as collaborative filtering algorithms or content-based recommendation systems, which aim to improve the relevance and accuracy of recommended content.
[0003] While these approaches have improved user experience to a certain extent, their main drawback is that they fail to fully consider the value of multi-source, heterogeneous data. When processing user data on social media platforms, in addition to basic demographic information, there is also a large amount of social interaction records, multimedia consumption habits, and geographic location data. This data is crucial for gaining a deeper understanding of user interests and preferences, but due to the lack of effective integration mechanisms, existing technologies struggle to fully utilize this rich information resource. Furthermore, current methods are mostly limited to static datasets and cannot update user profiles in real time to adapt to changes in user behavior, which limits the dynamic responsiveness of recommendation systems. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a shadow inference method to solve the problem of insufficient recommendation accuracy and responsiveness caused by incomplete user portraits and poor real-time performance in existing recommendation systems by integrating multi-source heterogeneous user data and combining it with dynamic behavior simulation technology.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for inference inference, which includes collecting multi-source user feature data and generating user portraits;
[0008] Based on user portraits, use simulation control technology to simulate user operations on the platform;
[0009] Based on the user's actions on the platform, the behavior path is adjusted to guide the push of target content. By analyzing the content of the platform's recommendation flow, the platform's recommendation logic is identified;
[0010] Adopt multi-device switching and proxy IP pool rotation technology to simulate the real access environment;
[0011] Capture recommended content on the platform by performing sentiment analysis and popularity evaluation on the recommended content on the platform;
[0012] Based on the recommended content on the platform, the risk level of the platform's video content is assessed in real time and an early warning report is generated.
[0013] As a preferred solution of the infiltration deduction method of the present invention, wherein: collecting multi-source user feature data and generating user portraits includes the following steps:
[0014] Use API to obtain basic information about users’ gender, age, and region to form user interaction behaviors;
[0015] Based on user interaction behavior, use the TF-IDF algorithm to analyze user interest preferences;
[0016] Integrate user interaction behaviors and user interest preferences to build user portraits.
[0017] As a preferred solution of the infiltration deduction method of the present invention, wherein: according to the user portrait, the user's operation on the platform is simulated by using simulation control technology, including the following steps:
[0018] Set operation probabilities for interest categories through user portraits to form a preliminary behavior simulation plan;
[0019] The preliminary behavior simulation plan is processed using simulation control technology to simulate user interaction behavior and generate operation sequences similar to real users.
[0020] As a preferred solution of the infiltration deduction method of the present invention, the following steps are included: based on the user's operation on the platform, the behavior path is adjusted to guide the push of target content, and the recommendation logic of the platform is identified by analyzing the content of the platform recommendation flow.
[0021] Structured log data is generated through similar operation sequences of real users;
[0022] Clean the structured log data and add labels according to the content type to form a cleaned feature dataset;
[0023] Extract content feature vectors from the feature data set, use the decision tree algorithm to train the recommendation rule model based on the content feature vectors to identify the recommendation logic, and obtain a preliminary recommendation rule model;
[0024] According to the recommendation rule model, the initial behavior path strategy of the simulated user is set, and the behavior path is dynamically adjusted through the Q-learning algorithm;
[0025] Simulate user behavior operations based on the behavior path, analyze the content of the platform recommendation flow and identify the platform's recommendation logic.
[0026] As a preferred solution of the infiltration deduction method of the present invention, wherein: multi-device switching and proxy IP pool rotation technology are used to simulate the real access environment, including the following steps:
[0027] Collect device features from the internal database and build a device fingerprint library;
[0028] Based on the anti-crawler mechanism, the disguised device configuration set is filtered out from the device fingerprint library;
[0029] Through the proxy IP service, an initial proxy IP pool is established, and each IP address in the proxy IP pool is tested for speed and anonymity, and the availability score is calculated to obtain the quality assessment results of the proxy IP;
[0030] Based on the quality assessment results of the proxy IP, HTTP header information is set to simulate access from specific devices and geographic locations, simulating the real access environment.
[0031] As a preferred solution of the infiltration deduction method of the present invention, wherein: by performing sentiment analysis and popularity evaluation on the recommended content on the platform, the recommended content on the platform is captured, including the following steps:
[0032] Extract the data types that need to be crawled from the recommended content on the platform and form a crawling list;
[0033] Calculate the popularity score of each content based on structured log data;
[0034] Based on the crawled list, the popularity score of each content is sorted in descending order to form a popularity list;
[0035] Clean the structured log data to form a vocabulary. Use a deep learning model to perform sentiment classification on the comments in the vocabulary, identify positive, negative, and neutral sentiment tendencies, and calculate the sentiment tendency score.
[0036] Combine the sentiment tendency score and popularity list to form recommended content on the platform.
[0037] As a preferred solution of the infiltration deduction method of the present invention, the risk level of the platform video content is evaluated in real time based on the recommended content on the platform, and an early warning report is generated, including the following steps:
[0038] Calculate the negative sentiment ratio of each piece of content based on the sentiment tendency score;
[0039] Set a heat threshold based on the heat list to judge abnormally high heat content;
[0040] Use the TF-IDF algorithm to extract keywords from recommended content and calculate the frequency of occurrence of sensitive words;
[0041] Based on the frequency of sensitive words, unusually high-profile content, and sentiment ratio, a risk scoring model is constructed to calculate the risk score for each piece of content.
[0042] An early warning report is generated based on the risk score of each content.
[0043] In a second aspect, the present invention provides a penetration inference system, comprising a user portrait construction module, which collects multi-source user feature data and generates a user portrait;
[0044] User behavior simulation module, based on user portraits, uses simulation control technology to simulate user operations on the platform;
[0045] The recommendation feedback guidance module adjusts the behavior path and guides the push of target content based on the user's actions on the platform. It also identifies the platform's recommendation logic by analyzing the content of the platform's recommendation flow.
[0046] The anti-crawler bypass module uses multi-device switching and proxy IP pool rotation technology to simulate the real access environment;
[0047] The content collection module captures recommended content on the platform by performing sentiment analysis and popularity evaluation on the recommended content on the platform;
[0048] The public opinion warning module evaluates the risk level of the platform's video content in real time based on the recommended content on the platform and generates a warning report.
[0049] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the infiltration deduction method described in the first aspect of the present invention is implemented.
[0050] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the infiltration deduction method described in the first aspect of the present invention is implemented.
[0051] The beneficial effects of the present invention are as follows: through user profiling, simulation control technology is used to simulate user operations on the platform, and multi-device switching and proxy IP pool rotation technology are used to simulate a real access environment. By setting operation probabilities for interest categories in user profiles through simulation control technology and generating operation sequences similar to real users, high-precision modeling of user behavior is achieved, providing structured and representative behavioral data support for subsequent recognition and recommendation logic. At the same time, through multi-device switching and proxy IP pool rotation technology, combined with a device fingerprint library and proxy IP availability evaluation mechanism, a highly simulated access environment is constructed, effectively circumventing platform access restrictions and improving the authenticity and stability of data collection. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0053] Figure 1 This is a flow chart of the infiltration deduction system.
[0054] Figure 2 Schematic diagram of the infiltration deduction method.
[0055] Figure 3 Flowchart providing background information for the analysis.
[0056] Figure 4 Flowchart of early warning report. DETAILED DESCRIPTION
[0057] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0058] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0059] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0060] Reference Figures 1 to 4, is an embodiment of the present invention, which provides a method for inference, comprising the following steps:
[0061] S1. Collect multi-source user feature data and generate user portraits.
[0062] S1.1. Use the API to obtain basic information about the user’s gender, age, and region to form user interaction behaviors.
[0063] Furthermore, the specific process of using API to obtain basic information about the user's gender, age and region to form user interactive behavior is as follows: first, the user's gender information is obtained through the platform's open interface or authorized access data source. This information comes from the user's active filling in during the registration or use of platform services or is obtained by system recognition; then, age data is extracted from the user's account information. This data is usually the year of birth entered by the user when registering or estimated by the platform based on usage behavior, and the user's region information is obtained. This information is based on the IP address positioning when the user logs in or the permanent residence identifier manually set by the user; the above three information are integrated and associated with the operation records generated by the user on the platform to form a complete user interactive behavior.
[0064] S1.2. Analyze user interest preferences using the TF-IDF algorithm based on user interaction behavior.
[0065] Furthermore, the specific process of using the TF-IDF algorithm to analyze user interest preferences based on user interactive behavior is as follows: first, the operation records generated by the user on the platform are textually processed, and the content title, description and label corresponding to each operation behavior are converted into analyzable text data; then the text data is segmented, and stop words and words without actual semantics are removed to generate a set of words used to represent content features; then the frequency of occurrence of each word in a single user behavior text is counted as the word frequency, that is, the number of user behavior texts containing the word. Through the above-mentioned ratio of word frequency to document frequency, the TF-IDF value of each word is used to measure its importance in the current user's interest expression. Finally, the TF-IDF values of all words are classified and summarized by content category to form a reflection of the user's interest preferences.
[0066] S1.3. Integrate user interaction behaviors and user interest preferences to build user portraits.
[0067] Furthermore, the user's gender, age and regional basic information obtained through the API is structurally associated with the user's operation records on the platform to form a description of user interaction behavior with demographic characteristics and behavioral trajectories; at the same time, the user interest preference results obtained by using the TF-IDF algorithm analysis, that is, the keyword weight distribution reflecting the content category, are matched with the content browsing, interaction and other behaviors in the above-mentioned user interaction behaviors; then, the data fusion method is used to splice the user's basic attributes, behavior sequence and interest weight according to the preset data dimensions to form a unified data representation containing multi-dimensional information; finally, by normalizing the information in the data representation and setting field identifiers according to the requirements of the platform recommendation task, a user portrait for subsequent simulation control technology use is generated.
[0068] S2. Based on the user portrait, use simulation control technology to simulate the user's operations on the platform.
[0069] S2.1. Set operation probabilities for interest categories through user portraits to form a preliminary behavior simulation plan.
[0070] Furthermore, based on the basic information of the user's gender, age and region contained in the constructed user portrait and the weight distribution of interest preference keywords obtained through TF-IDF algorithm analysis, the interest dimensions are divided according to content categories, and the keyword weights under each interest dimension are normalized to the interest intensity value of the corresponding category; then, according to the intensity ratio of each interest dimension, the probability distribution of selected operations during the simulation process is set for each type of content. For example, if a user has a higher total keyword weight for technology content, the operation probability of technology content will increase accordingly; then, combined with the behavioral characteristics such as content browsing time and click frequency involved in user interactive behavior, the execution frequency parameters of different operation types are set; finally, the interest category operation probability and the operation type execution frequency are combined to generate a behavior sequence planning scheme for guiding subsequent simulation control technology processing as a preliminary behavior simulation plan.
[0071] S2.2. Use simulation control technology to process the preliminary behavior simulation plan, simulate user interaction behavior, and generate operation sequences similar to real users.
[0072] Furthermore, the interest category operation probability and operation type execution frequency parameters set in the preliminary behavior simulation plan are loaded into the simulation control process as the driving basis for the generation of the behavior sequence; then, the operation path model is constructed according to the platform page structure and interaction logic, and each operation action is mapped to the corresponding functional control, and triggered in sequence according to the set time interval and sequence. On this basis, a random perturbation mechanism is introduced to fine-tune the operation time interval and click position to enhance the naturalness and authenticity of the behavior sequence; finally, by integrating the operation path model, behavior driving parameters and random perturbation results, an operation sequence with timestamp mark and operation type identification is generated.
[0073] S3. Based on the user's actions on the platform, adjust the behavior path to guide the push of target content, and identify the platform's recommendation logic by analyzing the content of the platform's recommendation flow.
[0074] S3.1. Structured log data is formed through similar operation sequences of real users.
[0075] Furthermore, the specific process of forming structured log data through similar operation sequences of real users is as follows: first, the operation sequence with timestamp mark and operation type identifier generated by simulation control technology is classified into event, and divided into basic behavior units such as page loading, content browsing, clicking, sliding, likes, and comments according to the type of operation action; then, a fixed format data field is defined for each type of behavior unit, including operation time, operation type, target space identifier, page location information and related content identifier; then, each operation record is organized into a key-value pair according to the preset data structure, and arranged in chronological order to form behavior log entries with contextual relationships, and all log entries are written into the log data in a unified coding format.
[0076] S3.2. Clean the structured log data and add labels according to the content type to form a cleaned feature dataset.
[0077] Furthermore, the structured log data is cleaned and tagged according to the content type to form a cleaned feature data set. The specific process is as follows: first, each log entry in the structured log data is parsed to extract the original fields such as operation time, operation type, target space identifier, page location information and associated content identifier; then, invalid operation records are removed through regular expression matching and keyword filtering methods, such as repeated clicks, abnormal time intervals or log entries with missing key fields; then, based on the associated content identifier, the corresponding content type information is obtained from the platform content database, and the content types include but are not limited to categories such as technology, entertainment, sports, and finance; the extracted content type information is then attached as a tag to the corresponding log record to form an enhanced data entry with behavioral characteristics and content classification tags; finally, all processed log entries are organized according to a unified field format and stored in the form of structured files to form a cleaned feature data set.
[0078] S3.3. Extract content feature vectors from the feature data set, use the decision tree algorithm to train the recommendation rule model based on the content feature vectors to identify the recommendation logic, and obtain a preliminary recommendation rule model.
[0079] Furthermore, content feature vectors are extracted from the feature data set, and a decision tree algorithm is used to train a recommendation rule model based on the content feature vectors to identify the recommendation logic. The specific process of obtaining a preliminary recommendation rule model is as follows: first, feature encoding is performed on each log record in the cleaned feature data set, and fields such as operation time, operation type, page location information, and content type are converted into numerical feature variables to form a content feature vector; then, the training samples are divided according to the user behavior sequence, and each sample contains a set of content feature vectors corresponding to continuous operations and the recommendation result labels that are ultimately triggered; then, the decision tree algorithm is used to model the mapping relationship between the content feature vectors and the recommendation result labels in all samples, and the optimal feature division node is selected through information gain calculation, and the decision path for judging the recommendation logic is gradually constructed; finally, a preliminary recommendation rule model consisting of multiple decision rules is output.
[0080] S3.4. According to the recommendation rule model, set the initial behavior path strategy of the simulated user and dynamically adjust the behavior path through the Q-learning algorithm.
[0081] Furthermore, according to the recommendation rule model, the initial behavior path strategy of the simulated user is set, and the specific process of dynamically adjusting the behavior path through the Q-learning algorithm is as follows: first, based on the preliminary recommendation rule model, the mapping relationship between the content feature vector and the recommendation result is analyzed, and the priority ranking of different content types in the recommendation system is determined, and the behavior preference weights of the simulated user in the initial stage are set accordingly; then the behavior preference weights are used as input parameters to construct the state space and action space, where the state represents the content browsing stage of the current operation, and the action is used to describe the type of operation that may be performed, such as clicking, sliding or staying; then, the Q-learning algorithm is introduced in the process of simulated user interaction with the platform, and the Q value table is updated according to the feedback signal obtained after each operation, and the feedback signal includes indicators such as content exposure time, changes in recommended content and operation response delay; finally, the behavior path is continuously optimized through multiple rounds of interaction.
[0082] S3.5. Simulate user behavior based on the behavior path, analyze the content of the platform recommendation flow, and identify the platform's recommendation logic.
[0083] Furthermore, based on the behavior path strategy optimized by the Q-learning algorithm, a user behavior operation sequence that conforms to the current recommendation rule model is first generated. The behavior operations include page loading, content browsing, clicking, sliding and interactive actions, and are executed in sequence according to the set time intervals. During the behavior simulation process, the recommended content stream returned by the platform is captured in real time. The recommended content stream contains continuously displayed content items and their accompanying metadata information, such as content title, description, tags and sorting position; then the captured recommended content stream is structured and parsed to extract the content feature vector of each content, and the input sample is constructed in combination with the behavioral state of the simulated user at that time point. Then, the input sample is matched and verified with the preliminary recommendation rule model to identify the association rules between the changes in recommended content and user behavior, further refine and supplement the conditional branches and weight distribution mechanism in the recommendation logic, and form an explainable platform recommendation logic through the inductive analysis of multiple rounds of simulation results.
[0084] S4. Use multi-device switching and proxy IP pool rotation technology to simulate the real access environment.
[0085] S4.1. Collect device features from the internal database and build a device fingerprint library.
[0086] Furthermore, by accessing the device information table stored in the internal database, the hardware and software configuration data related to the device are extracted. The device characteristics include but are not limited to the device model, operating system version, browser type and version, screen resolution, network connection type and device unique identifier; the extracted device characteristics are then standardized, the field format is unified and duplicate or missing data is removed to ensure that each record is complete and can be used for unique identification; the cleaned device characteristics are then organized into device fingerprint entries according to a preset data structure, and each device fingerprint entry contains a complete set of device attribute combinations for describing the access characteristics of a specific device; finally, all device fingerprint entry sets are stored in a dedicated database table to form a device fingerprint library that can be used for subsequent access environment simulation.
[0087] S4.2. Based on the anti-crawler mechanism, filter out the disguised device configuration set from the device fingerprint library.
[0088] Furthermore, based on the anti-crawler mechanism, the specific process of screening out disguised device configuration sets from the device fingerprint library is as follows: first, analyze the anti-crawler strategy adopted by the platform to identify key device feature dimensions used to detect abnormal access behavior, including but not limited to user agent strings, browser fingerprint consistency, IP address and device geographic location matching, and access frequency patterns; then compare each device fingerprint entry in the device fingerprint library with the above key features to detect whether there are abnormal feature combinations or situations that do not conform to normal user device behavior, such as user agent strings and operating system versions not matching, screen resolution not matching browser viewport size, etc.; then, score the device fingerprint entries according to the set anomaly judgment rules, and set thresholds to screen out device configurations whose feature combinations have obvious signs of disguise; finally, summarize the device configurations that meet the disguise characteristics to form a disguised device configuration set.
[0089] S4.3. Establish an initial proxy IP pool through the proxy IP service, perform speed and anonymity tests on each IP address in the proxy IP pool, calculate the availability score, and obtain the quality assessment results of the proxy IP.
[0090] Specifically, the expression is,
[0091]
[0092] Among them, S is the availability score of the proxy IP, α is the connection speed weight measured in the proxy IP, d is the connection speed measured in the proxy IP, Maxd is the maximum connection speed measured in the proxy IP, β is the script test weight for detecting the anonymity of the proxy, and A is the script test for detecting the anonymity of the proxy.
[0093] S4.4. Based on the quality assessment results of the proxy IP, set the HTTP header information to simulate access from specific devices and geographic locations, simulating the real access environment.
[0094] Furthermore, based on the quality assessment results of the proxy IP, HTTP header information is set to simulate access from specific devices and geographic locations. The specific process of simulating the real access environment is as follows: first, based on the availability score of each IP address in the proxy IP pool, high-quality proxy IPs with fast connection speed and high anonymity are screened out. The availability score is jointly determined by the connection speed weight, maximum connection speed, anonymity test weight and script test results; then the selected proxy IP is combined with the device fingerprint in the disguised device configuration set, and corresponding geographic location information is assigned to each combination. The geographic location information is based on the proxy IP's place of origin or simulated through the browser geolocation API; then, based on the characteristics of the target device, HTTP request header information that conforms to the device behavior pattern is constructed, including fields such as user agent string, accept language, screen resolution, browser plug-in list, etc., and ensures that it is consistent with the selected device fingerprint; finally, the constructed HTTP request header is used in combination with the selected proxy IP to initiate a network request, so that the access behavior received by the platform presents the characteristics of real user access in multiple dimensions such as device characteristics, IP ownership and geographic location, thereby realizing the platform access environment.
[0095] S5. Capture the recommended content on the platform by performing sentiment analysis and popularity evaluation on the recommended content on the platform.
[0096] S5.1. Extract the data types that need to be captured from the recommended content on the platform and form a capture list.
[0097] Furthermore, based on the structural layout of the platform content display page, the information modules contained in the recommended content are identified, including but not limited to content title, description text, tag classification, release time, author information, interactive data (such as number of likes, number of comments, number of reposts) and multimedia resource links; then, based on the requirements of subsequent tasks, various information modules are prioritized, among which the key fields used for sentiment analysis and popularity evaluation are set as mandatory crawling items, such as comment text, release time, interactive data and content body; then, by parsing the HTML source code of the page or calling the platform's public interface, the identifier path of the corresponding data field is extracted, and the identifier path includes XPath expression or JSON field key name, which is used to guide the automated crawling tool to locate the target data location; finally, all data types to be crawled and their corresponding identifier paths are organized into a structured file to form a crawling list that can be used for subsequent content collection tasks.
[0098] S5.2. Calculate the popularity score of each piece of content based on the structured log data.
[0099] Specifically, the expression is,
[0100] H(c)=γL like (c)+δL comment (c)+∈L share (c);
[0101] Among them, H(c) is the popularity score of content c, L like (c) is the degree of liking for the content, L comment (c) is the user’s attention to the content, L share (c) is the breadth of content dissemination, γ is the weight of the heat score, δ is the attention weight of the content, and ∈ is the weight of the breadth of content dissemination.
[0102] S5.3. Based on the crawled list, the popularity score of each content is sorted in descending order to form a popularity list.
[0103] Furthermore, according to the data type defined in the crawl list, the heat-related fields corresponding to each content are extracted from the platform's recommended content, including original heat indicators such as likeability, user attention, and content dissemination breadth; then the above indicators are weighted and summarized according to the preset heat calculation expression to obtain the heat score of each content, where likeability is obtained by adding the number of likes and the number of collections, user attention is measured by the frequency of fan interaction, and content dissemination breadth is reflected by the number of reposts and the number of shares; then the heat scores of all content are sorted according to the numerical size, and the descending order rule is applied to make the content with higher heat scores be ranked at the top of the list; finally, the sorting results are output in a structured form to generate a heat list containing content identifiers, heat scores and ranking numbers.
[0104] Clean the structured log data to form a vocabulary. Use a deep learning model to perform sentiment classification on the comments in the vocabulary, identify positive, negative, and neutral sentiment tendencies, and calculate the sentiment tendency score.
[0105] Specifically, the expression is,
[0106] E=w pos ·p pos -w neg ·p neg -w neu ·p neu ;
[0107] Among them, E is the sentiment tendency score, w pos is the positive sentiment weight, p pos For positive emotions, w neg is the negative sentiment weight, p neg For negative emotions, wneu is the neutral sentiment weight, p neu Neutral sentiment.
[0108] S5.4. Combine the sentiment tendency score and the popularity list to form recommended content on the platform.
[0109] Furthermore, the positive, negative and neutral sentiment tendency scores of each content calculated by the sentiment analysis model are normalized to make the sentiment tendency scores comparable on a unified scale; then the normalized sentiment tendency scores are weightedly fused with the popularity scores of the corresponding content in the popularity list, and the weighted fusion adopts a linear combination method, in which the sentiment tendency weight is used to reflect the influence of user emotions on the content dissemination, and the popularity score weight is used to reflect the contribution of content exposure; then the content is re-sorted according to the comprehensive score after weighted fusion to generate a new content sorting result that comprehensively considers user emotional feedback and dissemination popularity; finally, the sorting result is output in a structured form as recommended content on the platform.
[0110] S6. Based on the recommended content on the platform, the risk level of the platform’s video content is assessed in real time and an early warning report is generated.
[0111] S6.1. Calculate the negative sentiment ratio for each piece of content based on the sentiment tendency score.
[0112] Specifically, the expression is,
[0113]
[0114] Among them, N neg (c) is the negative sentiment ratio of each content c, c j is the comment of each content c, n is the total number of comments, j is the comment index, E(c j ) is the sentiment score of the jth comment, T neg The threshold for negative emotion judgment.
[0115] S6.2. Set a heat threshold based on the heat list to judge abnormally high heat content.
[0116] Furthermore, the heat score distribution of all the content in the heat list is statistically analyzed, the mean and standard deviation of the entire data set are sorted out, and the benchmark parameters for identifying abnormal heat content are determined in combination with historical data; then a statistical method is used to set a heat threshold, which is determined by adding several times the standard deviation to the average heat score, for example, it is set to the mean plus two times the standard deviation to distinguish between normal heat and content that exceeds the normal dissemination range; then the heat score of each content in the heat list is compared with the set heat threshold. If the heat score of a content exceeds the threshold, it is marked as a candidate for abnormally high heat content; finally, the content release time, growth rate and user interaction pattern are combined to further verify whether it belongs to sudden and fast-spreading content, thereby completing the identification and judgment of abnormally high heat content.
[0117] S6.3. Use the TF-IDF algorithm to extract keywords from recommended content and calculate the frequency of occurrence of sensitive words;
[0118] Specifically, the expression is,
[0119] F(v)=∑ b∈D I(v∈G);
[0120] Among them, F(v) is the frequency of occurrence of sensitive word v, G is the set of sensitive words, D is the total number of recommended contents, b is a piece of recommended content, v is the sensitive word index, and I is the indicator function.
[0121] S6.4. Based on the frequency of sensitive words, abnormally high-profile content, and sentiment ratio, a risk scoring model is constructed to calculate the risk score for each piece of content.
[0122] Specifically, the expression is,
[0123] R(c)=x1·F s +x2·Q abn +x3·N neg ;
[0124] Among them, R(c) is the risk score of each content, F s is the frequency of sensitive words, Q abn is abnormal heat, N neg is the negative emotion ratio, x1 is the sensitive word frequency coefficient, x2 is the abnormal heat coefficient, and x3 is the negative emotion ratio coefficient.
[0125] S6.5. Generate an early warning report based on the risk score of each content.
[0126] Furthermore, the three indicators of sensitive word frequency, abnormal heat index and negative emotion ratio of each content calculated by the risk scoring model are normalized to ensure that the indicators are comparable on a unified scale; then the normalized indicators are weighted and summarized according to the preset risk scoring expression to obtain the final risk scoring result. The sensitive word frequency coefficient in the weighted summary is used to reflect the impact of content compliance, the abnormal heat coefficient reflects the risk of dissemination and diffusion, and the negative emotion ratio coefficient measures the deviation of public opinion guidance. Then, according to the set risk level division threshold, the risk score is divided into different intervals, corresponding to the three warning levels of low risk, medium risk and high risk. Finally, the key information such as the risk score of each content, the warning level, the sensitive words involved, the heat score and the emotional tendency score are organized into a structured document, and a warning report is generated according to the time sorting and risk level priority principles.
[0127] This embodiment also provides a penetration deduction system, including: a user portrait construction module, which collects multi-source user feature data and generates user portraits;
[0128] User behavior simulation module, based on user portraits, uses simulation control technology to simulate user operations on the platform;
[0129] The recommendation feedback guidance module adjusts the behavior path and guides the push of target content based on the user's actions on the platform. It also identifies the platform's recommendation logic by analyzing the content of the platform's recommendation flow.
[0130] The anti-crawler bypass module uses multi-device switching and proxy IP pool rotation technology to simulate the real access environment;
[0131] The content collection module captures recommended content on the platform by performing sentiment analysis and popularity evaluation on the recommended content on the platform;
[0132] The public opinion warning module evaluates the risk level of the platform's video content in real time based on the recommended content on the platform and generates a warning report.
[0133] This embodiment also provides a computer device suitable for the infiltration deduction method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the infiltration deduction method proposed in the above embodiment.
[0134] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.
[0135] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the infiltration deduction method proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, disk or optical disk.
[0136] In summary, the present invention simulates a real access environment by using user profiles, utilizing simulation control technology to simulate user operations on the platform, and employing multi-device switching and proxy IP pool rotation technology. By setting operation probabilities for interest categories in user profiles through simulation control technology and generating operation sequences similar to those of real users, high-precision modeling of user behavior is achieved, providing structured and representative behavioral data support for subsequent identification and recommendation logic. At the same time, through multi-device switching and proxy IP pool rotation technology, combined with a device fingerprint library and proxy IP availability assessment mechanism, a highly realistic access environment is constructed, effectively circumventing platform access restrictions and improving the authenticity and stability of data collection.
[0137] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A seepage deduction method, characterized by: include, Collect multi-source user feature data and generate user portraits; Based on user portraits, use simulation control technology to simulate user operations on the platform; Based on the user's actions on the platform, the behavior path is adjusted to guide the push of target content. By analyzing the content of the platform's recommendation flow, the platform's recommendation logic is identified; Adopt multi-device switching and proxy IP pool rotation technology to simulate the real access environment; Capture recommended content on the platform by performing sentiment analysis and popularity evaluation on the recommended content on the platform; Based on the recommended content on the platform, the risk level of the platform's video content is assessed in real time and an early warning report is generated.
2. The infiltration deduction method according to claim 1, wherein: Collect multi-source user feature data and generate user portraits. The following steps are included: Use API to obtain basic information about users’ gender, age, and region to form user interaction behaviors; Based on user interaction behavior, use the TF-IDF algorithm to analyze user interest preferences; Integrate user interaction behaviors and user interest preferences to build user portraits.
3. The method for deducing seepage according to claim 2, wherein: According to the user portrait, the simulation control technology is used to simulate the user's operation on the platform, including the following steps: Set operation probabilities for interest categories through user portraits to form a preliminary behavior simulation plan; The preliminary behavior simulation plan is processed using simulation control technology to simulate user interaction behavior and generate operation sequences similar to real users.
4. The method for deducing seepage according to claim 3, wherein: Based on the user's actions on the platform, adjust the behavior path to guide the push of target content, and identify the platform's recommendation logic by analyzing the content of the platform's recommendation flow. The following steps are included: Structured log data is generated through similar operation sequences of real users; Clean the structured log data and add labels according to the content type to form a cleaned feature dataset; Extract content feature vectors from the feature data set, use the decision tree algorithm to train the recommendation rule model based on the content feature vectors to identify the recommendation logic, and obtain a preliminary recommendation rule model; According to the recommendation rule model, the initial behavior path strategy of the simulated user is set, and the behavior path is dynamically adjusted through the Q-learning algorithm; Simulate user behavior operations based on the behavior path, analyze the content of the platform recommendation flow and identify the platform's recommendation logic.
5. The method for deducing seepage according to claim 4, wherein: Use multi-device switching and proxy IP pool rotation technology to simulate the real access environment, including the following steps: Collect device features from the internal database and build a device fingerprint library; Based on the anti-crawler mechanism, the disguised device configuration set is filtered out from the device fingerprint library; Through the proxy IP service, an initial proxy IP pool is established, and each IP address in the proxy IP pool is tested for speed and anonymity, and the availability score is calculated to obtain the quality assessment results of the proxy IP; Based on the quality assessment results of the proxy IP, HTTP header information is set to simulate access from specific devices and geographic locations, simulating the real access environment.
6. The method for deducing seepage according to claim 5, wherein: By analyzing the sentiment and popularity of the recommended content on the platform, we can capture the recommended content on the platform. The following steps are included: Extract the data types that need to be crawled from the recommended content on the platform and form a crawling list; Calculate the popularity score of each content based on structured log data; Based on the crawled list, the popularity score of each content is sorted in descending order to form a popularity list; The structured log data is cleaned to form a vocabulary, and a deep learning model is used to perform sentiment classification on the comment texts in the vocabulary, identifying positive, negative, and neutral sentiment tendencies and calculating the sentiment tendency score.
7. The method for deducing seepage according to claim 6, wherein: Based on the recommended content on the platform, the risk level of the platform's video content is assessed in real time and an early warning report is generated. The following steps are included: Calculate the negative sentiment ratio of each piece of content based on the sentiment tendency score; Set a heat threshold based on the heat list to judge abnormally high heat content; Use the TF-IDF algorithm to extract keywords from recommended content and calculate the frequency of occurrence of sensitive words; Based on the frequency of sensitive words, unusually high-profile content, and sentiment ratio, a risk scoring model is constructed to calculate the risk score for each piece of content. An early warning report is generated based on the risk score of each content.
8. A seepage deduction system based on the seepage deduction method according to any one of claims 1 to 7, characterized in that: include, User portrait construction module collects multi-source user feature data and generates user portraits; User behavior simulation module, based on user portraits, uses simulation control technology to simulate user operations on the platform; The recommendation feedback guidance module adjusts the behavior path and guides the push of target content based on the user's actions on the platform. It also identifies the platform's recommendation logic by analyzing the content of the platform's recommendation flow. The anti-crawler bypass module uses multi-device switching and proxy IP pool rotation technology to simulate the real access environment; The content collection module captures recommended content on the platform by performing sentiment analysis and popularity evaluation on the recommended content on the platform; The public opinion warning module evaluates the risk level of the platform's video content in real time based on the recommended content on the platform and generates a warning report.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the infiltration deduction method described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the infiltration deduction method described in any one of claims 1 to 7 are implemented.