Dynamic personal file optimization method based on continuous feedback
By combining online experimental optimization methods with multi-armed slot machines and ε-greedy strategies, and integrating random forest regression and Shapley value analysis, the display weights of file variants are dynamically adjusted, solving the problem of lack of dynamic adaptability in traditional file management and improving the display effect of user files and the intelligence of the system.
Patent Information
- Application Number
- CN202511084621.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional personal profile management models lack dynamic adaptability and personalized adjustment, and cannot dynamically adjust based on the behavioral feedback of different visitors, resulting in poor user display effects and the inability to fully reflect potential capabilities.
An online experimental optimization method combining a multi-armed slot machine mechanism and an ε-greedy strategy is adopted. The display weight of file variants is dynamically adjusted through user interaction feedback. Personalized optimization suggestions are generated by combining a random forest regression model and Shapley value analysis.
It enables dynamic optimization of user profiles, enhances the attractiveness and conversion efficiency of profile display, strengthens the system's intelligent adaptive capabilities, and improves the user experience.
Smart Images

Figure CN120910244A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of personal profile optimization technology, and more specifically, to a dynamic personal profile optimization method based on continuous feedback. Background Technology
[0002] With the development of information technology and intelligentization, personal file systems have been widely used in education, recruitment, career development, and social platforms. Traditional file management models are mostly based on static information, lacking dynamic adaptability and personalized adjustment mechanisms, and often fail to accurately reflect a user's strengths and characteristics in different scenarios. Existing technologies have the following shortcomings:
[0003] Currently, the profile information submitted by users on recruitment platforms is displayed in a fixed manner once submitted, making it difficult to dynamically adjust based on feedback from different visitors. This can easily lead to poor user presentation or the inability to fully reflect potential abilities, resulting in a lack of support for feedback-based strategy evolution and a lack of real-time adjustment methods. Therefore, this paper proposes a dynamic personal profile optimization method based on continuous feedback.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a dynamic personal profile optimization method based on continuous feedback, which solves the problems mentioned in the background art by employing an online experimental optimization method that combines a multi-armed slot machine mechanism with an ε-greedy strategy.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a dynamic personal profile optimization method based on continuous feedback, comprising the following steps:
[0007] Step S1: The user uploads multiple candidate archive variant data through the archive element variant management module to construct an archive variant set for the archive element;
[0008] Step S2: The dynamic experiment and optimization module starts a multi-armed slot machine instance for the set of file variants. Using the ε-greedy algorithm, when other users access the user's file, it selects the target file variant to display based on the performance evaluation data, or randomly selects other file variants to explore and display.
[0009] Step S3: The user interaction feedback tracking module collects the interaction feedback data of other users in real time after the user profile is displayed, calculates the display reward value of each profile variant, and updates it to the dynamic experimentation and optimization module;
[0010] Step S4: The dynamic experiment and optimization module dynamically adjusts the display weight of each archive variant according to the display reward value, realizing the continuous optimization of the allocation strategy of the archive variant;
[0011] Step S5: The personalized archive consultant module generates personalized optimization suggestions using feature importance analysis method based on long-term accumulated interactive feedback data and multi-armed bandit instance results.
[0012] In a preferred embodiment, in step S1, the user inputs the candidate content of different versions corresponding to the archive element into the archive element variant management module, i.e., the archive variant;
[0013] Perform integrity check on the uploaded archive variant;
[0014] If the archive variant file passes the integrity check, it is marked as valid and added to the variant set corresponding to the current user archive element.
[0015] In a preferred embodiment, in step S2, a multi-armed bandit instance is created for the current user archive based on the variant set, each archive variant in the variant set is regarded as an independent pull rod, and performance evaluation data is constructed for each pull rod, including display count value, cumulative reward value, and average reward value;
[0016] After the multi-armed bandit instance is started, an ε-greedy algorithm is used as the archive variant selection strategy to dynamically determine the target archive variant to be displayed when other users access the current user archive;
[0017] When the target archive variant is determined, the dynamic experiment and optimization module returns the archive variant to the user interface layer and completes the display, while updating the display count value of the corresponding archive variant, making its value increase by one.
[0018] In a preferred embodiment, in step S3, the user interactive feedback tracking module monitors the currently displayed archive variant throughout the process to collect the interactive feedback data of other users on the archive browsing interface after each user archive is accessed and the archive variant is displayed by other users;
[0019] Based on the interactive feedback data, different weight coefficients are assigned to the interactive feedback data according to a predefined weight configuration table, and the display reward value of the archive variant is calculated.
[0020] In a preferred embodiment, in step S4, the dynamic experiment and optimization module updates the cumulative reward value and the average reward value of all archive variants based on the single display reward value;
[0021] According to the average reward value, the archive variants are arranged in descending order to generate an archive variant priority list, and the display weight factor of each archive variant is calculated.
[0022] In a preferred embodiment, in step S4, the dynamic experiment and optimization module introduces a smoothing mechanism in the weight adjustment logic, superimposes a certain proportion of historical weights on the basis of the current display weight factor;
[0023] The dynamic experiment and optimization module continuously optimizes the distribution strategy of the archive variant based on the display weight factor adjustment mechanism.
[0024] In a preferred embodiment, in step S5, the personalized archive consultant module first calls the performance evaluation data and weight factor of the archive variant;
[0025] All archive variants are classified and managed according to the corresponding archive elements, and a feature vector of the archive variant is constructed;
[0026] Based on the feature vector, a random forest regression model is used as a prediction model to predict the expected reward value under any archive feature combination through a training function.
[0027] In a preferred embodiment, in step S5, the personalized archive consultant module uses Shapley value analysis to explain the expected reward value output by the random forest model, and calculates the Shapley value of each archive feature;
[0028] The Shapley values of all features are analyzed and summarized to generate a ranking list of archive feature importance;
[0029] Archive features with a Shapley value greater than or equal to the filtering threshold in the archive feature importance ranking list are defined as high-impact features;
[0030] According to the high-impact features, a suggestion text corresponding to the archive element is generated;
[0031] After the user adopts the suggestion and adjusts the archive, the personalized archive consultant module submits the new archive variant to the archive element variant management module and continues the dynamic experiment and optimization process.
[0032] Technical effects and advantages of the present application:
[0033] The application is a dynamic personal profile optimization method based on continuous feedback, which realizes dynamic experiment and optimization of user profile elements. The user uploads multiple candidate profile variants through the profile element variant management module. The system checks the format, size and repeatability of the variants to form an effective variant set. The dynamic experiment and optimization module establishes a multi-armed bandit model based on the set, dynamically selects and displays variants when the user accesses the profile using the epsilon-greedy algorithm, collects interactive feedback in real time, calculates the display reward value and updates the weight. The weight adjustment introduces a smoothing mechanism to avoid local optimization caused by early data fluctuations. As data accumulates, the system continuously optimizes the display strategy to improve the attractiveness of the user profile. At the same time, the personalized profile consultant module builds a random forest regression model based on long-term data and multi-armed bandit results, calculates the predicted reward value of the profile features, and uses Shapley value analysis to quantify the contribution of each feature, generating personalized optimization suggestions. The suggestions cover picture selection, text keywords, label optimization, etc. The user can adjust the profile according to the prompts to ensure continuous optimization and user experience improvement, achieve high matching degree between profile content and user interaction intention, and enhance the display conversion efficiency of the profile and the intelligent adaptive ability of the system. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 The application is a dynamic personal profile optimization method based on continuous feedback.
[0035] Figure 2 The application is a dynamic personal profile optimization method based on continuous feedback. DETAILED DESCRIPTION
[0036] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0037] The application is a dynamic personal profile optimization method based on continuous feedback, which realizes dynamic experiment and optimization of user profile elements. The user uploads multiple candidate profile variants through the profile element variant management module. The system checks the format, size and repeatability of the variants to form an effective variant set. The dynamic experiment and optimization module establishes a multi-armed bandit model based on the set, dynamically selects and displays variants when the user accesses the profile using the ε-greedy algorithm, collects interactive feedback in real time, calculates the display reward value and updates the weight. The weight adjustment introduces a smoothing mechanism to avoid local optimization caused by early data fluctuations. As data accumulates, the system continuously optimizes the display strategy to improve the attractiveness of the user profile. At the same time, the personalized profile consultant module builds a random forest regression model based on long-term data and multi-armed bandit results, calculates the predicted reward value based on profile characteristics, and uses Shapley value analysis to quantify the contribution of each feature, generating personalized optimization suggestions. Suggestions cover picture selection, text keywords, label optimization, etc. Users can adjust the profile according to the prompts to ensure continuous optimization and improve user experience.
[0038] Embodiment 1, a dynamic personal profile optimization method based on continuous feedback, as shown in Figure 1 The steps include:
[0039] Step S1: The user uploads multiple candidate profile variant data through the profile element variant management module to build a profile variant set of profile elements;
[0040] Step S2: The dynamic experiment and optimization module starts a multi-armed bandit instance for the profile variant set. Using the ε-greedy algorithm, when other users access the user's profile, the target profile variant is selected for display based on performance evaluation data, or other profile variants are randomly selected for exploration display;
[0041] Step S3: The user interaction feedback tracking module collects real-time interaction feedback data of other users after the user's profile is displayed, and calculates the display reward value of each profile variant, which is updated to the dynamic experiment and optimization module;
[0042] Step S4: The dynamic experiment and optimization module dynamically adjusts the display weight of each profile variant based on the display reward value to realize the continuous optimization of the profile variant distribution strategy;
[0043] Step S5: The personalized profile consultant module generates personalized optimization suggestions based on long-term cumulative interaction feedback data and multi-armed bandit instance results using feature importance analysis methods.
[0044] The specific implementation is as follows:
[0045] In step S1, the user inputs candidate profile variant data for optimization to the profile element variant management module, the candidate profile variant data refers to the candidate content provided by the user in multiple different versions for the profile element, for subsequent dynamic optimization processing, the profile element refers to the structured component in the user profile, including user avatar, introduction title, introduction text, personal tag, the profile element variant management module receives each profile variant uploaded by the user and associates it with the corresponding profile element.
[0046] Then, the profile element variant management module performs integrity check on the uploaded profile variant, including file format verification, size ratio detection and repeatability detection, the file format verification is used to judge whether the profile variant conforms to the system preset format standard, for example, the avatar picture requires JPEG or PNG format, and the introduction text requires UTF-8 encoded text file; the size ratio detection is used to ensure that the pixel width-height ratio of the picture type profile variant is within the preset range, for example, the width-height ratio of the avatar picture should be within 1:1±5%; the repeatability detection compares the hash value of the file with the hash value of the existing profile variant to confirm whether the newly uploaded profile variant is a repeated submission of the existing version.
[0047] If the profile variant file format meets the standard, the size ratio is within the allowed range, and no repetition with other profile variants is detected, the profile variant is marked as valid and added to the variant set corresponding to the current user profile element, the variant set is composed of multiple valid profile variants, which is used as the input data of the subsequent dynamic experiment and optimization module for multi-armed bandit problem modeling, the variant set provides a material library for dynamic experiment, ensuring that only valid profile variants that pass quality audit are used in the subsequent optimization process, thereby improving the effectiveness of experimental data and the reliability of optimization results.
[0048] It should be noted that the hash value refers to a fixed-length string generated by processing the input data through a hash function, which can uniquely identify the content characteristics of the input data, which is used for repeatability detection of the profile variant in this embodiment.
[0049] In step S2, the dynamic experiment and optimization module receives the variant set constructed and output by the profile element variant management module, creates a multi-armed bandit instance for the current user profile based on the variant set, regards each profile variant in the variant set as an independent pull lever, and constructs performance evaluation data for each pull lever, including a display count value, a cumulative reward value, an average reward value, a unique identifier, and a state flag, wherein the display count value is used to record the total number of times the profile variant is displayed, and the initial value is set to zero; the cumulative reward value is used to record the cumulative value of all user interaction feedback since the profile variant joined the multi-armed bandit instance, and the initial value is set to zero; the average reward value is used to represent the unit display effect of the current profile variant, which is calculated by dividing the cumulative reward value by the display count value; the unique identifier records the system-assigned unique number of the profile variant, which is used for subsequent indexing and updating operations; and the state flag indicates whether the current profile variant is in an active displayable state, and the initial value is marked as Active.
[0050] After the multi-armed bandit instance is started, the dynamic experiment and optimization module uses an ε-greedy algorithm as the profile variant selection strategy to dynamically decide the target profile variant to be displayed when other users access the current user profile. The exploration rate parameter ε in the ε-greedy algorithm is a preset value, and its value is not unique. In this embodiment, it is set to 0.2, indicating that there is a 20% probability of performing exploration operation and an 80% probability of performing exploitation operation each time the display decision is made. When other users request to access the current user profile, the dynamic experiment and optimization module first calls a random number generation function to generate a floating-point random number ranging from zero to one to determine the mode of the current profile variant selection behavior. If the generated random number is less than or equal to zero point eight, the exploitation mode is entered, in which all profile variants with the state flag Active in the profile variant set are scanned, their average reward values are extracted, and the profile variant with the largest average reward value is selected as the target profile variant for display after comparing all average reward values. If the generated random number is greater than zero point eight, the exploration mode is entered, in which a profile variant with the state flag Active is randomly selected from the profile variant set as the target profile variant for display to ensure that all profile variants have a certain exploration opportunity and prevent the optimal profile variant from being missed due to insufficient early data.
[0051] When the target profile variant is determined, the dynamic experiment and optimization module returns the profile variant to the user interface layer and completes the display, and at the same time updates the display count value of the corresponding profile variant, so that its value increases by one.
[0052] It should be noted that the multi-armed bandit refers to a decision optimization problem modeling method derived from the field of probability theory and machine learning, which simulates a slot machine with multiple pull levers, each of which returns a reward value subject to a respective unknown probability distribution after being pulled, and the goal is to maximize the total reward value through a limited number of pull operations; the ε-greedy algorithm is a heuristic strategy for solving the multi-armed bandit problem, which introduces an exploration rate parameter ε, ε is a real number with a value range of 0 to 1, at each decision, with a probability of 1−ε, the best option in the current historical data is selected, and with a probability of ε, other options are randomly selected for exploration, so as to prevent falling into local optimization while maximizing overall revenue; the random number generation function refers to a software algorithm for generating an approximate random sequence of numerical values.
[0053] In step S3, the user interaction feedback tracking module starts a real-time data collection process after each user profile is accessed and displayed by other users, and monitors the currently displayed profile variant throughout to collect the interaction feedback data of other users on the profile browsing interface, including page dwell time, like behavior, comment behavior, collection behavior, and message sending behavior. Among them, the data form of like behavior, comment behavior, collection behavior and message sending behavior is Boolean value 1 or 0, if the behavior is detected, it is recorded as 1, if the behavior is not detected, it is recorded as 0.
[0054] For interaction feedback data, the user interaction feedback tracking module assigns different weight coefficients to the interaction feedback data according to a predefined weight configuration table, which is used to calculate the single display reward value of the profile variant. In the weight configuration table, each type of interaction feedback data corresponds to a fixed weight parameter. For example, the weight parameter of like behavior is 0.4, the weight parameter of comment behavior is 0.3, the weight parameter of collection behavior is 0.2, and the weight parameter of message sending behavior is 0.1. The event of page dwell time exceeding the set threshold value is counted as an additional reward, and the threshold value is set to 10 seconds. When calculating the single display reward value of the currently displayed profile variant, first, the total number of interaction behaviors with Boolean value 1 in the interaction feedback data of the profile variant in this display is counted, then the total number of interaction behaviors with Boolean value 1 is multiplied by the corresponding weight coefficient, and all the product results are added to form the single display reward value of the profile variant.
[0055] In step S4, the dynamic experiment and optimization module updates the cumulative reward value and the average reward value of all archive variants after receiving the single-show reward value of the user interaction feedback tracking module, and dynamically adjusts the display weight of each archive variant based on the updated data to realize the continuous optimization and distribution strategy of the archive variant. The dynamic experiment and optimization module uses a continuous weight adjustment mechanism to analyze the relative performance of each archive variant in the current variant set, and preferentially allocates display opportunities to the archive variant with the highest reward value, thereby improving the optimization efficiency of the overall archive attractiveness.
[0056] Specifically, the dynamic experiment and optimization module first accumulates the single-show reward value of each archive variant in the archive variant set to the cumulative reward value, and recalculates the average reward value according to the updated cumulative reward value.
[0057] After completing data update, the dynamic experiment and optimization module performs descending arrangement according to the average reward value of each archive variant to generate an archive variant priority list, and dynamically calculates the display weight factor of each archive variant based on the priority list. The display weight factor is used to determine the probability of each archive variant being selected for display in subsequent access requests. The calculation of the display weight factor follows the principle of proportional distribution, that is, the weight factor of the archive variant is proportional to its relative ranking in the priority list. For each archive variant, the total number of archive variants is first subtracted from the ranking of the variant in the priority list, and then 1 is added to obtain the original weight value of the archive variant. Then, the original weight values of all archive variants are summed. Finally, the original weight value of the archive variant is divided by the sum to obtain the display weight factor of the archive variant. This normalization calculation ensures that the sum of the display weight factors of all archive variants is 1, and the archive variants with higher rankings obtain higher display weight factors, thereby being preferentially selected in subsequent access.
[0058] In order to prevent the problem of local optimal convergence caused by insufficient early data, the dynamic experiment and optimization module introduces a smoothing processing mechanism in the weight adjustment logic, which realizes gradual adjustment of the weight factor by superimposing a preset proportion of historical weight on the current display weight factor, thereby avoiding the dramatic fluctuation of the weight caused by single user interaction feedback.
[0059] With the accumulation of interactive feedback data, the dynamic experiment and optimization module continuously optimizes the allocation strategy of the archive variant based on the above weight adjustment mechanism. For example, three avatar variants P1, P2 and P3 uploaded by a user are equally displayed in the initial stage. With the increase of interactive feedback data, the user like rate, comment rate and collection rate of P3 are significantly higher than those of P1 and P2. Therefore, the display frequency of P3 is gradually increased, and its weight factor is increased to 0.75, that is, there is a 75% probability of selecting P3 for display when the user visits. The weights of P1 and P2 are reduced to 0.15 and 0.1 respectively, so as to realize the dynamic tilt of display resources to high-performance archive variants, and ensure the effectiveness of the final optimization result and the overall improvement of user experience.
[0060] It should be noted that the smoothing mechanism refers to a technical method for avoiding irrational interference of individual variants on the overall optimization due to data fluctuations by introducing global historical statistics when processing archive variants with small statistical sample size or insufficient initial data.
[0061] In step S5, the personalized archive consultant module comprehensively analyzes the user interactive feedback data accumulated during the long-term operation of the dynamic experiment and optimization module and the experimental results of the multi-armed bandit instance, mines the contribution of each archive variant and its related features to the overall archive performance, and generates data-driven personalized optimization suggestions accordingly.
[0062] The personalized archive consultant module first calls the performance evaluation data and weight factor of the archive variant, classifies and manages all archive variants according to the corresponding archive elements, extracts metadata and context environment data associated with each archive variant, constructs a feature vector, including file type, image features, text features, display period, and access user portrait information. Among them, the image features include color distribution and shooting scene label, the text features include introduction length and keyword distribution, and the access user portrait information includes age, region and interest label.
[0063] Based on the feature vector and the historical data accumulated during the long-term operation of the dynamic experiment and optimization module, a mapping relationship between archive features and performance is constructed. A random forest regression model is used as a prediction model to model the relationship between the feature vector and the performance evaluation data. The random forest model is composed of multiple decision trees, has nonlinear fitting capability and feature selection capability, and can predict the expected reward value under any archive feature combination through a training function.
[0064] To further enhance the interpretability of the model, the personalized profile consultant module integrates interpretable AI technology, using Shapley value analysis to explain the expected reward value output by the random forest model, quantifying the marginal contribution of each profile feature to the predicted reward value. For each profile feature, its Shapley value is calculated to assess its positive or negative impact on the reward value in different profile variants.
[0065] It should be noted that the random forest regression model is an ensemble learning model used to solve nonlinear, multi-feature input regression prediction problems; Shapley value analysis is a feature importance measurement method derived from cooperative game theory, used to explain the prediction results of machine learning models.
[0066] By summarizing and analyzing the Shapley values of all features, a profile feature importance ranking list is generated, and based on the preset filtering threshold, a high-impact feature set is filtered out, i.e., profile features with Shapley values greater than or equal to the filtering threshold are defined as high-impact features. Based on the high-impact features and their historical performance data, combined with contextual information and user behavior patterns, data-driven suggestion texts are automatically generated, which clearly indicate the optimization direction of specific profile elements, such as:
[0067] The picture optimization suggestion is "According to recent interaction data analysis, your 'avatar C' (outdoor scene shooting) has an average interaction rate 30% higher than other avatars in the last 7 days, suggesting setting it as the first picture and considering uploading more photos reflecting outdoor activities to enhance the overall profile appeal."
[0068] The text optimization suggestion is "Data analysis shows that profiles containing interest keywords (such as 'travel' and 'photography') have a 15% higher interaction rate than those without keywords, suggesting adding relevant keywords to your profile."
[0069] The label optimization suggestion is "Adding the 'pet lover' label can significantly improve page dwell time, suggesting you add such a description to your labels."
[0070] In addition, for profile features that perform poorly, corrective suggestions are also provided, such as "Your profile length exceeds 200 words, and data analysis shows that overly long profiles may reduce page reading completion rates, suggesting shortening to within 150 words."
[0071] It should be noted that the suggestions generated by the personalized profile consultant module are based on statistical analysis of historical data and prediction model output, aiming to assist user decision-making, and the final adoption of the suggestions is determined by the user, which does not affect the core dynamic experiment and optimization functions of the system.
[0072] Finally, it should be noted that the terms "first", "second", and the like, herein do not necessarily have any actual meaning such as they are used to distinguish one element from another, but do not necessarily require or imply any actual such relationship or order between or among the elements referred to.
[0073] Also, the use of "including," "comprising," "having" and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless otherwise specified, "or" means "and / or." Unless otherwise noted, the use of the singular includes the plural.
[0074] As used in this document, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. As used in this document, the term "or" as used herein, unless otherwise indicated, is intended to mean "and / or," i.e., the term "or" when used in a list of two or more items, means that any one of the items can be present or none of the items are present. Unless otherwise noted, the use of the singular includes the plural.
[0075] The various embodiments described in this specification are presented by way of example, and are not intended to limit the scope of the disclosure. Each embodiment described herein can be combined with other embodiments in any way deemed suitable by those skilled in the art. Similarly, the various embodiments described herein can be combined with other embodiments, as well as with other operations, components, parts, or steps, to form other embodiments also within the scope of the disclosure.
[0076] The above description of disclosed embodiments is not intended to be exhaustive or to be unduly limited by the details given. Instead, various modifications and adaptations to the various embodiments can occur to those skilled in the art without departing from the spirit or scope of the disclosure. Thus, it is intended that the scope of the application be defined by the claims appended hereto rather than by the specific disclosure presented above.
Claims
1. A method for dynamic personal profile optimization based on persistent feedback, characterized by: The method comprises the following steps: Step S1: the user uploads a plurality of candidate profile variant data through the profile element variant management module to construct a profile variant set of the profile element; Step S2: the dynamic experiment and optimization module starts a multi-armed bandit instance for the profile variant set, adopts an ε-greedy algorithm, and selects a target profile variant to be displayed according to performance evaluation data or randomly selects other profile variants to be displayed for exploration when other users access the user profile; Step S3: the user interaction feedback tracking module collects the interaction feedback data of other users after the user profile is displayed in real time, calculates the display reward value of each profile variant, and updates the dynamic experiment and optimization module; Step S4: the dynamic experiment and optimization module dynamically adjusts the display weight of each profile variant according to the display reward value to realize a continuous optimization allocation strategy of the profile variant; Step S5: the personalized profile consultant module generates a personalized optimization suggestion by using a feature importance analysis method based on long-term accumulated interaction feedback data and multi-armed bandit instance results.
2. The dynamic personal profile optimization method based on continuous feedback according to claim 1, characterized in that: In step S1, the user inputs different versions of candidate content corresponding to the profile element into the profile element variant management module, that is, the profile variant; Perform integrity check on the uploaded profile variant, including detecting whether the profile variant meets the preset format standard, whether the pixel width-height ratio is within the preset range, and comparing the hash value of the profile variant with that of the existing profile variant; If the profile variant file passes the integrity check, it is marked as valid and added to the variant set corresponding to the current user profile element.
3. The dynamic personal profile optimization method based on continuous feedback according to claim 2, characterized in that: In step S2, a multi-armed bandit instance is created for the current user profile based on the variant set, each profile variant in the variant set is regarded as an independent pull rod, and performance evaluation data is constructed for each pull rod, including display count value, cumulative reward value, and average reward value; After the multi-armed bandit instance is started, an ε-greedy algorithm is adopted as the profile variant selection strategy to dynamically determine the target profile variant to be displayed when other users access the current user profile; After the target profile variant is determined, the dynamic experiment and optimization module returns the profile variant to the user interface layer and completes the display, and updates the display count value of the corresponding profile variant so that the value increases by one.
4. The dynamic personal profile optimization method based on continuous feedback according to claim 1, characterized in that: In step S3, the user interaction feedback tracking module monitors the currently displayed profile variant to collect the interaction feedback data of other users on the profile browsing interface after the user profile is accessed and the profile variant is displayed; Based on the interaction feedback data, different weight coefficients are assigned to the interaction feedback data according to a predefined weight configuration table, and the display reward value of the profile variant is calculated.
5. The dynamic personal profile optimization method based on continuous feedback according to claim 4, characterized in that: In step S4, the dynamic experiment and optimization module updates the cumulative reward value and the average reward value of all profile variants based on the single presentation reward value; According to the average reward value, the profile variants are arranged in descending order to generate a profile variant priority list, and the presentation weight factor of each profile variant is calculated.
6. The dynamic personal profile optimization method based on continuous feedback according to claim 5, characterized in that: In step S4, the dynamic experiment and optimization module introduces a smoothing mechanism in the weight adjustment logic, which superimposes a certain proportion of historical weight on the current presentation weight factor; The dynamic experiment and optimization module continuously optimizes the allocation strategy of the profile variants based on the presentation weight factor adjustment mechanism.
7. The dynamic personal profile optimization method based on continuous feedback according to claim 6, characterized in that: In step S5, the personalized profile consultant module first calls the performance evaluation data and weight factor of the profile variants; All profile variants are classified and managed according to the corresponding profile elements, and a feature vector of the profile variants is constructed; Based on the feature vector, a random forest regression model is used as a prediction model to predict the expected reward value under any profile feature combination through a training function.
8. The dynamic personal profile optimization method based on continuous feedback according to claim 7, characterized in that: In step S5, the personalized profile consultant module uses Shapley value analysis to explain the expected reward value output by the random forest model, and calculates the Shapley value of each profile feature; All feature Shapley values are analyzed to generate a profile feature importance ranking list; The profile features with Shapley values greater than or equal to the filtering threshold in the profile feature importance ranking list are defined as high-impact features; According to the high-impact features, the corresponding suggestion text of the profile elements is generated; After the user adopts the suggestions and adjusts the profile, the personalized profile consultant module submits the new profile variant to the profile element variant management module and continues the dynamic experiment and optimization process.