Cache management method and device, electronic equipment and medium

By obtaining the multi-dimensional attributes of comments and videos, using classification models to predict the display probability of comments, and dynamically managing the storage of comments in the cache, the problem of low cache hit rate is solved and more efficient cache resource utilization is achieved.

CN120640023APending Publication Date: 2025-09-12BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510823782.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The existing cache management strategy relies on a fixed time expiration mechanism, which results in a large number of comments stored in the cache that do not need to be displayed in a short period of time, resulting in low cache hit rate and low resource utilization.

Method used

By obtaining the multi-dimensional attributes of comments and videos, using the pre-trained classification model to predict the classification probability of comments, and determining the dynamic expiration time based on the probability, the comment storage in the cache is dynamically managed.

Benefits of technology

It improves the cache hit rate, reduces excessive cache usage by comments that do not need to be displayed in a short period of time, and improves cache utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640023A_ABST
    Figure CN120640023A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a cache management method and device, electronic equipment and a medium, and relates to the technical field of data storage, and the technical scheme of the embodiment of the invention comprises the following steps: for a comment in a cache, obtaining a first multi-dimensional attribute of the comment and a second multi-dimensional attribute of a video to which the comment belongs, and based on the first multi-dimensional attribute and the second multi-dimensional attribute, obtaining a video to which the comment belongs; the classification probability of the comments is obtained through a pre-trained classification model, and the classification probability represents the probability that the time difference between the current moment and the next display moment of the comments is smaller than or equal to the preset time difference. And according to a preset mapping relationship between the classification probability and the storage duration, taking the storage duration mapped by the classification probability of the comments as the expiration duration of the comments. When the used storage space of the cache reaches a preset threshold value, based on the cached moments and the expiration durations of all the comments in the cache, the comments needing to be deleted in the cache are selected, and the selected comments are deleted from the cache. The utilization rate of cache resources is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data storage technology, and in particular to a cache management method, device, electronic device and medium. Background Art

[0002] During live or on-demand video playback, users can comment on the video content. Comments can be displayed in the comment area of ​​the video playback page, or they can be displayed above the video screen. When comments are displayed above the video screen, they are also called barrage comments. The internet contains a vast number of videos, each of which has a large number of comments. Storing all comments in the cache would require a huge amount of cache space. Therefore, comments are typically stored in the cache after they are generated or displayed. When the cache storage time reaches a fixed expiration time, the comments are deleted from the cache, thereby reducing the storage space required for the cache and improving cache resource utilization.

[0003] When a comment needs to be displayed, it is first read from the cache. If the comment is not read, it is read from a designated storage location, such as memory or hard disk. The read comment is then displayed and stored in the cache. Because current comment cache management strategies rely primarily on fixed expiration mechanisms, the cache contains a large number of comments that do not need to be displayed within a short period of time, resulting in a low cache hit rate, that is, a low success rate in reading comments from the cache, and low cache resource utilization. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a cache management method, apparatus, electronic device, and medium to improve the utilization of cache resources. The specific technical solution is as follows:

[0005] A first aspect of an embodiment of the present application provides a cache management method, the method comprising:

[0006] For a comment in the cache, obtaining a first multidimensional attribute of the comment and a second multidimensional attribute of the video to which the comment belongs, wherein the first multidimensional attribute includes multiple attributes that affect the frequency of display of the comment, and the second multidimensional attribute includes multiple attributes that affect the frequency of access to the video;

[0007] Based on the first multidimensional attribute and the second multidimensional attribute, a pre-trained classification model is used to obtain a classification probability of the comment, wherein the classification probability represents a probability that the time difference between the current moment and the next display moment of the comment is less than or equal to a preset time difference; wherein the classification model is trained based on multiple positive samples and negative samples, the positive samples are: the first multidimensional attribute and the second multidimensional attribute of the first category of comments, the first category of comments being comments for which the time difference between the moment of deletion from the cache and the first display moment after deletion is greater than the preset time difference, the negative samples are: the first multidimensional attribute and the second multidimensional attribute of the second category of comments, the second category of comments being comments for which the time difference between the moment of deletion from the cache and the first display moment after deletion is less than or equal to the preset time difference;

[0008] According to a preset mapping relationship between classification probability and storage duration, the storage duration mapped to the classification probability of the comment is used as the expiration duration of the comment;

[0009] When the used storage space of the cache reaches a preset threshold, based on the cache time and expiration time of each comment in the cache, the comments to be deleted in the cache are selected, and the selected comments are deleted from the cache.

[0010] Optionally, the step of mapping the classification probability of the comment to the storage duration as the expiration duration of the comment based on a preset mapping relationship between the classification probability and the storage duration includes:

[0011] Determine the classification probability interval to which the classification probability of the comment belongs from a plurality of preset classification probability intervals, and obtain a target interval;

[0012] Determine the target storage duration corresponding to the target interval based on a preset mapping relationship between each classification probability interval and the storage duration;

[0013] If the target storage duration includes a duration, the target storage duration is used as the expiration duration;

[0014] If the target storage duration includes multiple durations, one duration is selected from the target storage durations as the expiration duration.

[0015] Optionally, the method further includes:

[0016] Obtaining a current status every preset time period, wherein the current status is status information that affects the expiration time of each comment in the cache;

[0017] Input the current state into the reinforcement learning model, generate a random number within a preset numerical range through the reinforcement learning model, determine whether the random number is less than a preset exploration probability, and if so, randomly generate a parameter update result, the parameter update result including: data representing each classification probability interval after the update; if not, obtain multiple candidate parameters from the preset parameter selection space, each candidate parameter including a data representing each classification probability interval, and use the policy network included in the reinforcement learning model to predict the reward value obtained by updating each candidate parameter in the current state, and select the candidate parameter corresponding to the maximum reward value as the parameter update result in the current state;

[0018] According to the parameter update result, each classification probability interval is updated.

[0019] Optionally, the classification model is specifically obtained by training based on multiple positive samples and negative samples and the weights of each sample; each candidate parameter also includes a set of candidate weights, and the candidate weights represent the weights of each sample. The parameter update result also includes: the updated weights of each sample, and the updated weights are used to update the classification model.

[0020] Optionally, after updating each classification probability interval according to the parameter update result, the method further includes:

[0021] Obtaining state change information, the state change information including at least one of the following information: a change in the cache hit rate, a change in the load of the device where the cache is located, a proportion of the remaining storage space of the cache in the total storage space, a change in the access frequency of comments in the cache, and a change in the access latency after the update is performed according to the parameter update result;

[0022] Determining an immediate reward based on the state change information;

[0023] Based on the immediate reward, network parameters of the policy network are updated.

[0024] Optionally, after obtaining the parameter update result, the method further includes:

[0025] According to the sorting method included in the parameter update result, the sorting method used when selecting comments to be deleted from the cache is updated;

[0026] And / or, updating the size of the total storage space of the cache according to the size of the total storage space included in the parameter update result.

[0027] Optionally, obtaining the classification probability of the review based on the first multidimensional attribute and the second multidimensional attribute using a pre-trained classification model includes:

[0028] If the first multidimensional attribute includes the publishing time of the comment, determining the time difference between the publishing time and the current time, and normalizing the time difference as a publishing time feature value;

[0029] If the first multidimensional attribute includes the number of interactions between the comment and the user, normalizing the number of interactions to use as a characteristic value of comment activity;

[0030] If the first multidimensional attribute includes the text length of the comment, normalizing the text length to use it as a text length feature value;

[0031] If the first multidimensional attribute includes the playback time of the comment in the video, then determining the playback time period of the comment in the video based on the playback time, encoding the playback time period, and using the encoded value as the progress feature value;

[0032] If the first multidimensional attribute includes the content type of the comment, encoding the content type and using the encoded value as the type feature value;

[0033] If the first multidimensional attribute includes the user level of the user who posted the comment, normalizing the user level to use it as a level feature value;

[0034] If the first multidimensional attribute includes the frequency of historical comments posted by the user who posted the comment, normalizing the frequency of historical comments posted as a feature value of the frequency of posting;

[0035] If the first multidimensional attribute includes the display frequency of the comment, normalizing the display frequency to obtain a display frequency feature value;

[0036] If the second multi-dimensional attribute includes the number of times the video has been played, normalizing the number of times the video has been played to obtain a play count feature value;

[0037] If the second multi-dimensional attribute includes the number of interactions between the video and the user, normalizing the number of interactions to use as a video activity feature value;

[0038] If the second multi-dimensional attribute includes the playback type of the video, encoding the playback type and using the encoded value as the playback type feature value;

[0039] Each eigenvalue is constructed into a eigenvector, and the eigenvector is input into the classification model to obtain the classification probability output by the classification model.

[0040] According to a second aspect of an embodiment of the present application, a cache management device is provided, the device comprising:

[0041] an acquisition module, configured to acquire, for a comment in a cache, a first multidimensional attribute of the comment and a second multidimensional attribute of the video to which the comment belongs, wherein the first multidimensional attribute includes a plurality of attributes that affect a frequency at which the comment is displayed, and the second multidimensional attribute includes a plurality of attributes that affect a frequency at which the video is accessed;

[0042] A classification module is configured to obtain a classification probability of the comment based on the first and second multidimensional attributes obtained by the acquisition module and using a pre-trained classification model, wherein the classification probability represents a probability that the time difference between the current moment and the next display moment of the comment is less than or equal to a preset time difference; wherein the classification model is trained based on multiple positive samples and negative samples, the positive samples being: the first and second multidimensional attributes of first-category comments, the first-category comments being comments for which the time difference between the moment of deletion from the cache and the first display moment after deletion is greater than the preset time difference, and the negative samples being: the first and second multidimensional attributes of second-category comments, the second-category comments being comments for which the time difference between the moment of deletion from the cache and the first display moment after deletion is less than or equal to the preset time difference;

[0043] A mapping module, configured to use the storage duration mapped to the classification probability of the comment obtained by the classification module as the expiration duration of the comment based on a preset mapping relationship between the classification probability and the storage duration;

[0044] The management module is used to select comments to be deleted from the cache based on the cache time and expiration time of each comment in the cache when the used storage space of the cache reaches a preset threshold, and delete the selected comments from the cache.

[0045] Optionally, the mapping module is specifically configured to:

[0046] Determine the classification probability interval to which the classification probability of the comment belongs from a plurality of preset classification probability intervals, and obtain a target interval;

[0047] Determine the target storage duration corresponding to the target interval based on a preset mapping relationship between each classification probability interval and the storage duration;

[0048] If the target storage duration includes a duration, the target storage duration is used as the expiration duration;

[0049] If the target storage duration includes multiple durations, one duration is selected from the target storage durations as the expiration duration.

[0050] Optionally, the device further includes:

[0051] The acquisition module is further configured to acquire a current status every preset time period, wherein the current status is status information that affects the expiration time of each comment in the cache;

[0052] a generation module, configured to input the current state acquired by the acquisition module into a reinforcement learning model, generate a random number within a preset numerical range through the reinforcement learning model, determine whether the random number is less than a preset exploration probability, and if so, randomly generate a parameter update result, the parameter update result including: data representing each updated classification probability interval; if not, obtain multiple candidate parameters from a preset parameter selection space, each candidate parameter including a data representing each classification probability interval, and use the policy network included in the reinforcement learning model to predict a reward value obtained by updating each candidate parameter in the current state, and select the candidate parameter corresponding to the maximum reward value as the parameter update result in the current state;

[0053] An updating module is used to update each classification probability interval according to the parameter update result generated by the generating module.

[0054] Optionally, the classification model is specifically obtained by training based on multiple positive samples and negative samples and the weights of each sample; each candidate parameter also includes a set of candidate weights, and the candidate weights represent the weights of each sample. The parameter update result also includes: the updated weights of each sample, and the updated weights are used to update the classification model.

[0055] Optionally, the device further includes:

[0056] The acquisition module is further configured to acquire state change information after updating each classification probability interval according to the parameter update result, the state change information including at least one of the following information: a change in the hit rate of the cache, a change in the load of the device where the cache is located, a proportion of the remaining storage space of the cache in the total storage space, a change in the access frequency of comments in the cache, and a change in the access delay after the update according to the parameter update result;

[0057] a determination module, configured to determine an immediate reward based on the state change information;

[0058] The updating module is further configured to update the network parameters of the policy network based on the immediate reward.

[0059] Optionally, the update module is further configured to:

[0060] After obtaining the parameter update result, updating the sorting method used when selecting comments to be deleted from the cache according to the sorting method included in the parameter update result;

[0061] And / or, after obtaining the parameter update result, updating the size of the total storage space of the cache according to the size of the total storage space included in the parameter update result.

[0062] Optionally, the classification module is specifically used to:

[0063] If the first multidimensional attribute includes the publishing time of the comment, determining the time difference between the publishing time and the current time, and normalizing the time difference as a publishing time feature value;

[0064] If the first multidimensional attribute includes the number of interactions between the comment and the user, normalizing the number of interactions to use as a characteristic value of comment activity;

[0065] If the first multidimensional attribute includes the text length of the comment, normalizing the text length to use it as a text length feature value;

[0066] If the first multidimensional attribute includes the playback time of the comment in the video, then determining the playback time period of the comment in the video based on the playback time, encoding the playback time period, and using the encoded value as the progress feature value;

[0067] If the first multidimensional attribute includes the content type of the comment, encoding the content type and using the encoded value as the type feature value;

[0068] If the first multidimensional attribute includes the user level of the user who posted the comment, normalizing the user level to use it as a level feature value;

[0069] If the first multidimensional attribute includes the frequency of historical comments posted by the user who posted the comment, normalizing the frequency of historical comments posted as a feature value of the frequency of posting;

[0070] If the first multidimensional attribute includes the display frequency of the comment, normalizing the display frequency to obtain a display frequency feature value;

[0071] If the second multi-dimensional attribute includes the number of times the video has been played, normalizing the number of times the video has been played to obtain a play count feature value;

[0072] If the second multi-dimensional attribute includes the number of interactions between the video and the user, normalizing the number of interactions to use as a video activity feature value;

[0073] If the second multi-dimensional attribute includes the playback type of the video, encoding the playback type and using the encoded value as the playback type feature value;

[0074] Each eigenvalue is constructed into a eigenvector, and the eigenvector is input into the classification model to obtain the classification probability output by the classification model.

[0075] According to a third aspect of the present application, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus.

[0076] Memory for storing computer programs;

[0077] The processor is configured to implement the cache management method described in any one of the first aspects when executing a program stored in the memory.

[0078] In a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the cache management method described in any one of the first aspects is implemented.

[0079] In a fifth aspect of the present application, there is also provided a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the cache management method described in any one of the first aspects above.

[0080] The cache management method, device, electronic device, and medium provided in the embodiments of the present application can obtain a first multidimensional attribute of a comment in the cache and a second multidimensional attribute of the video to which the comment belongs, wherein the first multidimensional attribute includes multiple attributes that affect the frequency of display of the comment, and the second multidimensional attribute includes multiple attributes that affect the frequency of video access. In the embodiments of the present application, the classification model is trained based on the multidimensional attributes of the first and second categories of comments. Since the first category of comments are not displayed within a short period of time after being deleted, it indicates that the display frequency of the first category of comments is low, and the second category of comments are displayed within a short period of time after being deleted, it indicates that the display frequency of the second category of comments is high. Therefore, the classification model trained in this way can more accurately analyze the display frequency of the comment, thereby predicting the probability of the comment being displayed within a short period of time, that is, more accurately determining the classification probability of the comment. By determining the expiration time based on the classification probability of the comment, it can be achieved that when the probability of the comment being displayed within a short period of time is higher, the expiration time is longer, and thus it is less likely to be deleted from the cache. Therefore, the embodiments of the present application can store more comments that are displayed more frequently in the cache, thereby reducing the excessive occupation of the cache by comments that do not need to be displayed within a short period of time, improving the cache hit rate, and thus improving the cache utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art.

[0082] Figure 1 A flowchart of the first cache management method provided in an embodiment of the present application;

[0083] Figure 2 A flowchart of a second cache management method provided in an embodiment of the present application;

[0084] Figure 3 A flowchart of a method for updating probability intervals of each classification provided in an embodiment of the present application;

[0085] Figure 4 A flowchart of a third cache management method provided in an embodiment of the present application;

[0086] Figure 5 An exemplary schematic diagram of a cache management process provided in an embodiment of the present application;

[0087] Figure 6 A schematic diagram of the structure of a cache management device provided in an embodiment of the present application;

[0088] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0089] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0090] In order to improve the utilization of cache resources, the embodiment of the present disclosure provides a cache management method, which is applied to electronic devices, such as servers, desktop computers, or laptop computers with data processing capabilities. Figure 1 As shown, the cache management method provided in the embodiment of the present application includes the following steps:

[0091] S101: For a comment in a cache, obtain a first multidimensional attribute of the comment and a second multidimensional attribute of the video to which the comment belongs.

[0092] It can be understood that the comments in the cache include: video evaluation information displayed in the comment display area of ​​the video playback page, and / or barrage displayed on the video screen.

[0093] Among them, the first multidimensional attribute includes multiple attributes that affect the frequency of comments being displayed, and the second multidimensional attribute includes multiple attributes that affect the frequency of video access. It can be understood that while the video is playing, the comments of the video can also be displayed. For example, if the user manually clicks the barrage opening button when watching a video, or if the barrage display is turned on by default when the video is played, the barrage can be displayed above the video screen. Therefore, the attributes that affect the frequency of video access will also affect the frequency of video barrage display. Therefore, the embodiment of the present application obtains the first multidimensional attribute and the second multidimensional attribute, so that the expiration time of the comments determined in this way is more accurate.

[0094] For ease of distinction, in the embodiments of this application, the electronic device that executes the cache management method is referred to as a management device, and the electronic device where the cache resides is referred to as a cache device. The management device and the cache device can be the same electronic device or different electronic devices, and this is not specifically limited in the embodiments of this application. The management device can obtain the first multidimensional attribute and the second multidimensional attribute by analyzing the logs of the cache device.

[0095] For example, the cache device can manage the expiration time of comments in its own cache by executing the cache management method provided in the embodiment of the present application, and perform cache management based on the expiration time.

[0096] For another example, in a distributed scenario, the management device may manage the expiration time of the comments cached in each cache device and perform cache management based on the expiration time.

[0097] S102: Based on the first multidimensional attribute and the second multidimensional attribute, using a pre-trained classification model, obtain the classification probability of the comment.

[0098] The classification probability represents the probability that the time difference between the current moment and the next time a comment is displayed is less than or equal to the preset time difference.

[0099] The classification model is obtained based on training of multiple positive samples and negative samples. The positive samples are: the first multidimensional attribute and the second multidimensional attribute of the first category of comments, and the first category of comments are comments whose time difference between the moment of deletion from the cache and the moment of first display after deletion is greater than the preset time difference. The negative samples are: the first multidimensional attribute and the second multidimensional attribute of the second category of comments, and the second category of comments are comments whose time difference between the moment of deletion from the cache and the moment of first display after deletion is less than or equal to the preset time difference. The preset time difference can be set according to the actual application scenario, for example, according to information such as the size of the storage space of the cache and the user's tolerance for comment display delays, etc., to set the preset time difference. Exemplarily, the preset time difference is 5 minutes.

[0100] The training process of the classification model includes: inputting each positive sample and each negative sample into the classification model respectively, and obtaining the classification probability output by the classification model for each input sample. The classification probability output by the classification model and the classification label corresponding to the sample are substituted into the preset loss function to calculate the loss value, wherein the classification label of the positive sample is 1 and the classification label of the negative sample is 0. The preset loss function can be cross entropy loss, 0-1 loss or absolute error loss, etc. Then, based on the gradient descent method, the model parameters of the classification model are adjusted using the loss value, and the step of inputting each positive sample and each negative sample into the classification model is returned until the classification model converges to obtain the trained classification model. Among them, the classification model can be updated periodically, or updated in real time through online learning. The updating process of the classification model is the same as the training process and will not be repeated here.

[0101] Optionally, the classification model may be a model with classification capabilities, such as Extreme Gradient Boosting (XGBoost), Gradient Boosting Decision Tree (GBDT), or a decision tree.

[0102] S103: Based on a preset mapping relationship between classification probability and storage duration, the storage duration mapped to the classification probability of the comment is used as the expiration duration of the comment.

[0103] It can be understood that the higher the classification probability, that is, the higher the probability that the time difference between the current moment and the next time the comment is displayed is less than or equal to the preset time difference, the higher the possibility that the comment will be displayed in a short time, and therefore it can correspond to a longer storage time; conversely, the lower the classification probability, that is, the lower the probability that the time difference between the current moment and the next time the comment is displayed is less than or equal to the preset time difference, the lower the possibility that the comment will be displayed in a short time, and therefore it can correspond to a shorter storage time.

[0104] Optionally, the embodiment of the present application can determine the expiration time of a new comment through S101 to S103 when a new comment is added to the cache; or, the expiration time of each comment in the cache can be re-determined through S101 to S103 at fixed intervals. Alternatively, the embodiment of the present application also supports determining the expiration time of cached comments in other situations. The embodiment of the present application does not specifically limit the triggering timing of S101 to S103.

[0105] S104: When the used storage space of the cache reaches a preset threshold, based on the cache time and expiration time of each comment in the cache, select the comments to be deleted from the cache, and delete the selected comments from the cache.

[0106] Optionally, when the used storage space of the cache reaches a preset threshold, expired comments in the cache may be identified, where an expired comment is one for which the time difference between the time it was cached and the current time exceeds the expiration period. All expired comments may then be deleted from the cache, or a preset number of expired comments may be deleted in ascending order of the number of likes received by the expired comments. Alternatively, other methods may be used to select the comments to be deleted, which are not specifically limited in this embodiment of the present application.

[0107] The preset threshold may be the product of the total storage space of the cache and a preset ratio. The preset ratio may be set according to actual application scenarios, for example, the preset ratio is 80%.

[0108] The cache management method provided in the embodiment of the present application can obtain a first multidimensional attribute of the comment and a second multidimensional attribute of the video to which the comment belongs, for comments in the cache, wherein the first multidimensional attribute includes multiple attributes that affect the frequency of comment display, and the second multidimensional attribute includes multiple attributes that affect the frequency of video access. In the embodiment of the present application, the classification model is trained based on the multidimensional attributes of the first category of comments and the multidimensional attributes of the second category of comments. Since the first category of comments are not displayed within a short period of time after being deleted, it indicates that the display frequency of the first category of comments is low, and the second category of comments are displayed within a short period of time after being deleted, it indicates that the display frequency of the second category of comments is high. Therefore, the classification model trained in this way can more accurately analyze the display frequency of comments, thereby predicting the probability of comments being displayed within a short period of time, that is, more accurately determining the classification probability of comments. By determining the expiration time based on the classification probability of comments, it can be achieved that when the probability of comments being displayed within a short period of time is higher, the expiration time is longer, and thus it is less likely to be deleted from the cache. Therefore, the embodiment of the present application can store more comments that are displayed more frequently in the cache, thereby reducing the excessive occupation of the cache by comments that do not need to be displayed within a short period of time, improving the cache hit rate, and thus improving the cache utilization rate.

[0109] The following is a detailed description of the cache management method provided in the embodiment of the present application:

[0110] The first attribute information in the above S101 includes at least two of the following attributes: the time of posting of the comment, text length, display frequency, content type, number of interactions between the comment and the user, playback time of the comment in the video, user level of the user who posted the comment, and historical comment posting frequency.

[0111] The comment publishing time includes: the system time when the user's device published the comment and / or the playback time of the video to which the comment belongs. For example, the comment publishing time includes: 2025.1.1 11:11:11 and 00:38:27. 2025.1.1 11:11:11 represents the system time when the mobile phone published the comment, and 00:38:27 represents the playback time of the video to which the comment belongs.

[0112] The text length indicates the number of characters or bytes included in the comment content.

[0113] The display frequency indicates the number of times a comment has been displayed within a preset time period before the current moment.

[0114] The content type indicates whether the comment content is text, a picture, or a symbol, or a combination of at least two of text, a picture, and a symbol. For example, the symbol may be an emoji.

[0115] The number of interactions between comments and users includes likes and / or dislikes. Likes represent the cumulative number of times a comment's corresponding "like" button is clicked, indicating user approval of the comment's content. Dislikes represent the cumulative number of times a comment's corresponding "dislike" button is clicked, indicating user disapproval of the comment's content. The number of interactions between comments and users can reflect a comment's popularity.

[0116] The comment's playback time in the video indicates the time period during which the comment was displayed. As you can see, the closer a comment's playback time is to the middle of the video, the more relevant the comment is to the video content. Therefore, higher-quality comments are likely to have a longer expiration period.

[0117] Among them, when the video to which the comment belongs is a non-live video, the length of the video to which the comment belongs is fixed. Therefore, the video playback time period to which the comment belongs can be obtained through the display time of the comment and the video length.

[0118] The management device may predetermine the number of time periods included in the video, and when obtaining the first attribute information, divide the video to which the comment belongs into the number of time periods, and determine the video playback time period to which the comment is displayed.

[0119] For example, a video is 90 minutes long and has three time segments, representing the beginning, middle, and end of the video. That is, the beginning segment is 0-30 minutes, the middle segment is 31-60 minutes, and the end segment is 61-90 minutes. Assume that the comment is displayed at the 26th minute of the video, and the video playback segment to which the comment is displayed is the beginning segment.

[0120] The user level of the comment poster indicates their level of activity on the video platform. For example, the user level is positively correlated with information such as the number of video views, cumulative viewing time, number of reposts, number of comments posted, and top-up amount on the platform that year.

[0121] The historical comment publishing frequency of the comment publishing user represents the number of comments published by the comment publishing user within a preset time period before the current moment.

[0122] The first attribute information may also include other information, which is not specifically limited in the embodiments of the present application. For example, the first attribute information also includes the display status of the comment. The display status includes: normal status, hidden status, and deleted status. The normal status indicates that the comment can be displayed normally when each user device plays the video to which the comment belongs. The hidden status indicates that the comment is only displayed when the client that posted the comment plays the video to which the comment belongs, and is hidden in other cases. The deleted status indicates that the comment is only displayed when the client that posted the comment plays the video to which the comment belongs, and is not displayed in other cases. The display status of the comment can be set according to the content of the barrage. For example, determine whether the comment contains sensitive words or whether it is negative; if any of the judgment results are yes, it is determined that the review has failed and the comment is determined to be deleted; if not, it is determined that the review has passed. For comments that have passed the review, if the quality of the comment content is lower than the preset quality, the comment is determined to be hidden; otherwise, the comment is determined to be normal. The expiration period of the comment decreases in the order of normal status, hidden status, and deleted status.

[0123] In an embodiment of the present application, the second attribute information in the above S101 includes at least two of the following attributes: the number of times the video is played, the number of times the video interacts with the user, and the playback type of the video.

[0124] The number of interactions between the video and users includes: the number of likes, dislikes, collections and / or reposts of the video.

[0125] The playback types of videos include: live broadcast, recorded broadcast, and on-demand. Among them, live broadcast means playing the currently captured video in real time; recorded broadcast means playing the completed video at a specified time; on-demand means playing the completed video based on the user's selection. For example, on-demand includes: playing movies, TV series, or live broadcast replays selected by the user. The expiration period decreases in the order of live broadcast, recorded broadcast, and on-demand. Since the life cycle and access mode of the comments included in videos of different playback types are different, the embodiment of the present application takes into account the playback type of the video and can set a more appropriate expiration period for the comments included in different types of videos, so that the videos to which the comments supported by the embodiment of the present application belong are more diversified, and the application scope of the embodiment of the present application is wider.

[0126] The second attribute information may also include other information, which is not specifically limited in the present embodiment. For example, the second attribute information may also include the release time of the video. The release time represents the time difference between the video release time and the current time.

[0127] On this basis, the method of obtaining the classification probability of the review based on the first multidimensional attribute and the second multidimensional attribute using a pre-trained classification model in S102 includes the following steps:

[0128] Step 1: If the first multidimensional attribute includes the publishing time of the comment, determine the time difference between the publishing time and the current time, and normalize the time difference as the publishing time feature value.

[0129] Step 2: If the first multidimensional attribute includes the number of interactions between the comment and the user, the number of interactions is normalized and used as the comment activity feature value.

[0130] Step 3: If the first multidimensional attribute includes the text length of the comment, normalize the text length as a text length feature value.

[0131] Step 4: If the first multidimensional attribute includes the playback time of the comment in the video, determine the playback time period of the comment in the video based on the playback time, encode the playback time period, and use the encoded value as the progress feature value.

[0132] The embodiments of the present application may use different encoding methods for different attributes, or the same encoding method may be used, which may be set according to the actual application scenario. For example, the encoding method includes: one-hot encoding, label encoding, or bag of words model.

[0133] Step 5: If the first multidimensional attribute includes the content type of the comment, encode the content type and use the encoded value as the type feature value.

[0134] Step 6: If the first multidimensional attribute includes the user level of the user who posted the comment, the user level is normalized and used as the level feature value.

[0135] Step 7: If the first multidimensional attribute includes the historical comment publishing frequency of the comment publishing user, the historical comment publishing frequency is normalized and used as the publishing frequency feature value.

[0136] Step 8: If the first multidimensional attribute includes the display frequency of the comments, the display frequency is normalized as the display frequency feature value.

[0137] Step 9: If the second multidimensional attribute includes the number of times the video is played, the number of times the video is played is normalized and used as a feature value of the amount of play.

[0138] Step 10: If the second multidimensional attribute includes the number of interactions between the video and the user, the number of interactions is normalized and used as a feature value of the video activity.

[0139] Step 11: If the second multi-dimensional attribute includes the playback type of the video, encode the playback type and use the encoded value as the playback type feature value.

[0140] Step 12: construct each eigenvalue into a eigenvector, input the eigenvector into the classification model, and obtain the classification probability output by the classification model.

[0141] The time a comment is published reflects how long it's been posted, thus indicating its freshness. The number of interactions between a comment and users reflects its popularity and how popular it is. The length of the text reflects the amount of effort users put into reading the comment, thus reflecting its attention. The duration of the comment within the video reflects its relevance to the video content, thus reflecting its quality. The type of content reflects its quality. The user level and frequency of historical comments posted by the comment-posting user reflect the quality of the user's comments. The frequency of comment display reflects its popularity. The number of video plays and the number of interactions between the video and users reflect its popularity and how popular it is. The playback type reflects the video playback scenario. Therefore, constructing a feature vector based on the first and second multidimensional attributes and using a classification model to output classification probabilities based on the feature vector can control the classification probability of comments that are fresher, more popular, more frequently displayed, and have higher content quality in different playback scenarios. Caching these comments better meets the real-time display requirements of comments in actual video playback scenarios. The availability of high-value comments is guaranteed, allowing them to be stored in the cache for longer periods of time, so that they can be prioritized at key points in their life cycle, such as being displayed first during video playback; and the pressure on the cache caused by low-value comments is reduced, thereby increasing the cache hit rate and improving the user's viewing experience.

[0142] In some embodiments of this application, see Figure 2 The above S103 uses the storage duration mapped to the classification probability of the comment as the expiration duration of the comment according to the preset mapping relationship between the classification probability and the storage duration, including the following steps:

[0143] S1031. Determine the classification probability interval to which the classification probability of the comment belongs from a plurality of preset classification probability intervals, and obtain a target interval.

[0144] Among them, there is no intersection between the classification probability intervals, and the union of the classification probability intervals is [0,1].

[0145] Each classification probability interval can be set according to actual business needs, and the embodiments of the present application do not make specific limitations on this.

[0146] For example, the classification probability intervals are: [0, 0.7] and (0.7, 1]. Assuming that the classification probability of the comment is 0.6, the target interval is [0, 0.7].

[0147] For another example, the classification probability intervals are: [0, 0.5], (0.5, 0.7], (0.7, 0.9] and (0.9, 1]. Assuming the classification probability of the comment is 0.6, the target interval is (0.5, 0.7].

[0148] S1032: Determine a target storage duration corresponding to the target interval based on a preset mapping relationship between each classification probability interval and the storage duration.

[0149] For example, the storage duration for the classification probability interval [0, 0.7] is 12 hours, and the storage duration for the classification probability interval (0.7, 1] is 0 hours.

[0150] For example, the storage duration for the classification probability interval [0, 0.5] is [0, 10 minutes], the storage duration for the classification probability interval (0.5, 0.7] is (30 minutes, 2 hours], the storage duration for the classification probability interval (0.7, 0.9] is (2 hours, 6 hours], and the storage duration for the classification probability interval (0.9, 1] is (6 hours, 24 hours].

[0151] S1033: Determine whether the target storage duration includes multiple durations. If the target storage duration includes one duration, execute S1034; if the target storage duration includes multiple durations, execute S1035.

[0152] If the target storage duration is a single value, it indicates that the target storage duration is a single duration. If the target storage duration is multiple values, or the target storage duration is a duration interval, for example, [30 minutes, 2 hours], it is determined that the target storage duration includes multiple durations.

[0153] S1034. Set the target storage duration as the expiration duration.

[0154] S1035: Select a duration from the target storage duration as the expiration duration.

[0155] Optionally, a duration can be randomly selected from the durations included in the target storage duration as the expiration duration. Alternatively, when the target storage duration is a duration interval, the duration at the corresponding position in the duration interval can be selected according to the position of the classification probability of the comment in the target interval as the expiration duration. For example, the classification probability of the comment is 0.8, the target interval is (0.7, 0.9], and the classification probability of the comment is in the middle of the target interval. Assuming that the target storage duration is (2 hours, 6 hours], the duration in the middle of (2 hours, 6 hours], that is, 4 hours, is selected as the expiration duration of the comment.

[0156] In traditional cache management systems, a fixed time-to-live (TTL) mechanism is used for comments in the cache, that is, the comments in the cache are all set to a fixed expiration time, such as an expiration time of 2 hours or 5 days, and when the cache time of each comment reaches the expiration time, it is deleted from the cache. This method does not consider whether the comment will be displayed in a short period of time, resulting in a large number of comments that will not be displayed in a short period of time occupying the cache for a long time, resulting in a low cache hit rate, high garbage collection (GC) pressure, and low cache management system performance. Among them, GC pressure refers to the pressure caused by garbage collection of the cache, and garbage collection is to delete comments from the cache.

[0157] The embodiment of the present application can dynamically determine the expiration time of a comment based on the classification probability of the comment, and can comprehensively consider various information such as the display frequency of the comment, the degree of user popularity, and the quality of the comment content, and more accurately and comprehensively estimate the probability of the comment being displayed in a short period of time, thereby increasing the retention of comments with a higher probability of being displayed in a short period of time in the cache, and reducing the situation where the cache is occupied for a long time by a large number of comments with a low probability of being displayed in a short period of time, thereby improving the cache hit rate and enhancing the performance of the cache management system. Moreover, the embodiment of the present application deletes some comments in the cache only when the used storage space of the cache reaches a preset threshold, rather than taking a deletion operation when the cache length of each comment exceeds the expiration time, thereby reducing the GC pressure.

[0158] In the embodiment of the present application, in order to more accurately determine the expiration time of the review, the agent in the management device can also be used to regularly optimize the probability intervals of each classification, see Figure 3 ,The optimization method includes the following steps.

[0159] S301. Obtain the current status every preset time period.

[0160] The current status is status information that affects the expiration time of each comment in the cache.

[0161] For example, the current state (State) includes cache state, comment state, and system state. The cache state includes at least one of the following information: cache hit rate, load on the cache device, remaining cache storage space, and the average expiration time of each comment in the cache. The cache hit rate is the ratio of the number of successful comment retrievals from the cache to the total number of comment searches from the cache within a preset time period before the current moment. The load on the cache device includes, among other things, central processing unit (CPU) utilization and memory utilization. The remaining cache storage space is the difference between the total cache storage space and the used storage space. The comment state includes the first and second multidimensional attributes of each comment in the cache. The system state includes at least one of the following information: the rate of new comments added to the cache, the rate of old comments deleted from the cache, and the display frequency of each type of comment in the cache. The rate of new comments added to the cache can be: the number of comments added to the cache within a specified time period before the current moment. The rate of old comments deleted from the cache can be: the number of comments deleted from the cache within a specified time period before the current moment. Comments can be classified according to content type, for example, comments can be classified into text, pictures or symbols, or comments can be classified according to the content type of the video to which they belong, for example, comments can be classified into education, technology and entertainment, or comments can be classified in other ways. The embodiments of the present application do not make specific limitations on this. On this basis, the display frequency of each type of comment in the cache within a specified time period before the current moment can be counted.

[0162] S302. Input the current state into the reinforcement learning model, generate a random number within a preset numerical range through the reinforcement learning model, and determine whether the random number is less than the preset exploration probability. If so, randomly generate a parameter update result. If not, obtain multiple candidate parameters from the preset parameter selection space, and use the policy network included in the reinforcement learning model to predict the reward value obtained by updating each candidate parameter in the current state, and select the candidate parameter corresponding to the maximum reward value as the parameter update result in the current state.

[0163] In the embodiments of the present application, the reinforcement learning model may be a Deep Q-Network (DQN) or a Policy Gradient Methods algorithm. The reinforcement learning model may be pre-initialized, including setting the learning rate, discount factor, and exploration probability of the reinforcement learning model. The exploration probability is denoted as ε.

[0164] After the current state is input into the reinforcement learning model, the reinforcement learning model can adopt an epsilon-greedy strategy to generate a random number within a preset numerical range and determine whether the random number is less than epsilon. For example, a random number is generated in the range of [0,1], and it is determined whether the random number is less than 0.7. If so, a parameter update result is randomly generated. The parameter update result includes: data representing the updated probability intervals of each classification. For example, the parameter update result includes multiple values, and every two adjacent values ​​form a classification probability interval. Exemplarily, the parameter update results include: 0, 0.2, 0.5 and 1, indicating that the updated classification probability intervals are: [0, 0.2], (0.2, 0.5] and (0.5, 1].

[0165] If not, the reinforcement learning model obtains multiple candidate parameters from the preset parameter selection space. Each candidate parameter includes a data representing the probability interval of each classification, and each candidate parameter is different. For example, candidate parameter 1 includes: 0, 0.1, 0.4 and 1, and candidate parameter 2 includes: 0, 0.3, 0.6 and 1. The policy network included in the reinforcement learning model is used to predict the reward value obtained by updating each candidate parameter in the current state. For example, DQN predicts the Q value corresponding to each candidate parameter in the current state by Q value estimation, and selects the candidate parameter corresponding to the maximum reward value. This candidate parameter represents the optimal action (Action) in the current state, so it is used as the parameter update result in the current state.

[0166] S303: Update each classification probability interval based on the parameter update result.

[0167] After the intelligent agent in the management device obtains the parameter update result, it can send the parameter update result to the environment execution module of the management device, so that the environment execution module executes the parameter update result, thereby updating each classification probability interval.

[0168] Through the above method, the embodiment of the present application can use the reinforcement learning model to adopt the ε-greedy strategy to determine the parameter update result, and realize the balance between exploration and utilization using the preset exploration probability. That is, when the random number is less than the preset exploration probability, the parameter update result is randomly generated, thereby exploring new strategies in the current state, preventing the decision of the reinforcement learning model from being limited to the parameter selection space, thereby ignoring potential better decisions. At the same time, when the random number is greater than or equal to the preset exploration probability, the policy network is used to predict the optimal decision in the current state, thereby realizing the screening of the optimal strategy from the existing strategies. It can be seen that the embodiment of the present application can continuously optimize the probability intervals of each classification using the reinforcement learning model, thereby improving the accuracy of determining the expiration time of the comments.

[0169] In some embodiments of the present application, the classification model for predicting the classification probability of comments in the above S102 is specifically obtained by training based on multiple positive samples and negative samples and the weights of each sample.

[0170] Specifically, each positive and negative sample is fed into the classification model. After obtaining the classification probability output by the classification model for each input sample, a preset loss function is used to determine the error between each sample's classification label and the classification probability output by the classification model for that sample. This error is then multiplied by the sample's weight to obtain the product. The loss value is then calculated based on the product calculated for each sample. The loss value is then used to adjust the model parameters of the classification model until the classification model converges.

[0171] In an embodiment of the present application, comments deleted from the cache during the previous cycle can be periodically determined to be positive or negative samples based on the time difference between the moment the comment was deleted and the moment it was first displayed after deletion. These samples are then added to a training set containing both positive and negative samples. The updated training set is then used to update the classification model. Updating the classification model can also be referred to as fine-tuning. The classification model updating method is the same as the training method and is described above. This description is omitted here.

[0172] The embodiment of the present application can also use a reinforcement learning model to optimize the weight of each sample. That is, each candidate parameter in the above S303 also includes a set of candidate weights, and each set of candidate weights is different. The candidate weights represent the weights of each sample. Accordingly, the parameter update result also includes: the updated weights of each sample, wherein the updated weights are used to update the classification model. That is, after obtaining the parameter update result in the above S302, the weights of each positive sample and each negative sample can also be updated based on the parameter update result, and then the classification model is updated based on each positive sample and its updated weight, as well as each negative sample and its updated weight.

[0173] Through the above method, the embodiment of the present application can update the weights of each positive sample and negative sample through the reinforcement learning model. For example, when the cache hit rate included in the current state is low and the load of the device where the cache is located is high, the weight of the negative sample is increased, so that the classification model can more accurately identify the negative sample, thereby reducing the comments stored in the cache with a low probability of being displayed in a short period of time, improving the cache hit rate, and reducing the load of the device where the cache is located. For another example, when the cache hit rate included in the current state is high and the load of the device where the cache is located is low, the proportion of positive samples in the samples added to the training set in this cycle is high. In order to reduce the classification model's excessive focus on positive samples, the weights of the samples added to the training set in this cycle can be set to 0, thereby reducing the impact of uneven sample distribution on the classification model's recognition accuracy. It can be seen that the embodiment of the present application can adjust the classification model's attention to each sample when updating by optimizing the weights of positive and negative samples, thereby more accurately predicting the expiration time of comments with the goal of optimizing the current state.

[0174] In some embodiments of the present application, after updating each classification probability interval based on the parameter update result in step S303, the agent in the management device may further update its decision strategy for the parameter update result based on the environmental state after the parameter update result is executed. The updating method includes the following steps:

[0175] Step 1: Get status change information.

[0176] Among them, the state change information includes at least one of the following information: after updating according to the parameter update result, the change in the cache hit rate, the change in the load of the device where the cache is located, the proportion of the remaining storage space of the cache in the total storage space, the change in the access frequency of comments in the cache, and the change in the access delay.

[0177] The change in cache hit rate is the difference between the cache hit rate before the update and the cache hit rate after the update.

[0178] The load change amount of the device where the cache is located is: the difference between the load change amount of the device where the cache is located before the update and the load change amount of the device where the cache is located after the update.

[0179] The change in the access frequency of comments in the cache is: the difference between the access frequency of comments in the cache before the update and the access frequency of comments in the cache after the update.

[0180] The change in access latency is the difference between the display latency of comments on the video before the update and the display latency of comments on the video after the update.

[0181] Step 2: Determine the immediate reward based on the status change information.

[0182] The state change information can be substituted into the preset reward function to obtain immediate rewards.

[0183] For example, the reward function includes: when the cache hit rate change is less than -5%, the reward value obtained is -1, when the cache hit rate change is in the range of [-5%, 5%], the reward value obtained is 0, when the cache hit rate change is in the range of (5%, 10%], the reward value obtained is 1, when the cache hit rate change is greater than 10%, the reward value obtained is 2. Also, when the load change of the cache device is less than -10%, the reward value obtained is 2, when the load change of the cache device is in the range of [-10%, -5%], the reward value obtained is 1, when the load change of the cache device is greater than -5%, the reward value obtained is 0. When the proportion of the remaining storage space of the cache in the total storage space is in the range of [30%, 70%] , the reward value obtained is 1. When the proportion of the remaining storage space in the cache to the total storage space is not in [30%, 70%], the reward value obtained is 0. When the change in the access frequency of comments in the cache is greater than 5%, the reward value obtained is 1. When the change in the access frequency of comments in the cache is in [-5%, 5%], the reward value obtained is 0. When the change in the access frequency of comments in the cache is less than -5%, the reward value obtained is -1. When the change in access delay is less than -10%, the reward value obtained is 1. When the change in access delay is in [-10%, 10%], the reward value obtained is 0. When the change in access delay is greater than 10%, the reward value obtained is -1. The reward values ​​obtained are then summed up as the immediate reward.

[0184] Step 3: Update the network parameters of the policy network based on the immediate reward.

[0185] The embodiment of the present application can select a policy network update method based on the actual application scenario. For example, when the reinforcement learning model is a DQN model, the Q value update formula is used to update the network parameters of the policy network.

[0186] Through this method, the intelligent agent in the management device can collect state change information based on the updated parameter update results, determine an immediate reward based on this state change information, and then use this information to update the network parameters of the policy network, so that the policy network can determine the parameter update results with the goal of higher immediate rewards. This shows that the intelligent agent can not only continuously update cache management strategies such as classification probability intervals and sample weights, but also continuously update its own selection strategy for parameter update results with the goal of higher immediate rewards, allowing it to continuously learn better decision-making methods, thereby continuously optimizing the cache management strategy.

[0187] In some embodiments of the present application, the reinforcement learning model may also update the sorting method used when selecting comments to be deleted from the cache.

[0188] For example, the sorting method includes sorting comments that have been in the cache for longer than the expiration time in descending order based on the time difference between the last time they were accessed and the current time. The strategy for selecting comments to be deleted based on this sorting method can be called the Least Recently Used (LRU) strategy.

[0189] For example, the sorting method includes sorting comments whose cache duration exceeds the expiration time in the cache by access frequency or cumulative number of accesses in the cache from smallest to largest. The strategy for selecting comments to be deleted based on this sorting method can be called the Least Frequently Used (LFU) strategy.

[0190] For example, the sorting method includes sorting comments that have been cached for longer than the expiration time by the number of likes or replies from low to high. The strategy of selecting comments to be deleted according to this sorting method can be called a popularity sorting strategy.

[0191] The above sorting methods are only examples provided in the embodiments of the present application, and the embodiments of the present application can also support other sorting methods.

[0192] Accordingly, the parameter update result may also include an updated sorting method. On this basis, after obtaining the parameter update result in S302, the management device may also update the sorting method used when selecting comments to be deleted from the cache based on the sorting method included in the parameter update result.

[0193] For example, when the cache hit rate is low and the load of the device where the cache is located is high, the sorting mode may be updated to a sorting mode corresponding to the LFU policy.

[0194] For another example, when the remaining storage space of the cache is small, the sorting method may be updated to a sorting method corresponding to the LRU strategy.

[0195] In some embodiments of the present application, the reinforcement learning model may also update the size of the total storage space of the cache. Accordingly, the parameter update result may also include the updated total storage space size. Based on this, after obtaining the parameter update result in S302 above, the management device may also update the total storage space size of the cache based on the total storage space size included in the parameter update result.

[0196] The parameter update result of the reinforcement learning model may also include other information, which is not specifically limited in the embodiments of the present application. For example, the reinforcement learning model may also update the cache retention ratio, where the cache retention ratio represents the ratio of the storage space occupied by the remaining comments to the total cache storage space after the cache elimination strategy is executed. Accordingly, the parameter update result may also include the updated cache retention ratio. On this basis, after obtaining the parameter update result in S302 above, the management device may also update the cache retention ratio based on the cache retention ratio included in the parameter update result.

[0197] For another example, the reinforcement learning model can also update the weight of the attribute of each dimension in the first multidimensional attribute and the second multidimensional attribute. Accordingly, the parameter update result can also include the weight of the attribute of each dimension after the update. On this basis, after obtaining the parameter update result in the above S302, the management device can also update the weight of the attribute of each dimension in the first multidimensional attribute and the second multidimensional attribute based on the weight of the attribute of each dimension included in the parameter update result. Among them, the weight of each dimensional attribute is used to multiply with the eigenvalue of the dimension, so as to construct a feature vector based on the product. By optimizing the weight of each dimensional attribute, the classification model can be adjusted to pay attention to the attributes of different dimensions when determining the classification probability, so that the classification model pays more attention to the dimensions that can improve the instant reward, so that when the cache management is performed based on the expiration time output by the classification model, a higher instant reward can be obtained.

[0198] The embodiment of the present application can utilize a reinforcement learning model to optimize multiple aspects such as the probability intervals of each classification, sample weights, the sorting method for selecting comments to be deleted from the cache, and the total cache storage space, thereby optimizing the cache management strategy from different angles, thereby optimizing the cache management more comprehensively, improving the cache hit rate, and reducing the load on the device where the cache is located.

[0199] See also Figure 4 , the overall process of the cache management method provided in the embodiment of the present application is described below.

[0200] S401: For the comments in the cache, obtain the first multi-dimensional attribute of the comments and the second multi-dimensional attribute of the video to which the comments belong. The specific implementation of S401 can refer to the above S101 and will not be repeated here.

[0201] S402: Convert each dimension of the first multi-dimensional attribute and the second multi-dimensional attribute into a feature value, and construct each feature value into a feature vector. The specific implementation of S402 can refer to the relevant description of S102 above, which will not be repeated here.

[0202] S403: Input the feature vector into XGBoost to obtain the classification probability output by XGBoost. Based on the preset mapping relationship between classification probability and storage duration, the storage duration mapped to the classification probability of the comment is used as the expiration duration of the comment. The specific implementation of S403 can be found in the relevant descriptions of S102 and S103 above and will not be repeated here.

[0203] S404: Determine whether to optimize XGBoost. If yes, execute S405; if no, execute S406.

[0204] Among them, XGBoost can be optimized periodically, that is, XGBoost is optimized at every specified interval.

[0205] Alternatively, the accuracy of the XGBoost output result can be evaluated, and when the accuracy is less than a preset accuracy threshold, it is determined that XGBoost is to be optimized, otherwise it is determined that XGBoost is not to be optimized. For example, the classification probability output by XGBoost can be obtained, and after setting the expiration time of the comments in the cache, whether the time difference between the moment when the comment is deleted from the cache and the moment when the comment is first displayed after being deleted is greater than the preset time difference, if so, the comment is determined to be a positive sample, otherwise the comment is determined to be a negative sample. Determine whether the proportion of negative samples in the positive samples and negative samples as a whole is greater than a preset proportion. If so, determine that the accuracy of the XGBoost output result is less than the preset accuracy threshold, otherwise determine that the accuracy of the XGBoost output result is greater than or equal to the preset accuracy threshold. Alternatively, the accuracy of the XGBoost output result can also be evaluated in other ways, and the embodiments of the present application do not make specific limitations on this.

[0206] S405. Update XGBoost.

[0207] S406: Obtain the current state every preset time interval, input the current state into the DQN, obtain the parameter update result under the current state output by the DQN, and update the parameters based on the parameter update result. The specific implementation of S406 can be found in the relevant description above and will not be repeated here.

[0208] S407: When the used storage space of the cache reaches a preset threshold, based on the cached time and expiration time of each comment in the cache, select the comments to be deleted from the cache and delete the selected comments from the cache. The specific implementation of S407 can be referred to the relevant description of S104 above and will not be repeated here.

[0209] S408: Obtain state change information, determine an immediate reward based on the state change information, and update the network parameters of the policy network included in the DQN based on the immediate reward. The specific implementation of S408 can be found in the description of steps 1 to 3 above and will not be repeated here.

[0210] It can be seen that the embodiment of the present application can dynamically determine the expiration time of comments based on the multi-dimensional attributes of the comments and the videos to which the comments belong, that is, it can analyze the probability of comments being displayed in a short period of time based on multiple aspects such as the popularity of the comments and videos and the quality of the content, and dynamically manage the comments in the cache based on this, so that more high-value comments with a higher probability of being displayed in a short period of time are stored in the cache, thereby improving the flexibility and accuracy of the cache management strategy, avoiding the static and rigid problems of traditional cache management strategies, and significantly improving the resource utilization, hit rate and user experience of the cache during the comment display process. Moreover, the embodiment of the present application can also use the reinforcement learning model to update the probability intervals of each classification and the weights of each sample based on the current state, thereby continuously optimizing the management strategy for comments in the cache, making the cache management strategy adaptive and able to be continuously optimized as business needs and state changes, thereby improving the efficiency and stability of cache management. In addition, the embodiment of the present application can also update the network parameters of the policy network included in the reinforcement learning model based on the updated state change information, so that the reinforcement learning model can make decisions on aspects such as the probability interval of each classification and the weight of each sample with the goal of higher cache hit rate, lower cache device load and more reasonable cache remaining storage space, so as to continuously optimize the cache hit rate and cache device load, etc., improve the overall system performance, and can alleviate the load pressure of the cache device in the high concurrency scenario of accessing comments, reduce the risk of system crash, and improve the stability and responsiveness of the system.

[0211] See also Figure 5 , the collaborative optimization process of the classification model and the reinforcement learning model in the embodiment of the present application is described in detail below.

[0212] First, historical information is collected, including the time when comments were deleted from the cache and the time when the comments were first displayed after being deleted. Based on this historical information, multiple positive and negative samples are determined, and initial weights corresponding to each positive and negative sample are obtained. The initial weights can be manually preset or fixed by default.

[0213] Then, the classification model is trained using the positive and negative samples and the weights corresponding to each sample. The classification model is then used to determine the classification probability of the comments based on the first and second multidimensional attributes of the cached comments. This is used to determine the expiration time of the comments, and the cached comments are then managed based on the expiration time.

[0214] The current state is obtained at preset intervals and fed into the reinforcement learning model. The model then outputs a parameter update for the current state, which includes the weights of each positive sample and each negative sample. Based on the parameter update, the weights of each positive sample and each negative sample are then updated.

[0215] The classification model is then trained using the positive samples, negative samples, and the updated weights of each sample.

[0216] It can be seen that the embodiment of the present application provides an intelligent, learnable and dynamically optimized cache management method, which can combine classification prediction and continuous optimization feedback to perform fine-grained control of the expiration period of comments, thereby continuously optimizing the management of comments in the cache and realizing the intelligence and adaptability of the cache management strategy.

[0217] Based on the same inventive concept, corresponding to the above method embodiment, the embodiment of the present application also provides a cache management device, such as Figure 6 As shown, the device includes: an acquisition module 601, a classification module 602, a mapping module 603 and a management module 604;

[0218] An acquisition module 601 is configured to acquire, for a comment in a cache, a first multidimensional attribute of the comment and a second multidimensional attribute of the video to which the comment belongs, wherein the first multidimensional attribute includes a plurality of attributes that affect the frequency with which the comment is displayed, and the second multidimensional attribute includes a plurality of attributes that affect the frequency with which the video is accessed;

[0219] Classification module 602 is used to obtain a classification probability of the comment based on the first and second multidimensional attributes obtained by acquisition module 601 and using a pre-trained classification model, where the classification probability represents the probability that the time difference between the current moment and the next time the comment is displayed is less than or equal to a preset time difference; wherein the classification model is trained based on multiple positive samples and negative samples, where the positive samples are: the first and second multidimensional attributes of first-category comments, where the first-category comments are comments where the time difference between the moment they are deleted from the cache and the moment they are first displayed after deletion is greater than the preset time difference, and the negative samples are: the first and second multidimensional attributes of second-category comments, where the time difference between the moment they are deleted from the cache and the moment they are first displayed after deletion is less than or equal to the preset time difference;

[0220] A mapping module 603 is configured to use the storage duration mapped to the classification probability of the comment obtained by the classification module 602 as the expiration duration of the comment based on a preset mapping relationship between the classification probability and the storage duration;

[0221] The management module 604 is used to select comments to be deleted from the cache based on the cache time and expiration time of each comment in the cache when the used storage space of the cache reaches a preset threshold, and delete the selected comments from the cache.

[0222] Optionally, the mapping module 603 is specifically configured to:

[0223] Determine the classification probability interval to which the classification probability of the comment belongs from a plurality of preset classification probability intervals, and obtain a target interval;

[0224] Determine the target storage duration corresponding to the target interval based on the preset mapping relationship between each classification probability interval and the storage duration;

[0225] If the target storage duration includes a duration, the target storage duration is used as the expiration duration;

[0226] If the target storage duration includes multiple durations, one duration is selected from the target storage duration as the expiration duration.

[0227] Optionally, the device may further include:

[0228] The acquisition module 601 is further configured to acquire the current status every preset time period, where the current status is status information that affects the expiration time of each comment in the cache;

[0229] A generation module is configured to input the current state obtained by the acquisition module 601 into the reinforcement learning model, generate a random number within a preset numerical range through the reinforcement learning model, determine whether the random number is less than a preset exploration probability, and if so, randomly generate a parameter update result, the parameter update result including data representing each updated classification probability interval; if not, obtain multiple candidate parameters from the preset parameter selection space, each candidate parameter including data representing each classification probability interval, and use the policy network included in the reinforcement learning model to predict the reward value obtained by updating each candidate parameter in the current state, and select the candidate parameter corresponding to the maximum reward value as the parameter update result in the current state;

[0230] The update module is used to update the results of the parameters generated by the generation module and update the probability intervals of each classification.

[0231] Optionally, the classification model is specifically trained based on multiple positive samples and negative samples and the weights of each sample; each candidate parameter also includes a set of candidate weights, which represent the weights of each sample, and the parameter update result also includes: the updated weights of each sample, and the updated weights are used to update the classification model.

[0232] Optionally, the device may further include:

[0233] The acquisition module 601 is further configured to acquire state change information after updating each classification probability interval based on the parameter update result, where the state change information includes at least one of the following information: a change in the cache hit rate, a change in the load of the device where the cache is located, a proportion of the remaining storage space in the cache to the total storage space, a change in the access frequency of comments in the cache, and a change in the access latency after the update based on the parameter update result;

[0234] A determination module is used to determine the immediate reward based on the state change information;

[0235] The update module is also used to update the network parameters of the policy network based on the immediate reward.

[0236] Optionally, the update module is also used to:

[0237] After obtaining the parameter update result, update the sorting method used when selecting comments to be deleted from the cache according to the sorting method included in the parameter update result;

[0238] And / or, after obtaining the parameter update result, the size of the total storage space of the cache is updated according to the size of the total storage space included in the parameter update result.

[0239] Optionally, the classification module 602 is specifically configured to:

[0240] If the first multidimensional attribute includes the posting time of the comment, then determining the time difference between the posting time and the current time, and normalizing the time difference as the posting time feature value;

[0241] If the first multidimensional attribute includes the number of interactions between comments and users, the number of interactions is normalized and used as the characteristic value of comment activity;

[0242] If the first multidimensional attribute includes the text length of the review, the text length is normalized as the text length feature value;

[0243] If the first multidimensional attribute includes the playback time of the comment in the video, then based on the playback time, determine the playback time period of the comment in the video, encode the playback time period, and use the encoded value as the progress feature value;

[0244] If the first multidimensional attribute includes the content type of the review, the content type is encoded and the encoded value is used as the type feature value;

[0245] If the first multidimensional attribute includes the user level of the user who posted the comment, the user level is normalized and used as the level feature value;

[0246] If the first multidimensional attribute includes the frequency of historical comments posted by the comment posting user, the frequency of historical comments posted is normalized and used as the frequency feature value;

[0247] If the first multidimensional attribute includes the display frequency of the review, the display frequency is normalized to serve as the display frequency feature value;

[0248] If the second multidimensional attribute includes the number of times the video has been played, the number of times the video has been played is normalized and used as the feature value of the amount of play;

[0249] If the second multidimensional attribute includes the number of interactions between the video and the user, the number of interactions is normalized and used as the video activity feature value;

[0250] If the second multi-dimensional attribute includes the playback type of the video, the playback type is encoded and the encoded value is used as the playback type feature value;

[0251] Each eigenvalue is constructed into a eigenvector, and the eigenvector is input into the classification model to obtain the classification probability output by the classification model.

[0252] The present application also provides an electronic device, such as Figure 7 As shown, it includes a processor 701, a communication interface 702, a memory 703 and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704.

[0253] Memory 703, for storing computer programs;

[0254] The processor 701 is configured to implement the method steps executed by the management device in the above method embodiment when executing the program stored in the memory 703 .

[0255] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.

[0256] The communication interface is used for communication between the above electronic device and other devices.

[0257] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.

[0258] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.

[0259] In another embodiment provided in the present application, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the cache management method described in any one of the above embodiments is implemented.

[0260] In another embodiment provided by the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute the cache management method described in any one of the above embodiments.

[0261] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0262] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0263] Each embodiment in this specification is described in a related manner. Similar portions between the embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so their description is relatively simple. For related portions, refer to the description of the method embodiments.

[0264] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the scope of protection of the present application.

Claims

1. A cache management method, characterized in that: The method comprises: For a comment in the cache, obtaining a first multidimensional attribute of the comment and a second multidimensional attribute of the video to which the comment belongs, wherein the first multidimensional attribute includes multiple attributes that affect the frequency of display of the comment, and the second multidimensional attribute includes multiple attributes that affect the frequency of access to the video; Based on the first multidimensional attribute and the second multidimensional attribute, a pre-trained classification model is used to obtain a classification probability of the comment, wherein the classification probability represents a probability that the time difference between the current moment and the next display moment of the comment is less than or equal to a preset time difference; wherein the classification model is trained based on multiple positive samples and negative samples, the positive samples are: the first multidimensional attribute and the second multidimensional attribute of the first category of comments, the first category of comments being comments for which the time difference between the moment of deletion from the cache and the first display moment after deletion is greater than the preset time difference, the negative samples are: the first multidimensional attribute and the second multidimensional attribute of the second category of comments, the second category of comments being comments for which the time difference between the moment of deletion from the cache and the first display moment after deletion is less than or equal to the preset time difference; According to a preset mapping relationship between classification probability and storage duration, the storage duration mapped to the classification probability of the comment is used as the expiration duration of the comment; When the used storage space of the cache reaches a preset threshold, based on the cache time and expiration time of each comment in the cache, the comments to be deleted in the cache are selected, and the selected comments are deleted from the cache.

2. The method according to claim 1, characterized in that The step of mapping the classification probability of the comment to the storage duration as the expiration duration of the comment based on a preset mapping relationship between the classification probability and the storage duration includes: Determine the classification probability interval to which the classification probability of the comment belongs from a plurality of preset classification probability intervals, and obtain a target interval; Determine the target storage duration corresponding to the target interval based on a preset mapping relationship between each classification probability interval and the storage duration; If the target storage duration includes a duration, the target storage duration is used as the expiration duration; If the target storage duration includes multiple durations, one duration is selected from the target storage durations as the expiration duration.

3. The method according to claim 2, characterized in that The method further comprises: Obtaining a current status every preset time period, wherein the current status is status information that affects the expiration time of each comment in the cache; Input the current state into the reinforcement learning model, generate a random number within a preset numerical range through the reinforcement learning model, determine whether the random number is less than a preset exploration probability, and if so, randomly generate a parameter update result, the parameter update result including: data representing each classification probability interval after the update; if not, obtain multiple candidate parameters from the preset parameter selection space, each candidate parameter including a data representing each classification probability interval, and use the policy network included in the reinforcement learning model to predict the reward value obtained by updating each candidate parameter in the current state, and select the candidate parameter corresponding to the maximum reward value as the parameter update result in the current state; According to the parameter update result, each classification probability interval is updated.

4. The method according to claim 3, characterized in that The classification model is specifically obtained based on multiple positive samples and negative samples and the weight training of each sample; each candidate parameter also includes a set of candidate weights, and the candidate weights represent the weights of each sample. The parameter update result also includes: the updated weights of each sample, and the updated weights are used to update the classification model.

5. The method according to claim 3 or 4, characterized in that After updating each classification probability interval according to the parameter update result, the method further includes: Obtaining state change information, the state change information including at least one of the following information: a change in the cache hit rate, a change in the load of the device where the cache is located, a proportion of the remaining storage space of the cache in the total storage space, a change in the access frequency of comments in the cache, and a change in the access latency after the update is performed according to the parameter update result; Determining an immediate reward based on the state change information; Based on the immediate reward, network parameters of the policy network are updated.

6. The method according to claim 5, characterized in that After obtaining the parameter update result, the method further includes: According to the sorting method included in the parameter update result, the sorting method used when selecting comments to be deleted from the cache is updated; And / or, updating the size of the total storage space of the cache according to the size of the total storage space included in the parameter update result.

7. The method according to claim 1, characterized in that The obtaining of the classification probability of the review based on the first multidimensional attribute and the second multidimensional attribute using a pre-trained classification model includes: If the first multidimensional attribute includes the publishing time of the comment, determining the time difference between the publishing time and the current time, and normalizing the time difference as a publishing time feature value; If the first multidimensional attribute includes the number of interactions between the comment and the user, normalizing the number of interactions to use as a characteristic value of comment activity; If the first multidimensional attribute includes the text length of the comment, normalizing the text length to use it as a text length feature value; If the first multidimensional attribute includes the playback time of the comment in the video, then determining the playback time period of the comment in the video based on the playback time, encoding the playback time period, and using the encoded value as the progress feature value; If the first multidimensional attribute includes the content type of the comment, encoding the content type and using the encoded value as the type feature value; If the first multidimensional attribute includes the user level of the user who posted the comment, normalizing the user level to use it as a level feature value; If the first multidimensional attribute includes the frequency of historical comments posted by the user who posted the comment, normalizing the frequency of historical comments posted as a feature value of the frequency of posting; If the first multidimensional attribute includes the display frequency of the comment, normalizing the display frequency to obtain a display frequency feature value; If the second multi-dimensional attribute includes the number of times the video has been played, normalizing the number of times the video has been played to obtain a play count feature value; If the second multi-dimensional attribute includes the number of interactions between the video and the user, normalizing the number of interactions to use as a video activity feature value; If the second multi-dimensional attribute includes the playback type of the video, encoding the playback type and using the encoded value as the playback type feature value; Each eigenvalue is constructed into a eigenvector, and the eigenvector is input into the classification model to obtain the classification probability output by the classification model.

8. A cache management device, characterized in that: The device comprises: an acquisition module, configured to acquire, for a comment in a cache, a first multidimensional attribute of the comment and a second multidimensional attribute of the video to which the comment belongs, wherein the first multidimensional attribute includes a plurality of attributes that affect a frequency at which the comment is displayed, and the second multidimensional attribute includes a plurality of attributes that affect a frequency at which the video is accessed; A classification module is configured to obtain a classification probability of the comment based on the first and second multidimensional attributes obtained by the acquisition module and using a pre-trained classification model, wherein the classification probability represents a probability that the time difference between the current moment and the next display moment of the comment is less than or equal to a preset time difference; wherein the classification model is trained based on multiple positive samples and negative samples, the positive samples being: the first and second multidimensional attributes of first-category comments, the first-category comments being comments for which the time difference between the moment of deletion from the cache and the first display moment after deletion is greater than the preset time difference, and the negative samples being: the first and second multidimensional attributes of second-category comments, the second-category comments being comments for which the time difference between the moment of deletion from the cache and the first display moment after deletion is less than or equal to the preset time difference; A mapping module, configured to use the storage duration mapped to the classification probability of the comment obtained by the classification module as the expiration duration of the comment based on a preset mapping relationship between the classification probability and the storage duration; The management module is used to select comments to be deleted from the cache based on the cache time and expiration time of each comment in the cache when the used storage space of the cache reaches a preset threshold, and delete the selected comments from the cache.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 7 when executing a program stored in a memory.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.