Video recommendation method, method and device for training deep learning model, and agent
Patent Information
- Application Number
- CN202610956843.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-08-18
AI Technical Summary
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.
Smart Images

Figure CN122594594A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of deep learning, intelligent recommendation, and intelligent search, especially to video recommendation methods, methods, devices, and intelligent agents for training deep learning models. Background Technology
[0002] With the rapid development of technology, users can browse video, web pages and other resources through smartphones and other terminal devices. Summary of the Invention
[0003] This disclosure provides a video recommendation method, a method for training deep learning models, an apparatus, and an intelligent agent.
[0004] According to one aspect of this disclosure, a video recommendation method is provided, comprising: receiving interactive behavior features of a target object and resource features related to candidate videos; detecting interactive behavior patterns of the target object based on the interactive behavior features and resource features to obtain browsing behavior distribution information, wherein the browsing behavior distribution information represents the probability distribution of the target object's browsing behavior for video content of at least one video segment in the candidate videos; determining a target video from the candidate videos based on the browsing behavior distribution information, and recommending the target video to the target object.
[0005] According to another aspect of this disclosure, a method for training a deep learning model is provided, comprising: receiving training samples, the training samples including sample interaction behavior features of sample objects, sample resource features related to sample videos, and the tag browsing duration of sample objects for sample videos; using the deep learning model to perform the following operations to determine the sample browsing duration: detecting the interaction behavior pattern of sample objects based on the sample interaction behavior features and sample resource features to obtain sample browsing behavior distribution information, the sample browsing behavior distribution information representing the behavior probability distribution of sample objects performing browsing behavior for video content of at least one video segment in the sample video; determining the sample browsing duration based on the sample browsing behavior distribution information; and training the deep learning model based on the sample browsing duration and the tag browsing duration to obtain the trained deep learning model.
[0006] According to another aspect of this disclosure, a video recommendation apparatus is provided, comprising: a first receiving module, configured to receive interactive behavior features of a target object and resource features related to candidate videos; a detection module, configured to detect the interactive behavior patterns of the target object based on the interactive behavior features and resource features, and obtain browsing behavior distribution information, wherein the browsing behavior distribution information represents the probability distribution of the target object's browsing behavior for video content of at least one video segment in the candidate videos; and a recommendation module, configured to determine a target video from the candidate videos based on the browsing behavior distribution information, and recommend the target video to the target object.
[0007] According to another aspect of this disclosure, an apparatus for training a deep learning model is provided, comprising: a second receiving module for receiving training samples, the training samples including sample interaction behavior features of sample objects, sample resource features related to sample videos, and tag browsing duration of sample objects for sample videos; a determining module for using the deep learning model to perform the following operations to determine the sample browsing duration: detecting the interaction behavior pattern of sample objects based on the sample interaction behavior features and sample resource features to obtain sample browsing behavior distribution information, the sample browsing behavior distribution information representing the probability distribution of the behavior of sample objects performing browsing behavior for video content of at least one video segment in the sample video; determining the sample browsing duration based on the sample browsing behavior distribution information; and a training module for training the deep learning model based on the sample browsing duration and tag browsing duration to obtain the trained deep learning model.
[0008] According to another aspect of this disclosure, an artificial intelligence agent is provided, comprising: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a large model based on the target task, and obtaining output information by calling the large model to execute the method provided in the embodiments of this disclosure; and an output module for outputting the output information obtained by the processing module.
[0009] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method provided according to an embodiment of this disclosure.
[0010] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform a method provided according to embodiments of this disclosure.
[0011] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to embodiments of this disclosure.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0014] Figure 1The illustration schematically depicts an exemplary system architecture to which video recommendation methods and apparatus can be applied according to embodiments of the present disclosure.
[0015] Figure 2 A flowchart illustrating a video recommendation method according to an embodiment of the present disclosure is shown schematically.
[0016] Figure 3 A schematic diagram of a deep learning model according to an embodiment of the present disclosure is shown.
[0017] Figure 4 A schematic diagram of a Gaussian distribution curve according to an embodiment of the present disclosure is shown.
[0018] Figure 5 A schematic diagram of an exponential distribution curve according to an embodiment of the present disclosure is shown.
[0019] Figure 6 The schematic diagram illustrates the principle of a distribution information detection network for a deep learning model according to an embodiment of the present disclosure.
[0020] Figure 7 A flowchart illustrating a method for training a deep learning model according to an embodiment of the present disclosure is shown schematically.
[0021] Figure 8 A block diagram of a video recommendation apparatus according to an embodiment of the present disclosure is shown schematically.
[0022] Figure 9 A block diagram of an apparatus for training a deep learning model according to an embodiment of the present disclosure is shown schematically.
[0023] Figure 10 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.
[0024] Figure 11 A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0025] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0026] In the technical solutions disclosed herein, the acquisition, storage, and application of any type of information, such as user personal information, comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals.
[0027] The inventors discovered that in the process of recommending resources such as videos to users, there is a problem with low video recommendation accuracy, making it difficult to accurately recommend video resources that users are satisfied with based on their preferences.
[0028] Embodiments of this disclosure provide a video recommendation method, a method for training a deep learning model, an apparatus, and an intelligent agent. The video recommendation method includes: receiving interactive behavior features of a target object and resource features related to candidate videos; detecting interactive behavior patterns of the target object based on the interactive behavior features and resource features to obtain browsing behavior distribution information, wherein the browsing behavior distribution information represents the probability distribution of the target object's browsing behavior for video content of at least one video segment in the candidate videos; determining a target video from the candidate videos based on the browsing behavior distribution information, and recommending the target video to the target object.
[0029] According to embodiments of this disclosure, by detecting the interaction behavior patterns of the target object based on interaction behavior features and resource features, the obtained browsing behavior distribution information can represent the probability distribution of the target object's browsing behavior for video content displayed in the video segment of the candidate video. This allows the browsing behavior distribution information to more accurately and finely represent the interaction methods of the target object when watching the candidate video. As a result, the target video can be determined from the candidate video based on the browsing behavior distribution information, enabling the target video to more accurately meet the diverse viewing needs of the target object for video content and improve the accuracy of target video recommendation.
[0030] Figure 1 The illustration schematically depicts an exemplary system architecture to which video recommendation methods and apparatus can be applied according to embodiments of the present disclosure.
[0031] It is important to note that Figure 1 The examples shown are merely illustrative of system architectures applicable to embodiments of this disclosure, intended to help those skilled in the art understand the technical content of this disclosure. They do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For instance, in another embodiment, an exemplary system architecture for applying the video recommendation method and apparatus may include a terminal device. However, the terminal device can implement the video recommendation method and apparatus provided by embodiments of this disclosure without interacting with a server.
[0032] like Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0033] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0034] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0035] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0036] It should be noted that the video recommendation method provided in this embodiment can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the video recommendation device provided in this embodiment can also be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0037] Alternatively, the video recommendation method provided in this embodiment can generally be executed by server 105. Correspondingly, the video recommendation device provided in this embodiment can generally be located in server 105. The video recommendation method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the video recommendation device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0038] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0039] Figure 2 A flowchart illustrating a video recommendation method according to an embodiment of the present disclosure is shown schematically.
[0040] like Figure 2 As shown, the resource recommendation method includes operations S210~S230.
[0041] In operation S210, the interactive behavior characteristics of the target object and the resource characteristics related to the candidate video are received.
[0042] In operation S220, the interaction behavior pattern of the target object is detected based on the interaction behavior characteristics and resource characteristics to obtain browsing behavior distribution information.
[0043] In operation S230, the target video is determined from the candidate videos based on the browsing behavior distribution information, and the target video is recommended to the target audience.
[0044] According to embodiments of this disclosure, interactive behavior features characterize the interactive behavior of a target object. For example, they can represent any type of interactive behavior of the target object, such as the duration of video content viewing, liking actions, or dragging the video playback progress bar.
[0045] Resource features represent information related to the candidate video, such as video content, video title, and video category. Alternatively, resource features may include information related to historical interactions with the candidate video, such as likes, comments, and shares. The embodiments of this disclosure do not limit the specific type of resource features, as long as they are relevant to the candidate video.
[0046] Browsing behavior distribution information represents the probability distribution of a target object's browsing behavior on video content within at least one video segment of a candidate video. For example, the browsing behavior distribution information represents the probability that a target object will browse video content within the first 10 seconds of a candidate video, or, for another example, the browsing behavior distribution information also represents the probability of browsing video content within a video segment between 5 and 20 minutes in the candidate video.
[0047] In some examples, the browsing behavior distribution information is a probability distribution curve corresponding to the video time period of the candidate video. By using the behavior probability corresponding to the video time period in the probability distribution curve, the probability of the target object browsing the video content of each video time period in the candidate video can be represented in a more granular way. In this way, the attention and preference of the target object for the candidate video can be determined based on the browsing behavior distribution information, thereby improving the recommendation accuracy of the target video.
[0048] In some embodiments, the interaction behavior patterns of a target object are detected based on interaction behavior features and resource features, including using a trained deep learning model to process the interaction behavior features and resource features, and outputting browsing behavior distribution information. Alternatively, the interaction behavior features and resource features can be processed based on a deep learning model with a large number of model parameters, such as a large model, to output browsing behavior distribution information.
[0049] Interactive behavior features can represent the interactive behaviors corresponding to multiple interaction time periods. Interactive time periods can be of any type, such as long-term periods, recent periods, or target periodic periods. Therefore, interactive behavior features represent the interactive behaviors performed by the target object at multiple time scales. By detecting the target object's interactive behavior patterns with candidate videos based on interactive behavior features and resource features, it is possible to analyze the target object's interactive behaviors at multiple interaction time periods with finer granularity. This allows for the determination of the target object's interactive behavior patterns with candidate videos, enabling browsing behavior distribution information to more accurately represent the probability distribution of browsing behavior for at least one video segment. This allows for fine-grained analysis of the target object's preference for video content at each video time period within the candidate videos. Ultimately, this ensures that the target videos determined based on browsing behavior distribution information can more accurately meet the target object's actual needs for video resources, improving the user experience.
[0050] In some embodiments, detecting the interaction behavior pattern of a target object based on interaction behavior features and resource features may include: fusing interaction behavior features and resource features to obtain target fusion features; performing parameter detection on the target fusion features to determine parameter information for the probability density function; and processing the parameter information according to the probability density function to determine browsing behavior distribution information.
[0051] For example, by fusing interaction behavior features and resource features based on an attention mechanism, target fusion features are obtained, which can more accurately represent the interaction patterns of the target object with the candidate video.
[0052] The probability density function may include the exponential distribution function, the exponential-log-normal mixture distribution, the Gamma-Gaussian mixture distribution function, the Weibull-Gaussian mixture distribution function, the truncated Gaussian mixture distribution function, etc. The embodiments of this disclosure do not limit the specific type of probability density function.
[0053] Parameter information, as variables in the probability density function, can be processed using this information to obtain browsing behavior distribution information representing the probability density distribution curve. Therefore, based on the probability density information corresponding to the video segments of candidate videos, as represented by the probability density distribution curve, the probability distribution of a target audience browsing one or more video segments can be determined. This allows for the use of probability density distribution curves to determine recommendation metrics such as the target audience's preference for candidate videos or the probability of performing browsing behaviors. This enables the precise selection of the most satisfactory target video from multiple candidate videos, improving the accuracy of video resource recommendations and enhancing the user experience for the target audience by representing the diverse interaction patterns of the target audience at a finer granular level through browsing behavior distribution information.
[0054] According to embodiments of this disclosure, interactive behavior features may include multiple features, which characterize the interactive behavior of the target object in multiple interactive time periods, and the multiple interactive time periods correspond to multiple time period types.
[0055] Interaction periods can include long-term interaction periods, short-term interaction periods, and interaction periods within the same phase. Long-term interaction periods correspond to long-term interaction behaviors that represent the interaction behaviors performed by the target object within the first historical period. For example, long-term interaction behaviors represent the interaction behaviors of the target object within one year.
[0056] The short-term interaction period corresponds to the short-term interaction behavior performed by the target object in the second historical period. For example, the long-term interaction behavior refers to the interaction behavior of the target object within one week. The first historical period is longer than the second historical period.
[0057] The same-stage interaction period refers to a time period that has the same stage attribute as the current time period of the target object. For example, if the current interaction time period of the target object is 9 pm, the same-stage interaction behavior corresponding to the same-stage interaction period refers to the interaction behavior performed by the target object in the target stage from 9 pm to 10 pm in the historical time period.
[0058] It should be noted that the information acquisition involved in the embodiments of this disclosure, including but not limited to interactive behaviors, interactive features, and resource features, is all obtained under the condition of obtaining authorization from relevant personnel and organizations. The embodiments of this disclosure perform necessary data processing measures such as encryption and de-identification on the acquired information to avoid information leakage.
[0059] In some examples, the following operations are used to construct the long-term interaction features, short-term interaction features, and same-stage interaction features corresponding to long-term interaction behaviors, short-term interaction behaviors, and same-stage interaction behaviors, thereby realizing the construction of multi-time-scale interaction behavior features.
[0060] Obtain the set H of historical interaction behaviors related to the target object. u , Each historical interaction behavior It includes various information related to interactive behaviors, such as behavior content identifiers, interaction behavior type tags, browsing duration, the inherent duration of the viewed video, and the execution time of the interaction behavior. Based on the historical set of interactive behaviors, the following three types of interaction behavior sequences are constructed as long-term interactive behaviors, short-term interactive behaviors, and interactive behaviors in the same stage.
[0061] Long-term interaction behavior sequence This is used to reflect the continuous and stable interactive interests and preferences of a target object over a period of time, spanning days, weeks, or even longer. This long-term interactive behavior sequence can be derived from the historical interactive behavior set H. u Extract the most recent L valid historical interaction behaviors, or extract historical interaction behaviors from the set of historical interaction behaviors according to a preset time window to obtain the long-term interaction behavior sequence. Where M > L > 1.
[0062] A sequence of interactive behaviors at the same stage This is used to characterize recurring, cyclical consumption needs of users under the same time period. For example, if the time period corresponding to the current login of the target user is τ, it can be determined from the historical interaction behavior set H. u The system filters historical interaction habits that meet the same time period conditions, such as historical interaction behaviors that all occur on Mondays or Sundays, to form a sequence of interaction behaviors in the same period. .
[0063] Here, `slot()` is the time slot mapping function for the same period. The time slots in this same period can be designed according to hour, morning / afternoon, weekday, rest day, holiday, or other time slot attribute rules.
[0064] Short-term interaction behavior sequence This represents the target object's immediate interests and short-term intentions within its current continuous consumption chain. The target object's historical interaction behavior set is divided into multiple subsets based on the time intervals between adjacent behaviors. From these subsets, one or more time intervals adjacent to the target object's current request time are identified as short-term interaction behavior sequences. Short-term interaction sequences can reflect the target user's rapidly switching focus on video topics, emotional preferences, or consumption intentions at the current request stage.
[0065] By extracting features from long-term interaction behavior sequences, short-term interaction behavior sequences, and same-stage interaction behavior sequences, multiple interaction behavior features are obtained, including long-term interaction behavior features, short-term interaction behavior features, and same-stage interaction behavior features.
[0066] In some examples, fusing interaction behavior features and resource features to obtain the target fused feature may include: processing interaction behavior features using an expert network corresponding to the time period type and outputting expert fused features; determining the expert weights corresponding to each of the multiple expert fused features based on multiple interaction behavior features and resource features; and fusing the multiple expert fused features based on the multiple expert weights to obtain the target fused feature.
[0067] In some embodiments, the expert network is constructed based on the attention network algorithm, which is used to fuse resource features with long-term interaction behavior features, short-term interaction behavior features and same-stage interaction behavior features, respectively, and output the expert fusion features corresponding to long-term interaction behavior, short-term interaction behavior and same-stage interaction behavior.
[0068] According to embodiments of this disclosure, expert weights represent the degree of influence of the target object's interactive behavior during the interaction period with the type on the browsing behavior pattern of candidate videos. Multiple interactive behavior features and resource features are processed using any type of network algorithm, such as attention network algorithms or convolutional neural network algorithms, to output expert weights corresponding to each of the multiple expert fusion features.
[0069] Figure 3 A schematic diagram of a deep learning model according to an embodiment of the present disclosure is shown.
[0070] like Figure 3 As shown, the various interactive behavior features are categorized into long-term interactive behavior features, short-term interactive behavior features, and same-period interactive behavior features. Long-term interactive behavior features represent the target object's stable interests across days and weeks; short-term interactive behavior features characterize the target object's immediate interests in the current time period's resource consumption chain; and same-period interactive behavior features represent the target object's periodic consumption needs within the same time period. Resource features include the target object's basic characteristics, candidate video content characteristics, and the target object's interaction scenario characteristics.
[0071] The input features of a deep learning model are represented as resource features, long-term interaction behavior features, short-term interaction behavior features, and same-stage interaction behavior features. These input features can be determined by feature extraction from the interaction behavior sequence using a trained encoder. For example, input data x is represented as... Where u represents the basic features of the target object, v represents the candidate video content features, and c represents the interaction scene features. This represents a long-term sequence of interactive behaviors. S represents the sequence of interactive behaviors in the same stage. R This represents a short-term sequence of interactive behaviors.
[0072] After obtaining the above-mentioned multi-timescale interaction behavior sequences, each interaction behavior sequence is encoded and represented by formula (1) to represent the long-term interaction behavior characteristics, short-term interaction behavior characteristics and same-stage interaction behavior characteristics at different time scales.
[0073] (1).
[0074] in, , , All of them are encoders built based on attention network algorithms. , , These three encoders can be different network layers with independent parameters, or they can be encoders that share some model parameters, so as to extract features of the differences in the interaction behavior patterns of the target object at different time scales and obtain long-term interaction behavior features h. L Short-term interactive behavior characteristics h R and interactive behavior characteristics h at the same stage p .
[0075] The long-term interaction behavior features h are input into the first expert network, and the first expert fusion features are output. ; short-term interaction behavior characteristics h R Input the second expert network, output the second expert fusion feature ; The interactive behavior characteristics h at the same stage p Input the third expert network to obtain the third expert fusion features. For example, the fusion characteristics of each expert can be represented by formula (2).
[0076] (2).
[0077] Among them, Expert L ( ) represents the first expert network, Expert P( ) represents the third expert network, Expert R ( ) indicates the second expert network.
[0078] Will , , The target object's basic features u, candidate video content features v, and interaction scene features c are input into the gating network, which outputs the expert weights corresponding to each expert network. , , The sum of the weights of each expert is 1. The processing of the gating network is based on formula (3).
[0079] (3).
[0080] in, and These are the model parameters of the gated network. ; used to normalize the gated output into a probabilistic weight distribution.
[0081] Using feature fusion networks based on expert weights , , Feature fusion is performed on multiple expert fusion features to output the target fusion feature.
[0082] For example, the processing procedure of the feature fusion network can be represented by formula (4).
[0083] (4).
[0084] Among them, h G Features are fused to target specific features.
[0085] The distribution information detection network processes target fusion features to output parameter information for the probability density function, and uses the probability density function to process the parameter information to output one or more browsing behavior distribution information.
[0086] By using a gating network to generate expert weights corresponding to different time periods of interaction, the expert weights can more accurately represent the contribution of any given time period interaction to the video content of different time periods in the candidate videos. For example, when the target audience consistently performs interactive behaviors to consume resource information at the same time each day, the expert weights output by the gating network can be... It can be adaptively increased. When the long-term interaction behavior of the target object fluctuates little and its interests are relatively stable, the expert weight corresponding to the long-term interaction behavior is increased. The weight value is higher. Therefore, the fusion process of multi-timescale interaction behavior features can be transformed into a dynamically changing interaction pattern preference learning process through deep learning models. This makes the target fusion features more consistent with the target object's interest preferences for candidate videos in real recommendation scenarios, and realizes the dynamic adjustment of the target object's interaction behavior pattern for candidate videos based on the target object's interaction behavior.
[0087] In some embodiments, each expert network can also be implemented based on a cross-attention algorithm, which queries long-term interaction behavior features, same-stage interaction behavior features, and short-term interaction behavior sequences based on resource features, respectively, and outputs multiple expert fusion features. For example, a first expert network is used to fuse resource features and long-term interaction behavior features to output a first expert fusion feature. A second expert network is used to fuse resource features and short-term interaction behavior features to output a second expert fusion feature. A third expert network is used to fuse resource features and same-stage interaction behavior features to output a third expert fusion feature.
[0088] Deep learning models extract interactive behavior features at different time scales by setting up expert networks corresponding to different time-segment types. They then use gating networks to fuse multiple interactive behavior features with resource features to determine multiple expert weights, thus assessing the influence of interactive behavior features across different time-segment types on the interactive behavior patterns of candidate videos. This allows for a hybrid expert architecture that dynamically generates expert weights, adaptively determining the contribution of interactive behavior features corresponding to different time-segment types to interactive behavior patterns. This makes the parameters used in the probability density function more closely reflect the interaction state between the target object and the candidate video. By explicitly generating resource features and interactive behavior features corresponding to multiple time scales during the feature fusion stage of the deep learning model, and using expert weights to represent the contribution ratio of each time-scale interactive behavior feature to the target object's preference, the deep learning model's understanding of the target fusion features becomes interpretable and controllable. This improves the accuracy and granularity of the deep learning model's intermediate decision-making process in recognizing diverse interactive behavior patterns of the target object towards candidate videos, thereby enhancing the detection accuracy of the target video.
[0089] In some embodiments, the expert network corresponding to the time period type includes multiple behavioral expert networks, each corresponding to an interaction behavior of multiple behavior types. Processing interaction behavior features using the expert network corresponding to the time period type may include: using the behavioral expert network corresponding to the behavior type to process the interaction behavior sub-features corresponding to the interaction behaviors of the behavior type, thereby obtaining expert fusion sub-features.
[0090] In this system, multiple expert weights are associated with expert fusion sub-features corresponding to various behavior types within the expert fusion features. For example, a gating network is used to process interaction behavior features and resource features corresponding to multiple time periods, outputting expert weights corresponding to multiple behavior expert networks. Then, a behavior expert network is used to process the interaction behavior sub-features corresponding to the behavior types, outputting expert fusion sub-features corresponding to the behavior types. Finally, the expert fusion sub-features corresponding to the various time periods are weighted and fused using multiple expert weights to obtain the target fusion feature.
[0091] Behavior types can include any type of interaction, such as liking, rewatching, and commenting. For example, behavior expert networks can be set up for long-term interaction behavior characteristics, specifically for liking and commenting. Or, for short-term interaction behavior characteristics, behavior expert networks can be set up to represent multiple interaction behavior types.
[0092] In some embodiments, the expert network is constructed based on a single or multiple feedforward network layers. The expert network can also be a lightweight network layer based on residual network layer connections. The gating network can be constructed based on a multilayer perceptron, attention-gated network algorithm, or other network algorithms used for output normalization weights. The number of expert networks can be arbitrary. The behavior network within the expert network performs feature fusion on the interaction behavior features corresponding to the time period type, and the expert weights fuse multiple expert fusion sub-features of different time period types. This improves the accuracy of the target fusion features in understanding the diverse interaction behaviors performed by the target object in multiple time period types, further enhancing the prediction accuracy of browsing behavior distribution information for the target object's interaction behaviors with video content during video periods, and improving the matching degree between the target video and the target object's needs.
[0093] In the embodiments of this disclosure, the browsing behavior distribution information is a probability density curve determined by processing parameter information using a probability density function. The probability density curve represents the probability density distribution corresponding to each moment within the overall playback duration of the candidate video. Therefore, the probability distribution of browsing behavior during video segments within the candidate video can be determined by analyzing the shape of the probability density curve.
[0094] In some embodiments, the probability density function includes a Gaussian distribution function, and the parameter information for the Gaussian distribution function includes a first reference parameter representing a reference time of the video segment, and a second reference parameter representing the degree of dispersion of the behavior probability. The first and second reference parameters can be, for example, the mean parameter μ detected by the deep learning model based on the target fusion features. k and standard deviation parameter σ k .
[0095] The browsing behavior distribution information includes Gaussian distribution curves, which represent the probability distribution of the target object's browsing behavior for video content within a video segment.
[0096] Figure 4 A schematic diagram of a Gaussian distribution curve according to an embodiment of the present disclosure is shown.
[0097] like Figure 4 As shown, the Gaussian distribution curve, which serves as information on browsing behavior distribution, represents the probability density of the target object's behavior in performing browsing interaction behavior on the video content corresponding to video time period ta1 in the candidate video. Figure 4 The horizontal axis of the Gaussian distribution curve shown represents the playback duration of the candidate video, and the vertical axis represents the probability density of the browsing behavior. Therefore, the Gaussian distribution curve can be used to determine the probability distribution of the target object's browsing behavior for the video content corresponding to video segment ta1 in the candidate video.
[0098] It should be understood that when both the first and second reference parameters include multiple parameters, there are multiple Gaussian distribution curves. The time periods corresponding to the peak regions of multiple Gaussian distribution curves can represent the probability distribution of browsing behavior performed in multiple different video time periods in the candidate video.
[0099] In some examples, the deep learning model processes the target fusion features by inputting multiple first reference parameters and multiple second reference parameters. The multiple first reference parameters are k mean parameters μ. k and k standard deviation parameters σ k Therefore, we can determine k Gaussian distribution curves representing the browsing behavior of the target object for video content corresponding to k different video time periods.
[0100] For example, based on formula (5), the mean parameter μ is processed using the Gaussian distribution function. k and k standard deviation parameters σ k .
[0101] (5)
[0102] Among them, f gauss ( ) represents the Gaussian distribution function.
[0103] By utilizing a deep learning model to determine multiple mean parameters and multiple standard deviation parameters for the Gaussian distribution function based on the target fusion features, the mean parameter μ can be optimized. k The center time of a video segment is used as the reference time. This is based on the standard deviation parameter σ. kThe degree of dispersion of the corresponding Gaussian curve represents the fluctuation of the duration of the video segment and the probability density of the behavior. Based on the Gaussian distribution curve, the probability of the target object's browsing behavior in multiple video segments of the candidate video can be identified more accurately. This allows the Gaussian distribution curve to accurately characterize the target object's diverse browsing behaviors such as exiting browsing, resuming browsing from a breakpoint, and dragging the progress bar. Multiple Gaussian distribution curves can precisely depict the diverse interaction patterns of the target object in the candidate video, accurately representing the multi-peak browsing pattern of the target object in multiple video segments of the candidate video. Therefore, the target video can be determined based on multiple Gaussian distribution curves, which can more accurately meet the actual needs of the target object.
[0104] In some embodiments, the probability density function further includes an exponential distribution function, and the parameter information includes a rate parameter λ for the exponential distribution function. The deep learning model can output the parameter value of the rate parameter by processing the target fusion features. By processing the parameter value of the rate parameter using the exponential distribution function, an exponential distribution curve is determined as information about the browsing behavior distribution.
[0105] For example, the exponential distribution curve is determined based on formula (6). It should be noted that t in formulas (5) and (6) can represent the values of multiple moments corresponding to the browsing duration, and the duration between the end of the video segment and the start of the candidate video can be represented as the browsing duration.
[0106] (6).
[0107] The browsing behavior distribution information includes an exponential distribution curve, which represents the probability distribution of a target user's behavior of exiting the browsing of candidate videos after browsing video content within a preset video time period. The preset video time period is the period after the candidate video begins playing; for example, the preset video time period is within 5 seconds, 10 seconds, etc., after the candidate video begins playing. It should be understood that the video time period of the candidate video includes the preset video time period.
[0108] Figure 5 A schematic diagram of an exponential distribution curve according to an embodiment of the present disclosure is shown.
[0109] like Figure 5 As shown, the indicator distribution curve represents the probability distribution of the target object's behavior of exiting browsing within a preset video time period ta2 after the candidate video starts playing.
[0110] By using a deep learning model to determine the rate parameter λ based on the target fusion features, and then processing the parameter λ according to the exponential distribution function, the exponential distribution curve can more accurately represent the probability of the target object performing low-dwelling behaviors such as fast swipe interaction or instant rewind in response to candidate videos. Thus, the target video with a lower probability of performing low-dwelling behaviors can be determined from multiple candidate videos based on the exponential distribution curve, thereby improving the target object's satisfaction with the target video.
[0111] In some embodiments, the browsing behavior distribution information includes exponential distribution curves and Gaussian distribution curves. This allows for a more accurate determination of target videos that meet the diverse interactive behavior needs of the target audience, based on the probability of the target audience's low-dwelling behavior on candidate videos and the probability distribution of browsing behavior across multiple video segments, thereby improving the accuracy of video recommendations.
[0112] It should be noted that deep learning models process target fusion features to output values for the first and second reference parameters, as well as the rate parameter of the exponential distribution function. This avoids the model illusion and data regression error caused by deep learning models regressing individual tokens or score labels. Furthermore, it uses the probability density function to process the parameter information represented by the numerical values to determine the target object's browsing time for candidate videos. This allows it to identify target videos with longer browsing times from multiple candidate videos, thereby improving the accuracy of video recommendations for the target object.
[0113] In some embodiments, determining a target video from candidate videos based on browsing behavior distribution information may include: fusing multiple browsing behavior distribution information according to their respective distribution weights to determine the browsing duration corresponding to the candidate video; and determining the target video from at least one candidate video based on the browsing duration.
[0114] In this embodiment, the distribution weights are determined by weight detection of the target fusion features derived from the fusion of interactive behavior features and resource features. For example, the target fusion features can be processed based on the distribution weight detection layer in a deep learning model, outputting the distribution weights corresponding to each of the multiple browsing behavior distribution information.
[0115] The distributed weight detection layer can be built based on the attention network algorithm, but it is not limited to this. It can also be built based on other types of neural network algorithms such as convolutional neural networks and multilayer perceptrons.
[0116] By performing distribution weight detection on target fusion features, multiple distribution weights can be learned from multi-time-type interaction behavior features to more accurately represent the contribution of the target object to browsing behavior in multiple video time periods, as well as the contribution of the target object to performing low-dwelling behavior in candidate videos. Thus, the distribution weights can be adaptively adjusted based on the diverse interaction behaviors of the target object to fuse multiple browsing behavior distribution information, accurately determine the browsing duration of the target object for candidate videos, and more accurately select target videos that match the needs of the target object, thereby improving the user experience of the target object.
[0117] In some embodiments, the browsing behavior distribution information includes a browsing behavior curve, which represents the probability density of a target object performing browsing behavior for a target video segment.
[0118] Based on the distribution weights corresponding to the distribution information of multiple browsing behaviors, the fusion of multiple browsing behavior distribution information includes: fusing multiple browsing behavior curves according to the multiple distribution weights to determine the target behavior curve; and detecting the target behavior curve to determine the browsing duration.
[0119] For example, browsing behavior curves include at least one of Gaussian and exponential distribution curves. The target behavior curve can be output by weighted fusion of multiple Gaussian and exponential distribution curves.
[0120] In one example, the probability of behavior corresponding to each preset browsing duration is determined by the probability density distribution information represented by the target behavior curve. Then, the preset browsing duration with the highest behavior probability can be selected from multiple preset browsing durations as the browsing duration.
[0121] In some embodiments, detecting the target behavior curve and determining the browsing duration may further include: determining a fluctuation curve segment with a peak attribute from the target behavior curve; determining the behavior probability corresponding to the target video time period based on the fluctuation curve segment; and determining the browsing duration based on the behavior probabilities corresponding to each of the multiple target video time periods.
[0122] The fluctuation curve segment is related to the target video time period of the candidate video. For example, the fluctuation curve segment can be a segment in the probability density curve that forms a peak region. By calculating the area formed between the fluctuation curve segment (as its envelope) and the horizontal axis, the probability of the target object performing a browsing behavior during the target video time period can be determined. Therefore, when there are multiple fluctuation curve segments, the behavior probability corresponding to each of the multiple target video time periods can be determined. Then, based on the behavior probabilities, the multiple target video time periods can be sorted or weighted and fused to determine the browsing duration.
[0123] Figure 6The schematic diagram illustrates the principle of a distribution information detection network for a deep learning model according to an embodiment of the present disclosure.
[0124] like Figure 6 As shown, the distribution information detection network includes a distribution parameter detection layer, a distribution weight detection layer, and a distribution curve detection layer. The deep learning model also includes a distribution curve fusion layer. The distribution curve fusion layer is used to fuse multiple browsing behavior curves to obtain the target behavior curve and output the browsing duration.
[0125] The target fusion feature input distribution weight detection layer outputs exponential and Gaussian distribution weights. The target fusion feature input distribution parameter detection layer outputs a parameter array consisting of k sets of mean and variance parameters. Parameter array 1 is (μ1, σ1), parameter array 2 is (μ2, σ2)... parameter array k is (μk, σk). The distribution parameter detection layer also outputs the rate parameter λ used for the exponential distribution function.
[0126] The distribution curve detection layer uses the exponential distribution function to process the rate parameter λ and outputs the exponential distribution curve 421. It also uses the Gaussian distribution function to process the k parameter arrays and outputs the first Gaussian distribution curve 411, the second Gaussian distribution curve 412, ..., the kth Gaussian distribution curve 41k.
[0127] The distribution curve fusion layer is based on the Gaussian distribution weights corresponding to the k Gaussian distribution curves output by the distribution weight detection layer. , and the exponential distribution weights used for the exponential distribution curve 421. Combine k Gaussian distribution curves and exponential distribution curves. Output the target behavior curve 431.
[0128] For example, the distribution curve fusion layer performs the processing based on the following formula (7).
[0129] (7)
[0130] Where p() represents the browsing duration, The weights of the exponential distribution corresponding to the exponential distribution curve are: Let be the weight of the k-th Gaussian distribution corresponding to the k-th Gaussian distribution curve. The values of the exponential distribution weight and the k-th Gaussian distribution weight satisfy formula (8).
[0131] (8).
[0132] The browsing duration was determined by detecting the target behavior curve 431 using a deep learning model. Figure 6 The dashed line indicates the time interval between the start time of the candidate video and the time interval between the start time of the video.
[0133] In some embodiments, the browsing behavior curve may also include other types of probability density distribution curves, such as exponential-log-normal mixture distribution curves, Gamma-Gaussian mixture distribution curves, Weibull-Gaussian mixture distribution curves, and truncated Gaussian mixture distribution curves. By utilizing deep learning models to process the numerical values of the parameter information corresponding to the diverse probability density distribution curves, these diverse probability density distribution curves can more accurately represent the probability of the target object's interactive behavior towards candidate videos, as well as the accuracy of detecting the video segment where the interactive behavior is performed. Therefore, by dynamically determining the distribution weights to fuse the diverse probability density curves, the expected browsing time of the target object for candidate videos can be more accurately determined, thereby improving the accuracy of recommending target videos to the target object and increasing the target object's satisfaction.
[0134] By detecting the parameter information used in the probability density function based on the interactive behavior characteristics of the target object, and further generating multiple probability density curves, the multiple probability density curves, which serve as browsing distribution information, can more accurately represent the near-zero skewness and fine-grained multi-peak features exhibited by the target object in interacting with candidate videos. This can further improve the perception of the target object's preference for candidate videos and browsing behavior, improve the accuracy of browsing duration detection, and thus push satisfactory target videos to the target object, improving the matching accuracy of resource recommendations.
[0135] In some embodiments, multiple candidate videos are ranked according to the viewing duration of the candidate videos output by the deep learning model and recommendation metrics such as completion rate and interaction rate related to the candidate videos. Based on the ranking results, the target object's preference for the target video can be identified more accurately, and video resources can be accurately pushed to the target object to meet the actual needs of the target object and reduce the computational overhead and communication resource consumption caused by redundantly pushing video resources to the target object.
[0136] Based on the video recommendation method provided in the above embodiments, embodiments of this disclosure also provide a method for training a deep learning model.
[0137] Figure 7 A flowchart illustrating a method for training a deep learning model according to an embodiment of the present disclosure is shown schematically.
[0138] like Figure 7 As shown, the method for training the deep learning model includes operating S710~S730.
[0139] The S710 is used to receive training samples.
[0140] According to embodiments of this disclosure, the training samples include sample interaction behavior features of sample objects, sample resource features related to sample videos, and sample objects' tag browsing duration for sample videos.
[0141] When operating the S720, the following operations are performed using a deep learning model to determine the sample browsing duration.
[0142] Based on the characteristics of sample interaction behavior and sample resource features, the interaction behavior patterns of sample objects are detected to obtain sample browsing behavior distribution information. The sample browsing behavior distribution information represents the probability distribution of the sample object's behavior of performing browsing behavior for video content of at least one video segment in the sample video. The sample browsing duration is determined based on the sample browsing behavior information.
[0143] When operating the S730, a deep learning model is trained based on the sample browsing time and tag browsing time to obtain the trained deep learning model.
[0144] It should be noted that the technical terms involved in the method for training deep learning models provided in this disclosure, including but not limited to sample interaction behavior features, sample resource features, and sample browsing behavior distribution information, have the same or similar meanings as the technical terms involved in the video recommendation method provided in this disclosure, including but not limited to interaction behavior features, resource features, and browsing behavior distribution information. The embodiments of this disclosure will not be repeated here.
[0145] The trained deep learning model determined by the method for training deep learning models provided in this disclosure can be used in the video recommendation method provided in this disclosure. The model structure of the deep learning model involved in the method for training deep learning models provided in this disclosure is the same as or similar to the model structure of the deep learning models involved in the above embodiments, and will not be described again in the embodiments of this disclosure.
[0146] In some embodiments, the sample browsing duration is determined by fusing multiple sample browsing behavior distribution information using multiple sample distribution weights. The multiple sample distribution weights are determined by using a deep learning model to process interaction behavior features and resource features.
[0147] For example, the distributed weight detection layer of a deep learning model can be used to process the target fusion features determined by fusing interactive behavior features and resource features, and output multiple sample distribution weights.
[0148] It should be noted that the various networks or model layers in a deep learning model can be constructed based on any type of neural network algorithm. The embodiments disclosed herein will not be described in detail here.
[0149] For example, training a deep learning model based on sample browsing time and tag browsing time includes: determining first loss information based on sample browsing time and tag browsing time; processing the distribution weights of multiple samples according to a loss function to obtain second loss information; and training a deep learning model based on the first loss information and the second loss information.
[0150] In one example, the first loss information is determined based on the following formulas (9) and (10).
[0151] (9).
[0152] (10).
[0153] in, For sample browsing duration, t i L represents the duration of tag browsing. reg This is the first piece of information indicating a loss.
[0154] The second loss function can utilize comparative loss functions such as entropy regularization loss function to process the distribution weights of multiple samples and output the second loss information.
[0155] For example, the second loss information can be determined based on the following formula (11).
[0156] (11).
[0157] Where, x i L represents the interactive behavior features and resource features input to the deep learning model. entropy This indicates the second loss information.
[0158] Therefore, the second loss information can be used to maintain the difference between the weights of multiple sample distributions to avoid the deep learning model focusing on one of the multiple sample browsing behavior distributions, which would lead to a lack of perception of the diverse browsing behavior patterns of the target object. Thus, the second loss information encourages the weights of multiple sample distributions to reduce the difference, thereby improving the accuracy of the deep learning model in perceiving diverse browsing behavior patterns and improving the detection accuracy of the target object's browsing duration of candidate videos.
[0159] The joint loss information is determined based on the first loss information and the second loss information. Then, the model parameters of the deep learning model are adjusted based on the joint loss information until the training conditions are met, and the trained deep learning model is determined.
[0160] In some embodiments, the deep learning model determines the detection probability of multiple candidate browsing durations based on sample browsing behavior distribution information, and determines the sample browsing duration from multiple candidate browsing durations based on multiple detection probabilities, wherein at least one candidate browsing duration is associated with a tag browsing duration.
[0161] For example, detection is performed using the probability density distribution curve of the samples, and the detection probability of the candidate browsing duration is output. This allows the determination of the detection probability for each of the multiple sample browsing durations.
[0162] Training a deep learning model based on first loss information and second loss information includes: training a deep learning model based on first loss information, second loss information and third loss information.
[0163] The third loss information is determined by processing the detection probability corresponding to the tag browsing time using a loss function.
[0164] For example, the third loss information can be determined based on the following formula (12).
[0165] (12).
[0166] in, For the training sample x of the i-th input i In this case, the deep learning model outputs the detection probability corresponding to the browsing duration of the i-th tag. When the training samples include N samples, the loss information corresponding to the N detection probabilities is fused to determine the third loss information, where N ≥ i > 1. Based on the third loss information...
[0167] Training a deep learning model based on the first, second, and third loss information can include determining the joint loss information L based on the following formula (13). total By using joint loss information to adjust the model parameters of the deep learning model, a trained deep learning model is obtained.
[0168] (13).
[0169] Wherein, α, β, and γ are the weighting coefficients used for each loss information, and the loss information includes first loss information, second loss information, third loss information, and business objective loss information L. biz The business objective loss information is calculated based on the detection probability of target interaction behavior for candidate videos output by the deep learning model, and the label probability related to the target interaction behavior. The deep learning model processes resource features and interaction behavior features to output the detection probability of the target interaction behavior. Target interaction behavior includes any type of interaction behavior such as liking and commenting.
[0170] By utilizing diverse loss information to determine joint loss information, and then training a deep learning model based on the target loss information, the probability distribution of the target object's diverse browsing behavior on candidate videos can be learned more accurately. This allows the deep learning model to understand the interaction behavior patterns of the target object on candidate videos more accurately, thereby improving the recommendation accuracy of the target video.
[0171] Figure 8 A block diagram of a video recommendation apparatus according to an embodiment of the present disclosure is shown schematically.
[0172] like Figure 8 As shown, the video recommendation device 800 includes: a first receiving module 810, a detection module 820, and a recommendation module 830.
[0173] The first receiving module 810 is used to receive the interactive behavior features of the target object and the resource features related to the candidate video.
[0174] The detection module 820 is used to detect the interaction behavior pattern of the target object based on the interaction behavior characteristics and resource characteristics, and obtain browsing behavior distribution information. The browsing behavior distribution information represents the probability distribution of the target object's browsing behavior for video content of at least one video segment in the candidate video.
[0175] The recommendation module 830 is used to determine the target video from the candidate videos based on the browsing behavior distribution information and recommend the target video to the target audience.
[0176] According to embodiments of this disclosure, the detection module includes: a fusion unit, a parameter detection unit, and a first determination unit.
[0177] The fusion unit is used to fuse interactive behavior features and resource features to obtain the target fusion features.
[0178] The parameter detection unit is used to perform parameter detection on the target fusion features and determine the parameter information used for the probability density function.
[0179] The first determining unit is used to process parameter information based on the probability density function to determine the browsing behavior distribution information.
[0180] According to embodiments of this disclosure, the probability density function includes a Gaussian distribution function, and the parameter information for the Gaussian distribution function includes a first reference parameter characterizing a reference time of the video segment, and a second reference parameter characterizing the degree of dispersion of the behavior probability; wherein, the browsing behavior distribution information includes a Gaussian distribution curve, which characterizes the probability distribution of the target object browsing the video content of the video segment.
[0181] According to embodiments of this disclosure, the probability density function includes an exponential distribution function, and the parameter information includes a rate parameter for the exponential distribution function; wherein, the browsing behavior distribution information includes an exponential distribution curve, which represents the probability distribution of the behavior of the target object to exit browsing the candidate video after browsing the video content of a preset video period in the candidate video.
[0182] According to embodiments of this disclosure, the interactive behavior features include multiple features, which characterize the interactive behavior of the target object in multiple interactive time periods, and the multiple interactive time periods correspond to multiple time period types; wherein, the fusion unit includes: an expert fusion feature acquisition subunit, an expert weight acquisition subunit, and a target fusion feature acquisition subunit.
[0183] The expert fusion feature acquisition subunit is used to process interactive behavior features using expert networks corresponding to time period types and output expert fusion features.
[0184] The expert weight acquisition subunit is used to determine the expert weights corresponding to each of the multiple expert fusion features based on multiple interaction behavior features and resource features. The expert weights represent the degree of influence of the interaction behavior performed by the target object during the interaction period with the type on the browsing behavior pattern of the candidate video.
[0185] The target fusion feature acquisition sub-unit is used to fuse multiple expert fusion features based on multiple expert weights to obtain the target fusion feature.
[0186] According to embodiments of this disclosure, the expert network corresponding to the time period type includes multiple behavioral expert networks, each of which corresponds to the interactive behavior of multiple behavioral types. The expert fusion feature acquisition sub-unit is configured to: use the behavioral expert network corresponding to the behavioral type to process the interactive behavior sub-features corresponding to the interactive behavior of the behavioral type in the interactive behavior features to obtain expert fusion sub-features, and multiple expert weights are respectively related to the expert fusion sub-features corresponding to the multiple behavioral types in the expert fusion features.
[0187] According to embodiments of this disclosure, the recommendation module includes: a browsing duration determination unit and a target video determination unit.
[0188] The browsing duration determination unit is used to determine the browsing duration corresponding to the candidate video by fusing multiple browsing behavior distribution information according to their respective distribution weights. The distribution weights are determined by weight detection of the target fusion features determined by fusing interactive behavior features and resource features.
[0189] The target video determination unit is used to determine the target video from at least one candidate video based on the browsing duration.
[0190] According to embodiments of this disclosure, the browsing behavior distribution information includes a browsing behavior curve, which represents the probability density of a target object performing browsing behavior for a target video time period; wherein, the browsing duration determination unit includes: a first determination subunit and a second determination subunit.
[0191] The first determining sub-unit is used to determine the target behavior curve by fusing multiple browsing behavior curves based on multiple distribution weights.
[0192] The second determining subunit is used to detect the target behavior curve and determine the browsing duration.
[0193] According to an embodiment of this disclosure, the second determining subunit is configured to: determine a fluctuation curve segment with a peak attribute from the target behavior curve, the fluctuation curve segment being related to the target video time period of the candidate video; determine the behavior probability corresponding to the target video time period based on the fluctuation curve segment; and determine the browsing duration based on the behavior probabilities corresponding to each of the multiple target video time periods.
[0194] Figure 9 A block diagram of an apparatus for training a deep learning model according to an embodiment of the present disclosure is shown schematically.
[0195] like Figure 9 As shown, the apparatus 900 for training deep learning models includes:
[0196] The second receiving module 910 is used to receive training samples, which include sample interaction behavior features of sample objects, sample resource features related to sample videos, and sample object's tag browsing time for sample videos.
[0197] The determination module 920 is used to perform the following operations using a deep learning model to determine the sample browsing duration: detect the interaction behavior pattern of the sample object based on the sample interaction behavior characteristics and sample resource characteristics, and obtain the sample browsing behavior distribution information. The sample browsing behavior distribution information represents the probability distribution of the sample object's behavior of performing browsing behavior for at least one video segment of the video content in the sample video; determine the sample browsing duration based on the sample browsing behavior distribution information.
[0198] Training module 930 is used to train a deep learning model based on sample browsing time and tag browsing time to obtain a trained deep learning model.
[0199] According to embodiments of this disclosure, the sample browsing duration is determined by fusing multiple sample browsing behavior distribution information using multiple sample distribution weights. The multiple sample distribution weights are determined by processing interaction behavior features and resource features using a deep learning model. The training module includes: a first loss information determination unit, a second loss information acquisition unit, and a training unit.
[0200] The first loss information determination unit is used to determine the first loss information based on the sample browsing time and the tag browsing time.
[0201] The second loss information acquisition unit is used to process the distribution weights of multiple samples according to the loss function to obtain the second loss information.
[0202] The training unit is used to train a deep learning model based on the first loss information and the second loss information.
[0203] According to embodiments of this disclosure, a deep learning model determines the detection probability of multiple candidate browsing durations based on sample browsing behavior distribution information, and determines the sample browsing duration from multiple candidate browsing durations based on multiple detection probabilities, wherein at least one candidate browsing duration is associated with a tag browsing duration; wherein, the training unit includes a training subunit.
[0204] The training subunit is used to train a deep learning model based on the first loss information, the second loss information, and the third loss information. The third loss information is determined by processing the detection probability corresponding to the tag browsing time using a loss function.
[0205] Figure 10 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.
[0206] In embodiments of this disclosure, such as Figure 10 As shown, the AI agent 1000 may include an input module 1010, a processing module 1020, and an output module 1030.
[0207] Input module 1010 is used to receive input information;
[0208] Processing module 1020 is used to determine the target task based on the input information received by the input module, determine the large model based on the target task, and obtain output information by calling the large model to execute the video recommendation method or the method for training a deep learning model according to the embodiments of this disclosure.
[0209] Output module 1030 is used to output the output information obtained by the processing module.
[0210] According to embodiments of this disclosure, the input module 1010 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the AI agent 1000 can understand and process. The input module 1010 is the primary link for the AI agent 1000 to interact with the outside world, enabling the AI agent 1000 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0211] In the example, input module 1010 can input the interactive behavior features and training samples described above.
[0212] In the example, processing module 1020 is the core support for the AI agent 1000's ability to handle complex tasks. Processing module 1020 can execute the video recommendation method or the method for training deep learning models described above.
[0213] In the example, the performance of the processing module 1020 is closely related to the large model on which the AI agent 1000 is based. To fully leverage the capabilities of the large model, the internal structure of the processing module 1020 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.
[0214] In the example, after the AI agent 1000 obtains the required voice, the processing module 1020 can use a large model to process the interactive behavior features and resource features, obtain browsing behavior distribution information, determine the target video based on the browsing behavior distribution information, and pass the target video to the output module 1030.
[0215] Understandably, while large language models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks is limited without the aid of any tools. Once the AI agent 1000 is given the ability to invoke tools, it can perform tasks such as using a calculator to complete mathematical calculations, using Python to perform data analysis, and using a search engine to create weather forecasts.
[0216] In the example, output module 1030 can output the target video or the trained deep learning model described above.
[0217] The AI agent 1000 according to the embodiments of this disclosure can simply and effectively improve the level of intelligence, as well as enhance flexibility and versatility.
[0218] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0219] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.
[0220] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the method described above.
[0221] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.
[0222] Figure 11 A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0223] like Figure 11 As shown, the electronic device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded into a random access memory (RAM) 1103 from a storage unit 1108. The RAM 1103 may also store various programs and data required for the operation of the electronic device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0224] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0225] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as video recommendation methods or methods for training deep learning models. For example, in some embodiments, the video recommendation method or the method for training deep learning models can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the video recommendation method or the method for training deep learning models described above can be performed. Alternatively, in other embodiments, computing unit 1101 may be configured by any other suitable means (e.g., by means of firmware) to perform a video recommendation method or a method for training a deep learning model.
[0226] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0227] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0228] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0229] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0230] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0231] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.
[0232] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0233] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A video recommendation method, comprising: Receive the interactive behavior characteristics of the target object and the resource characteristics related to the candidate video; Based on the interaction behavior features and the resource features, the interaction behavior pattern of the target object is detected to obtain browsing behavior distribution information. The browsing behavior distribution information represents the probability distribution of the target object performing browsing behavior for video content of at least one video segment in the candidate video. The target video is determined from the candidate videos based on the browsing behavior distribution information, and the target video is recommended to the target object.
2. The method according to claim 1, wherein, The step of detecting the interaction behavior pattern of the target object based on the interaction behavior features and the resource features includes: By fusing the interaction behavior features and the resource features, a target fusion feature is obtained; Parameter detection is performed on the target fusion features to determine the parameter information used for the probability density function; and The parameter information is processed according to the probability density function to determine the browsing behavior distribution information.
3. The method according to claim 2, wherein, The probability density function includes a Gaussian distribution function, and the parameter information for the Gaussian distribution function includes a first reference parameter characterizing the reference time of the video segment, and a second reference parameter characterizing the degree of dispersion of the behavior probability. The browsing behavior distribution information includes a Gaussian distribution curve, which represents the probability distribution of the target object's browsing behavior for the video content of the video segment.
4. The method according to claim 2 or 3, wherein, The probability density function includes an exponential distribution function, and the parameter information includes a rate parameter for the exponential distribution function; The browsing behavior distribution information includes an exponential distribution curve, which represents the probability distribution of the target object's behavior of exiting the browsing behavior after browsing the video content of a preset video period in the candidate videos.
5. The method according to claim 2, wherein the interactive behavior features include multiple features, the multiple interactive behavior features characterize the interactive behavior of the target object in multiple interactive time periods, and the multiple interactive time periods correspond to multiple time period types; in, The process of fusing the interaction behavior features and the resource features to obtain the target fused features includes: The interactive behavior features are processed using an expert network corresponding to the time period type, and expert fusion features are output; and Based on multiple interactive behavior features and resource features, expert weights are determined for each of the multiple expert fusion features, wherein the expert weights represent the degree of influence of the interactive behavior performed by the target object during the interaction period with the type on the browsing behavior pattern of the candidate video; and The target fusion feature is obtained by fusing multiple expert fusion features based on multiple expert weights.
6. The method according to claim 5, wherein, The expert network corresponding to the time period type includes multiple behavioral expert networks, and each of the multiple behavioral expert networks corresponds to the interactive behavior of multiple behavioral types; The step of processing the interactive behavior features using an expert network corresponding to the time period type includes: The interaction behavior sub-features corresponding to the interaction behavior type in the interaction behavior features are processed by the behavior expert network corresponding to the behavior type to obtain expert fusion sub-features. The multiple expert weights are respectively related to the expert fusion sub-features corresponding to the multiple behavior types in the expert fusion features.
7. The method according to claim 1 or 2, wherein, The step of determining the target video from the candidate videos based on the browsing behavior distribution information includes: Based on the distribution weights corresponding to the various browsing behavior distribution information, the multiple browsing behavior distribution information are fused to determine the browsing duration corresponding to the candidate video. The distribution weights are determined by weight detection of the target fusion feature determined by fusing the interaction behavior features and the resource features. The target video is determined from at least one of the candidate videos based on the browsing duration.
8. The method according to claim 7, wherein, The browsing behavior distribution information includes a browsing behavior curve, which represents the probability density of the target object performing browsing behavior for a target video time period; The step of fusing multiple browsing behavior distribution information based on their respective distribution weights includes: The target behavior curve is determined by fusing multiple browsing behavior curves based on multiple distribution weights; and The target behavior curve is detected to determine the browsing duration.
9. The method according to claim 8, wherein, The step of detecting the target behavior curve and determining the browsing duration includes: From the target behavior curve, determine the fluctuation curve segment with the peak attribute, the fluctuation curve segment being related to the target video time period of the candidate video; Based on the fluctuation curve segment, determine the behavioral probability corresponding to the target video time period; and The browsing duration is determined based on the behavioral probabilities corresponding to each of the multiple target video time periods.
10. A method for training a deep learning model, comprising: Receive training samples, the training samples including sample interaction behavior features of sample objects, sample resource features related to sample videos, and the tag browsing duration of the sample objects for the sample videos; The deep learning model is used to perform the following operations to determine the sample browsing time: Based on the sample interaction behavior features and the sample resource features, the interaction behavior pattern of the sample object is detected to obtain sample browsing behavior distribution information. The sample browsing behavior distribution information represents the probability distribution of the sample object performing browsing behavior for video content of at least one video segment in the sample video. The sample browsing duration is determined based on the sample browsing behavior distribution information; The deep learning model is trained based on the sample browsing time and the tag browsing time to obtain the trained deep learning model.
11. The method according to claim 10, wherein, The sample browsing duration is determined by fusing multiple sample browsing behavior distribution information using multiple sample distribution weights. The multiple sample distribution weights are determined by processing the interaction behavior features and resource features using the deep learning model. The step of training the deep learning model based on the sample browsing duration and the tag browsing duration includes: The first loss information is determined based on the sample browsing duration and the tag browsing duration; The second loss information is obtained by processing the distribution weights of multiple samples according to the loss function; and The deep learning model is trained based on the first loss information and the second loss information.
12. The method according to claim 11, wherein, The deep learning model determines the detection probability of multiple candidate browsing durations based on the sample browsing behavior distribution information, and determines the sample browsing duration from multiple candidate browsing durations based on the multiple detection probabilities, wherein at least one of the candidate browsing durations is associated with the tag browsing duration; The step of training the deep learning model based on the first loss information and the second loss information includes: The deep learning model is trained based on the first loss information, the second loss information, and the third loss information, wherein the third loss information is determined by processing the detection probability corresponding to the tag browsing time using a loss function.
13. A video recommendation device, comprising: The first receiving module is used to receive the interactive behavior features of the target object and the resource features related to the candidate video; The detection module is used to detect the interaction behavior pattern of the target object based on the interaction behavior features and the resource features, and obtain browsing behavior distribution information. The browsing behavior distribution information represents the probability distribution of the target object performing browsing behavior for video content of at least one video segment in the candidate video. The recommendation module is used to determine the target video from the candidate videos based on the browsing behavior distribution information, and recommend the target video to the target object.
14. An apparatus for training a deep learning model, comprising: The second receiving module is used to receive training samples, which include sample interaction behavior features of sample objects, sample resource features related to sample videos, and the tag browsing duration of the sample objects for the sample videos. The determination module is used to perform the following operations using the deep learning model to determine the sample browsing duration: Based on the sample interaction behavior features and the sample resource features, the interaction behavior pattern of the sample object is detected to obtain sample browsing behavior distribution information. The sample browsing behavior distribution information represents the probability distribution of the sample object performing browsing behavior for video content of at least one video segment in the sample video. The sample browsing duration is determined based on the sample browsing behavior distribution information; The training module is used to train the deep learning model based on the sample browsing time and the tag browsing time to obtain the trained deep learning model.
15. An intelligent agent of artificial intelligence, comprising: The input module is used to receive input information; The processing module is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and obtain output information by calling the large model to execute the method of any one of claims 1 to 12. An output module is used to output the output information obtained by the processing module.
16. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 12.
17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 12.
18. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 12.