Task processing method and device, equipment, storage medium and product
Through the multi-task processing model, multi-dimensional features of video, object and user are extracted, combined with sub-model processing of different task types, the problem of poor adaptability to the target user in the prior art is solved, and higher adaptability determination accuracy is achieved.
Patent Information
- Application Number
- CN202510614961.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, it is determined that the adaptability of the video to be pushed to the target user only depends on the video content or the content of the target object, resulting in poor adaptability.
By obtaining video attribute information, target object attribute information and target user association information, the pre-trained multi-task processing model is used for feature analysis, feature splicing sequences, cross features, adaptive features and preference features are extracted, combined with sub-model processing of different task types, and task processing results are output.
The accuracy and reliability of adaptability determination between the video to be pushed and the target user is improved, and the problem of low accuracy in single-dimensional adaptability determination is solved.
Smart Images

Figure CN120475201A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to a task processing method, apparatus, device, storage medium, and product. Background Art
[0002] As the audience for long and short videos grows, corresponding target objects are usually attached to the videos. In other words, in video scenarios, it is crucial to determine the videos to be pushed that are suitable for the target users and include the target objects. The target object mainly refers to the object that the target user wants to obtain.
[0003] Currently, the main method for determining the video to be pushed is to determine its compatibility with the target user based on the content of the video to be pushed, or to determine its compatibility with the target user based on the object content of the target object.
[0004] When implementing the technical solution based on the above method, the inventors found the following problems:
[0005] The above method only relies on the video content of the video to be pushed or the object content of the target object to determine its compatibility with the target user. There is a single reference factor, which leads to the problem that the determined video to be pushed or target object has poor compatibility with the target user. Summary of the Invention
[0006] The embodiments of the present invention provide a task processing method, apparatus, equipment, storage medium and product to determine the cross-features between the video to be pushed, the target object mounted in the video to be pushed and the target user from multiple angles, thereby obtaining the processing results corresponding to multiple task types, and improving the technical effect of the accuracy of determining the processing results.
[0007] In a first aspect, an embodiment of the present invention provides a task processing method, the method comprising:
[0008] Obtaining video attribute information of a video to be pushed that is mounted with a target object, object attribute information of the target object, and user association information of a target user; wherein the target user is the user to whom the video to be pushed is to be pushed;
[0009] Based on a pre-trained multi-task processing model, feature analysis processing is performed on the video attribute information, the object attribute information, and the user association information to obtain at least one feature to be used; wherein, the feature to be used includes a feature splicing sequence, a cross feature, an adaptation feature, or a preference feature, the feature splicing sequence is used to characterize the features obtained by splicing the video attribute information, the object attribute information, and the user association information after feature extraction, the cross feature is used to characterize the cross information between any two of the video to be pushed, the target object, and the target user, the adaptation feature is used to characterize the adaptation information of the target object and the video to be pushed relative to the target user, and the preference feature is used to characterize the target user's preference information for the video and the object;
[0010] The at least one feature to be used is processed based on the task processing sub-models corresponding to different task types in the multi-task processing model, and task processing results corresponding to different task types are output.
[0011] Furthermore, the method further comprises:
[0012] Based on a first preset condition, determining first video information, first object information, and first user information for determining a feature sequence from the video attribute information, the object attribute information, and the user association information; and / or,
[0013] Based on a second preset condition, obtaining second video information, second object information, and second user information for determining the cross feature from the video attribute information, the object attribute information, and the user association information; and / or,
[0014] Based on a third preset condition, third video information, third object information, and third user information for determining adaptation features are obtained from the video attribute information and the user association information; wherein the third user information includes historical viewing behavior information of the target user within a historical preset time period, and historical interaction behavior information of the target user with the displayed object.
[0015] Furthermore, the feature to be used corresponds to a feature splicing sequence, the feature splicing sequence includes a first feature sequence and / or a second feature sequence, and the multi-task processing model based on pre-training performs feature analysis processing on the video attribute information, the object attribute information, and the user association information to obtain at least one feature to be used, including:
[0016] Normalizing the first video information, the first object information, and the first user information based on a normalization module, and concatenating the first video features, the first object features, and the first user features obtained after the normalization to obtain a first feature sequence; and / or,
[0017] Based on the first feature extraction module, the first video information, the first object information and the first user information are respectively extracted for feature extraction, and the extracted second video features, second object features and second user features are spliced to obtain a second feature sequence.
[0018] Furthermore, the feature to be used corresponds to a cross feature, and the cross feature includes a first cross feature between the target object and the video to be pushed, a second cross feature between the video to be pushed and the target object, and a third cross feature between the target object and the target user. The multi-task processing model based on pre-training performs feature analysis processing on the video attribute information, the object attribute information, and the user association information to obtain at least one feature to be used, including:
[0019] Based on the prediction sub-model in the multi-task processing model, cross-feature extraction is performed on the second video information and the second object information to obtain the first cross-feature, cross-feature extraction is performed on the second video information and the second user information to obtain the second cross-feature, and cross-feature extraction is performed on the second object information and the second user information to obtain the third cross-feature.
[0020] Furthermore, the features to be used also include fusion features, and the multi-task processing model also includes a fusion sub-model, which is used to perform feature fusion on the second video information, the second object information and the second user information to obtain the fusion features.
[0021] Furthermore, the feature to be used corresponds to an adaptation feature, the adaptation feature includes a video adaptation feature and an object adaptation feature, the multi-task processing model also includes a first model and a second model with the same model structure, and the multi-task processing model based on pre-training performs feature analysis processing on the video attribute information, the object attribute information, and the user association information to obtain at least one feature to be used, including:
[0022] Performing feature extraction on the third video information and the historical viewing behavior information in the third user information based on the first model to obtain a video adaptation feature of the video to be pushed relative to the target user;
[0023] Based on the second model, feature extraction is performed on the historical interaction behavior information in the third video information and the third user information to obtain an object adaptation feature of the target object relative to the target user.
[0024] Furthermore, the multi-task processing model includes a long-term preference extraction sub-model for determining the preference feature, and determining the preference feature among the features to be used includes:
[0025] Obtaining the video single tag and video multi-tag in the video attribute information;
[0026] The first processing unit in the long-term preference extraction sub-model processes the video single label and the single label sequence corresponding to the video single label to determine a first label sequence, so as to determine a first to-be-weighted video feature of the corresponding historical video based on the first label sequence by a second processing unit; wherein the historical video is a video related to the first label sequence;
[0027] The third processing unit in the long-term preference extraction sub-model processes the video multi-label and the multi-label sequence corresponding to the video multi-label to obtain a second label sequence, and determines a second to-be-weighted video feature of the corresponding historical video based on the second label sequence by the fourth processing unit;
[0028] The video preference feature is obtained by weighting the first video feature to be weighted and the second video feature to be weighted.
[0029] Furthermore, the method further comprises:
[0030] By processing the video preference feature and the object preference feature in the preference feature, a video preference attribute corresponding to the video preference feature and an object preference attribute corresponding to the object preference feature are obtained, so as to update the video preference feature based on the video preference attribute and update the object preference feature based on the object preference attribute.
[0031] Furthermore, processing the at least one feature to be used based on the task processing sub-models corresponding to different task types in the multi-task processing model and outputting task processing results corresponding to different task types includes:
[0032] In a case where the features to be used include at least two of a feature splicing sequence, a cross feature, an adaptation feature, and a preference feature, splicing the features to be used to obtain a target splicing feature;
[0033] The target splicing features are processed based on the multi-task learning sub-model in the multi-task processing model, and task processing results under different task types are output.
[0034] In a second aspect, an embodiment of the present invention further provides a task processing device, the device comprising:
[0035] An information acquisition module is used to obtain video attribute information of a video to be pushed that is mounted with a target object, object attribute information of the target object, and user association information of a target user; wherein the target user is the user to whom the video to be pushed is to be pushed;
[0036] a feature determination module, configured to perform feature analysis on the video attribute information, the object attribute information, and the user association information based on a pre-trained multi-task processing model to obtain at least one feature to be used; wherein the feature to be used includes a feature splicing sequence, a cross feature, an adaptation feature, or a preference feature; the feature splicing sequence is used to characterize the features obtained by splicing the video attribute information, the object attribute information, and the user association information after feature extraction; the cross feature is used to characterize the cross information between any two of the video to be pushed, the target object, and the target user; the adaptation feature is used to characterize the adaptation information of the target object and the video to be pushed relative to the target user; and the preference feature is used to characterize the target user's preference information for the video and the object;
[0037] The result output module is used to process the at least one feature to be used based on the task processing sub-models corresponding to different task types in the multi-task processing model, and output task processing results corresponding to different task types.
[0038] In a third aspect, an embodiment of the present invention provides an electronic device, the electronic device comprising:
[0039] one or more processors;
[0040] a memory for storing one or more programs;
[0041] When the one or more programs are executed by the one or more processors, the one or more processors implement the task processing method provided by any embodiment of the present invention.
[0042] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the task processing method provided by any embodiment of the present invention.
[0043] In a fifth aspect, an embodiment of the present invention further provides a computer program product, comprising a computer program, characterized in that when the computer program is executed by a processor, it implements the task processing method as described in any one of the embodiments of the present invention.
[0044] The technical solution provided by the embodiment of the present invention can input the acquired video attribute information of the video to be pushed with the target object mounted thereon, the object attribute information of the target object, and the user association information of the target user into a pre-trained multi-task processing model, and can determine multiple feature sequences, cross-features after information cross-processing, video adaptation features of the target user corresponding to the video to be pushed, object adaptation features of the target object relative to the target user, video preference features of the target user, and object preference features based on the multi-task processing model. Subsequently, the at least one feature is processed to obtain task processing results under different task types. Since the method provided by this embodiment can determine the features between the video and the user, the features between the user and the object, the features between the object and the video, the video adaptation features of the video relative to the user, the object adaptation features of the object relative to the user, and the target user's preference features and object preference features for the video, that is, the adaptation information of the video to be pushed and the target object relative to the target user is determined from multiple angles, thereby improving the accuracy and reliability of determining the target processing results, and solving the problem of low accuracy when determining the adaptability of the video or object relative to the target user from a single dimension in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings introduced here only illustrate some of the embodiments to be described by the present invention, and are not exhaustive. A person skilled in the art can derive other drawings based on these drawings without inventive effort.
[0046] Figure 1 A flowchart of a task processing method provided by an embodiment of the present invention;
[0047] Figure 2 A schematic diagram of a system architecture corresponding to a task processing method provided in an embodiment of the present invention;
[0048] Figure 3 A flowchart of a task processing method provided by an embodiment of the present invention;
[0049] Figure 4 A flowchart of a task processing method provided by an embodiment of the present invention;
[0050] Figure 5A flowchart of a task processing method provided by an embodiment of the present invention;
[0051] Figure 6 A flowchart of a task processing method provided by an embodiment of the present invention;
[0052] Figure 7 A schematic diagram of data processing based on the long-term preference extraction sub-model provided by an embodiment of the present invention;
[0053] Figure 8 A flowchart of a task processing method provided by an embodiment of the present invention;
[0054] Figure 9 A schematic diagram of the structure of a task processing device provided by an embodiment of the present invention;
[0055] Figure 10 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0057] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0058] Before introducing the technical solutions provided by the embodiments of the present invention, an example application scenario may be first described.
[0059] In the short video market, it's crucial to push appropriate short videos to target users. Accordingly, to increase sales of certain target audiences, you can embed the corresponding target audience in short videos. Therefore, determining the compatibility between short videos, target audiences, and target users becomes crucial.
[0060] Based on this, the solution provided by the embodiment of the present invention can be used to determine the cross-features between videos, objects, and users, thereby obtaining processing results corresponding to different task types. The processing results can be sent to downstream systems, so that the downstream systems can determine whether to push the video to the target user based on the multiple processing results.
[0061] The above method can consider the cross-features between users, objects and videos from multiple perspectives. Therefore, the obtained processing results are relatively accurate. Therefore, when used in the target downstream system, the accuracy of downstream task processing can be improved.
[0062] Figure 1 This is a flowchart of a task processing method provided in an embodiment of the present invention. This embodiment can be applied to scenarios where videos to be pushed are processed and mounted with target objects. The task processing method provided in this embodiment can be executed by the server, or by the client, or by the client or the server in coordination. The specific execution method can be set according to actual needs and is not limited in this embodiment.
[0063] like Figure 1 As shown, the method specifically includes the following steps:
[0064] S110: Obtain video attribute information of the video to be pushed that is mounted with the target object, object attribute information of the target object, and user association information of the target user.
[0065] The video to be pushed to the target user is referred to as the video to be pushed. Accordingly, the target user is the user to whom the video to be pushed is to be pushed. The object to be acquired can be mounted in the video to be pushed, and the object to be acquired can be mounted in the video to be pushed as the target object. The target object can be any physical object or virtual object. Optionally, the physical object can be a specific item, and the virtual object can be a coupon, shopping card, membership benefits, etc. The mounted target object can be set according to actual needs and is not limited in this embodiment.
[0066] Among them, the video attribute information may include the author identification of the video to be pushed, the video playback length, the video category label, etc. As long as it is attribute information related to the video to be pushed, it is all content in the video attribute information. Object attribute information may include the item category of the target object, which may include primary category, secondary category, tertiary category, etc.; it also includes the click information, exposure information, value attribute information, transaction volume information, etc. of the target object within a preset time period. User association information mainly refers to information associated with the target user. The user association information may include the interactive behavior data of the target user within a historical time period, as well as the historical interactive video corresponding to the generation of the interactive behavior data, the video attribute information corresponding to the historical interactive video, and the object information of the object to be obtained when the interactive behavior data is generated. Such information can be used as user association information of the target user.
[0067] In this embodiment, when determining the video to be pushed, the video attribute information of the video to be pushed, the object attribute information of the target object mounted in the video to be pushed, and the user association information of the target user to whom the video to be pushed will be pushed can be obtained.
[0068] It should be noted that the background can determine multiple videos to be pushed to the target user. For each video to be pushed, the specific processing method is the same. The processing of one of the videos to be pushed can be used as an example to illustrate.
[0069] It should also be noted that the solution provided by the embodiment of the present invention may be executed after detecting that an application has logged into a certain application program, or in an offline phase.
[0070] S120 , performing feature analysis on the video attribute information, the object attribute information, and the user association information based on a pre-trained multi-task processing model to obtain at least one feature to be used.
[0071] Among them, the multi-task processing model is composed of multiple sub-models and / or modules. The multi-task processing model includes at least one of a sub-model for feature extraction, a sub-model for determining cross-features, a sub-model for determining fusion features, a sub-model for determining the short-term interests of the target user, and a sub-model for determining the long-term interests of the user. The relevant information in the video attribute information, object attribute information and user association information can be processed based on different sub-models to obtain features adapted to each sub-model. The features to be used include feature splicing sequences, cross-features, adaptation features or preference features. The feature splicing sequence is used to characterize the features spliced after feature extraction of the video attribute information, object attribute information and user association information. The cross-feature is used to characterize the cross-information between any two of the video to be pushed, the target object and the target user. The adaptation feature is used to characterize the adaptation information of the target object and the video to be pushed relative to the target user. The preference feature is used to characterize the preference information of the target user for the video and the object.
[0072] In this embodiment, the features to be used correspond to a feature splicing sequence, and the feature splicing sequence includes a first feature sequence and / or a second feature sequence. The first feature sequence is a feature sequence extracted after model processing, and the second feature sequence can be a feature sequence obtained without data processing, that is, a feature sequence obtained after feature extraction of the original data. The cross-features include a first cross-feature, a second cross-feature, and a third cross-feature. The first cross-feature is used to characterize the cross-features between the video to be pushed and the target user, the second cross-feature is used to characterize the cross-features between the video to be pushed and the target object, and the third cross-feature is used to characterize the cross-features between the target object and the target user. The adaptation features include video adaptation features and object adaptation features. The video adaptation features are mainly based on the features corresponding to the short-term interests of the video to be pushed relative to the target user. The object adaptation features are mainly features determined by the short-term interests of the target object relative to the target user. The video preference features are features corresponding to the processing of video attribute information, and the object preference features are features obtained after processing object attribute information.
[0073] Specifically, video attribute information, object attribute information and user association information can be input into a pre-trained multi-task processing model, and can be processed based on the sub-models or sub-modules used for processing different data in the multi-task processing model to obtain feature information corresponding to different dimensions.
[0074] S130 : Process the at least one feature to be used based on the task processing sub-models corresponding to different task types in the multi-task processing model, and output task processing results corresponding to different task types.
[0075] The multi-task processing model may include multiple task processing sub-models, and different task processing sub-models are used to process the features output by the multi-task processing model to obtain task processing results corresponding to corresponding task types.
[0076] It can be understood that multiple task types can be pre-set, and different task types correspond to different task processing sub-models. After obtaining the features output by the task processing model, different task processing sub-models can be used to process all features to obtain processing results corresponding to different task types.
[0077] The technical solution provided by the embodiment of the present invention can input the acquired video attribute information of the video to be pushed with the target object mounted thereon, the object attribute information of the target object, and the user association information of the target user into a pre-trained multi-task processing model, and can determine multiple feature sequences, cross-features after information cross-processing, video adaptation features of the target user corresponding to the video to be pushed, object adaptation features of the target object relative to the target user, video preference features of the target user, and object preference features based on the multi-task processing model. Subsequently, the at least one feature is processed to obtain task processing results under different task types. Since the method provided by this embodiment can determine the features between the video and the user, the features between the user and the object, the features between the object and the video, the video adaptation features of the video relative to the user, the object adaptation features of the object relative to the user, and the target user's preference features and object preference features for the video, that is, the adaptation information of the video to be pushed and the target object relative to the target user is determined from multiple angles, thereby improving the accuracy and reliability of determining the target processing results, and solving the problem of low accuracy when determining the adaptability of the video or object relative to the target user from a single dimension in the prior art.
[0078] Figure 2 A flowchart of a task processing method provided in an embodiment of the present invention. Based on the aforementioned embodiment, the multi-task processing model includes a first normalization module and a first feature extraction module. Accordingly, this embodiment can further refine the "determining the first feature sequence and the second feature sequence based on the multi-task processing model". The specific implementation method can refer to the detailed explanation of the present technical solution. Among them, the technical terms that are the same as or corresponding to the above embodiments are not repeated in this embodiment.
[0079] like Figure 2 As shown, the method provided in this embodiment includes:
[0080] S210: Determine first video information, first object information, and first user information that meet a first preset condition from the video attribute information, the object attribute information, and the user association information.
[0081] It should be noted that, based on the above, the video attribute information, object attribute information, and user association information include a relatively large amount of content. The data required by different sub-models in the multi-task processing model is different. That is, different sub-models in the multi-task processing model process portions of the attribute information, object attribute information, and user association information to obtain feature information under a certain dimension. At this point, information extraction can be performed on the video attribute information, user association information, and object attribute information based on the first preset condition to obtain data to be input into the normalization module and the first feature extraction module for feature extraction.
[0082] Among them, the first preset condition is pre-set and is used to obtain the constraints that can determine the first feature sequence and the second feature sequence. Optionally, the first preset condition can be a condition for extracting continuity data. For example, for video attribute information, features such as the number of video views and the number of viewing users can be considered to meet the first preset condition; for object attribute information, the number of times an object is obtained, etc., can all be used as continuity data, and such features can be obtained. The data in the video attribute information that meets the first preset condition is used as the first video information. The data in the object attribute information that meets the first preset condition is used as the first object information; the data in the user association information that meets the first preset condition is used as the first user information.
[0083] Specifically, according to the first preset condition, data with continuity features can be obtained from the video attribute information, the object attribute information, and the user association information, and used as the first user information, the first video information, and the first object information.
[0084] S220: Normalize the first video information, the first object information, and the first user information respectively based on the normalization module, and concatenate the first video features, the first object features, and the first user features obtained after the normalization to obtain a first feature sequence.
[0085] Among them, the multi-task processing model includes a normalization processing module, and the normalization processing module can be understood as a sub-model in the multi-task processing model. The function of the normalization processing module is mainly to normalize the input data to obtain a normalized feature vector. The first video feature is the feature obtained after normalizing the numerical value corresponding to the first video information; the first object feature is the feature obtained after normalizing the numerical value corresponding to the first object information; the first user feature is the feature obtained after normalizing the numerical value corresponding to the first user information. Correspondingly, the first feature sequence can be understood as the feature sequence obtained after splicing the first video feature, the first object feature and the first user feature in a specific way.
[0086] Specifically, after obtaining first user information, first video information, and first object information that meet the first preset condition, the data corresponding to the first user information, the first video information, and the first object information can be normalized using a normalization module to obtain features corresponding to the different information. After feature concatenation of the first video features, the first object features, and the first user features, a first feature sequence can be obtained.
[0087] S230: Extract the first video information, the first object information, and the first user information based on the first feature extraction module, perform feature extraction, and concatenate the extracted second video features, second object features, and second user features to obtain a second feature sequence.
[0088] The multi-task processing model includes a first feature extraction module. The first feature extraction module is used to extract video features, object features, and user features. The first feature extraction module and the normalization module use the same input data, but differ in their functions and the features they extract.
[0089] Specifically, the first video information feature can be extracted based on the first feature extraction module to obtain the second video feature. The first feature extraction module extracts the first object information feature to obtain the second object feature. The first feature extraction module extracts the first user information feature to obtain the second user feature. The second video feature, the second object feature and the second user feature can be determined sequentially or simultaneously. The specific processing order is related to the model structure of the first feature extraction module. As long as the relevant features can be obtained, the specific processing method and order are not limited in this embodiment. The second video feature, the second user feature and the second object feature are spliced and processed in a preset manner to obtain a second feature sequence.
[0090] Based on the above, it can be seen that the first feature sequence is a feature sequence obtained by directly normalizing the data that meets the first preset condition, and the second feature sequence is a feature sequence obtained by extracting features from the data that meets the first preset condition.
[0091] For example, see Figure 3 , will be as Figure 3The model corresponding to the system architecture diagram shown is called a multi-task processing model. The input data of the multi-task processing model is video attribute information, object attribute information, and user association information. Data that meets a first preset condition can be extracted from the video attribute information, object attribute information, and user association information. Specifically, data that meets the first preset condition includes first video information, first object information, and first user information. The normalization module and first feature extraction module in the multi-task processing model process the data, respectively, so that the normalization module outputs a first feature sequence and the first feature extraction module outputs a second feature sequence.
[0092] It should be noted that the normalization module directly performs normalization processing on the data to obtain feature information of the data, and the result output by the first feature extraction module is the feature data obtained after extracting the data features.
[0093] It should be noted that the above steps S220 and S230 can be used selectively or simultaneously. The above steps are only limited in order from the perspective of introduction, and are not limitations on the execution steps.
[0094] The technical solution provided by the embodiment of the present invention can analyze and process the first video information, first object information and first user information that meet the first preset condition to obtain a first feature sequence and a second feature sequence, thereby facilitating the subsequent determination of task processing results corresponding to different task types based on the first feature sequence and the second feature sequence, thereby improving the accuracy of the task processing results.
[0095] Figure 4 A flowchart of a task processing method provided in an embodiment of the present invention. Based on the aforementioned embodiment, the multi-task processing model includes a prediction sub-model and a feature fusion sub-model based on a retrospective memory mechanism. Accordingly, this embodiment can further refine: "Determining the first cross-feature, the second cross-feature, the third cross-feature and the fusion feature based on the multi-task processing model". The specific implementation method can refer to the detailed explanation of the present technical solution, wherein the technical terms that are the same as or corresponding to the above embodiments are not repeated in this embodiment.
[0096] like Figure 4 As shown, the method provided in this embodiment includes:
[0097] S310: Determine second video information, second object information, and second user information that meet a second preset condition from the video attribute information, the object attribute information, and the user association information.
[0098] The second preset condition is a constraint condition for determining the second video information, the second object information, and the second user information. The second preset condition may be a condition where the data content can be enumerated. Video attribute information that meets the second preset condition is used as the second video information. Object attribute information that meets the second preset condition is used as the second object information. User association information that meets the second preset condition is used as the second user information.
[0099] Specifically, based on the second preset condition, corresponding second video information, second object information and second user information may be screened out from the video attribute information, object attribute information and user association information.
[0100] For example, for videos, enumerable features include category tags, categories, and so on. That is, any enumerable features of the data content can serve as video information that meets the second preset condition. For object attribute information, this can be category information, primarily referring to enumerable object information. Accordingly, for user-related information, information corresponding to enumerable features is also used as the second user information.
[0101] S320. Based on the prediction sub-model in the multi-task processing model, cross-feature extraction is performed on the second video information and the second object information to obtain a first cross-feature, cross-feature extraction is performed on the second video information and the second user information to obtain a second cross-feature, and cross-feature extraction is performed on the second object information and the second user information to obtain a third cross-feature.
[0102] The multi-task processing model includes a prediction sub-model. The prediction sub-model is at the same level as the normalization module and the first feature extraction module mentioned above. Being at the same level primarily means that the processing order is the same. That is, after acquiring data that meets the second preset condition, the prediction sub-model is required to process the acquired data.
[0103] In this embodiment, the prediction sub-model can be any existing model that can determine cross-features. Optionally, the prediction sub-model can be MemoNet, which is a trajectory prediction sub-model based on a retrospective memory mechanism. After the second video information, the second object information, and the second object information are input into the prediction sub-model, the prediction sub-model can perform cross-feature feature extraction on the second video information and the second object information to obtain a first cross-feature between the target object and the video to be pushed. Based on the prediction sub-model, the second video information and the second user information are cross-feature extracted to obtain a second cross-feature between the video to be pushed and the target user, and the second object information and the second user information are processed based on the prediction sub-model to obtain a third cross-feature between the target object and the target user.
[0104] Exemplarily, the expression of the target object cannot be separated from the expression of the video to be pushed, and the expression of the video to be pushed cannot be separated from the compatibility between the target user and the target object. Correspondingly, the expression of the target object cannot be separated from the compatibility between the target user and the target object. Based on this, cross-features can be extracted in pairs. That is, second-order cross-features are used to capture the corresponding information expression. Optionally, three groups of cross-content are designed, namely the target user and the target object, the target user and the video to be pushed, and the target object and the video to be pushed. Based on the model architecture of the prediction sub-model, for the above three groups of cross-content, 5 hash buckets can be used for each group to represent the feature expression after feature cross-content. The feature elements to be crossed are fixed within a preset number range. Optionally, the preset number can be 20. The purpose is to avoid introducing useless cross-features to bring noise, thereby affecting the expression of information. Finally, the output result of the prediction sub-model includes cross-information features between the target object, the target user and the video to be pushed. At this time, the cross-information features have a certain correlation.
[0105] It should also be noted that although the input of the prediction sub-model includes the second video information, the second object information and the second user information, when processing it based on the prediction sub-model, it is also possible to perform cross-feature extraction on part of the information in the second video information, the second object information and the second user information according to the feature cross-field corresponding to the prediction sub-model, that is, further improving the accuracy of the cross-feature extraction.
[0106] S330 : Perform feature fusion on the second video information, the second object information, and the second user information based on the fusion sub-model to obtain fusion features.
[0107] It should be noted that the above steps S310 and S320 are executed in parallel.
[0108] The feature fusion sub-model can fuse features of the target user, target object, and the video to be pushed. Optionally, the feature fusion sub-model can be a machine learning model designed to address high-dimensional sparse data. This model can improve its predictive capabilities by introducing cross-feature terms and stabilization vectors. Accordingly, the fused features are features that fuse the aforementioned video attribute information, object attribute information, and user association information.
[0109] Specifically, after obtaining the second video information, second object information and second user information that meet the second preset conditions, feature extraction can be performed on the second video information, second object information and second user information based on the feature fusion sub-model to obtain fusion features after fusing the information of the above three dimensions.
[0110] For example, after obtaining the second user information, second video information, and second object information that meet the second preset condition, the second video information, second object information, and second user information can be subjected to feature fusion processing based on the feature fusion sub-model to obtain fused features. Accordingly, at the same time, the second user information, second object information, and second video information are cross-processed based on the prediction sub-model to obtain three sets of cross-features.
[0111] The technical solution provided by the embodiment of the present invention can use a prediction sub-model and a feature fusion sub-model to perform feature processing on the second video information, second object information and second user information that meet the second preset conditions, and obtain multiple cross-features and fusion features. This method can learn the cross-features, and accordingly, when the corresponding results are obtained based on the cross-features, the accuracy of determining the task processing results can be improved.
[0112] Figure 5 A flowchart of a task processing method provided in an embodiment of the present invention, based on the aforementioned embodiment, the multi-task processing model includes a first model and a second model, and the model structure of the first model and the second model are the same. Accordingly, this embodiment can further refine: "Determining video adaptation features and object adaptation features based on the multi-task processing model". Its specific implementation method can refer to the detailed explanation of this technical solution, wherein the technical terms that are the same as or corresponding to the above embodiments are not repeated in this embodiment.
[0113] like Figure 5 As shown, the method provided in this embodiment includes:
[0114] S410. Based on a third preset condition, third video information, third object information, and third user information for determining adaptation features are obtained from the video attribute information and the user association information, wherein the third user information includes historical viewing behavior information of the target user within a historical preset time period, and historical interaction behavior information of the target user with respect to the displayed object.
[0115] The third preset condition may be a preset condition for extracting short-term interest features. Optionally, the third preset condition corresponds to a condition for extracting feature information within a first preset time period before the current moment. Optionally, the video attribute information, user association information, and object attribute information correspond to data within a target time period. Feature extraction may be performed on the video attribute information, user association information, and object attribute information, respectively, based on the third preset condition to obtain third object information, third video information, third user information, and fourth user information that meet the third preset condition.
[0116] It should be noted that the third user information and the fourth user information can be the same or different, and whether they are the same is primarily related to the third preset condition. Accordingly, due to the differences in the features processed by the first and second models, the third user information is the user information corresponding to the target user's corresponding video to be pushed, that is, the information corresponding to the historical videos viewed by the target user within the first preset duration. The fourth user information is the user information of the target user relative to the target object, that is, the information corresponding to the objects clicked or accessed by the target user within the first preset duration.
[0117] That is to say, the third user information related to the video to be pushed can be extracted from the user association information based on the third preset condition. At this time, the third user information related to the video to be pushed can be a video feature sequence of historical videos watched by the target user within the first preset time period; at the same time, the fourth user information associated with the target object can be extracted from the user association information based on the third preset condition. At this time, the fourth user information associated with the target object can be an object sequence of objects clicked, browsed or obtained by the target user within the first preset time period.
[0118] Specifically, content extraction can be performed on the video attribute information, object attribute information, and user-related information based on the third preset condition to obtain third video information that meets the third preset condition and third user information related to the video to be pushed. Simultaneously, third object information that meets the third preset condition and fourth user information related to the target object can be obtained.
[0119] It can be understood that the third user information and the fourth user information that meet the third preset condition are user behavior sequences. In this case, the user behavior sequences mainly include behavior sequences on objects and behavior sequences on videos.
[0120] S420: Extract features of the third video information and the historical viewing behavior information in the third user information based on the first model to obtain video adaptation features of the video to be pushed relative to the target user.
[0121] Among them, the third user information is the historical viewing behavior information of the target user within the historical preset time period. The third user information is the video identification sequence or video sequence corresponding to the historical videos watched by the target user within the historical preset time period. The historical preset time period is the first preset time period in the third preset condition. The first model is a model for determining the degree of adaptability between the video to be pushed and the target user. The first model can be any existing model that can determine the adaptability characteristics between two sets of data. Optionally, the first model can be a DIN model. The video adaptation feature can be understood as the feature of the degree of adaptability of the video to be pushed relative to the target user. The video adaptation feature can be represented by a vector.
[0122] Specifically, the third video information and the third user information can be input into the first model. The first model can analyze and process the third video information and the third user information to obtain a video adaptation feature of the third video information relative to the third user information. In this case, the video adaptation feature is the adaptation feature of the video to be pushed relative to the target user.
[0123] S430: Extract features of the historical interaction behavior information in the third video information and the third user information based on the second model to obtain object adaptation features of the target object relative to the target user.
[0124] Among them, the historical interactive behavior information is the object information of clicking, browsing or obtaining the object before the current moment. The model architecture of the second model is the same as the model architecture of the first model. The difference is that the sample features of the training samples used in training are different. Accordingly, the content features input during application are different, and the output content features are also different. The display object can be an object displayed in the application. When the target user browses the application, the displayed display object can be clicked, and the data of the displayed display object can be collected and used as the user association information of the target user. Accordingly, the operation behavior data on the display object within the first preset time length is used as the historical behavior interaction information in the third user information. The object adaptation feature can be understood as the adaptation feature of the target object relative to the target user.
[0125] Specifically, the third object information and the fourth user may be analyzed and processed based on the second model to output an object adaptation feature of the target object relative to the target user.
[0126] For example, see Figure 3 The third video information of the video to be pushed that meets the third preset condition and the video sequence of the historical videos watched by the target user within the historical preset time period are input into the first model, so that the first model outputs the video adaptation feature. At the same time, the third object information (target item category) of the target object that meets the third preset condition and the target user's operational behavior data on the displayed object within the historical preset time period are used to determine the object sequence as the fourth user information. The fourth user information and the third object information are input into the second model, and the object adaptation feature is output.
[0127] The technical solution provided by the embodiment of the present invention can obtain the video adaptation features of the video to be pushed relative to the target user by inputting the third user information and the third video information that meet the third preset conditions into the first model, and can obtain the object adaptation features of the target object relative to the target user by inputting the fourth user information and the third object information that meet the third preset conditions into the second model, so as to facilitate the subsequent combination of the video adaptation features and the object adaptation features to determine the accuracy and efficiency of the target processing results corresponding to different task types.
[0128] Figure 6 A flowchart of a task processing method provided in an embodiment of the present invention. Based on the aforementioned embodiment, the multi-task processing model includes a long-term preference extraction sub-model. Accordingly, this embodiment can further refine: "Determining the video preference characteristics based on the multi-task processing model." Its specific implementation method can refer to the detailed explanation of this technical solution. Among them, technical terms that are the same as or corresponding to the aforementioned embodiments are not repeated in this embodiment.
[0129] It should be noted that the methods of determining video preference features and object preference features based on the multi-task processing model are the same. It is only necessary to replace the video attribute information with the object attribute information. In this embodiment, the determination of video preference features is taken as an example to illustrate.
[0130] like Figure 6 As shown, the method provided in this embodiment includes:
[0131] S510: Obtain a video single tag and a video multi-tag in the video attribute information.
[0132] The videos to be pushed often have certain tags, such as history tags, life tags, and so on. Each video can have multiple tags, with a unique tag specific to a video being used as a single video tag. For example, a video has a unique author, so the author can be used as a single video tag. A video's genre, theme, and other tags can be used as multiple video tags.
[0133] For example, see Figure 7To discover the target user's preferences for different videos, we can use one one-dimensional sequence and four two-dimensional sequences. One one-dimensional sequence can be a sequence of the creators of the videos to be pushed. This sequence primarily explores the similarities between the videos the target user has watched and the videos to be pushed, primarily from the creator's perspective. In this case, the single video tag primarily represents the tag of the creator of the video to be pushed. The other four two-dimensional sequences can be sequences composed of tags such as themes, interests, and genres. Based on this, we can determine single and multiple video tags.
[0134] S520. The first processing unit in the long-term preference extraction sub-model processes the video single label and the single label sequence corresponding to the video single label to determine the first label sequence, so as to determine the first video feature to be weighted of the corresponding historical video based on the first label sequence by the second processing unit.
[0135] Among them, the long-term preference extraction sub-model includes a first processing unit. The first processing unit is a processing unit for extracting video single labels and single label sequence features. The first label sequence includes label sequences corresponding to multiple watched videos. At this time, it can be based on user association information to determine the single label sequence composed of author information corresponding to the videos watched by the target user within a historical preset time period. The number of labels in the first label sequence is less than the number of labels in the single label sequence. The author sequence in the first label sequence is a label sequence composed of multiple authors with a high degree of adaptability to the target user after processing the single label sequence and the video single label. The second processing unit is used to determine the first video feature to be weighted of the historical video corresponding to the first label sequence. The feature of the historical video related to the first label sequence is used as the first video feature to be weighted.
[0136] It should also be noted that single-label video, multi-label video, single-label sequence and multi-label sequence are all information that can be directly obtained from user-related information.
[0137] For example, see Figure 7 After obtaining the video single label of the video to be pushed, a single label sequence can be constructed based on the author of the historical videos watched in the user association information. By matching the single label sequence and the video single label, a first label sequence that matches the video single label can be obtained. That is, after gsu single label matching, the similarity between each author label and the video single label in the single label sequence is obtained. The author labels with the top k similarities are selected to form the first label sequence. The first label sequence is sent to the second processing unit, and the second processing unit can obtain the historical videos corresponding to each author label in the first label sequence, and extract the video features of the historical videos as the first video features to be weighted.
[0138] S530. The third processing unit in the long-term preference extraction sub-model processes the video multi-label and the multi-label sequence corresponding to the video multi-label to obtain a second label sequence, so as to determine the second video feature to be weighted of the corresponding historical video based on the second label sequence by the fourth processing unit.
[0139] The third processing unit can be understood as a unit for processing multi-label videos. The multi-label sequence is a sequence of multi-labels corresponding to the historical videos watched by the target user within a preset historical duration. The second label sequence is a label sequence that ranks in the top k in terms of similarity to the multi-labels of the video. The fourth processing unit is a unit for extracting features from the historical videos associated with the second label sequence. The features obtained after extracting the features of the historical videos corresponding to the second label sequence are used as the second video features to be weighted.
[0140] Specifically, the third processing unit in the length preference extraction sub-model extracts the video multi-label and the multi-label sequence associated with the video multi-label, and performs similarity determination to obtain the labels in the multi-label sequence that are similar to the video multi-label and rank in the top k in terms of similarity, thereby forming a second label sequence. The fourth processing unit can extract historical videos associated with the second label sequence and extract the second to-be-weighted video features corresponding to the historical videos.
[0141] For example, see Figure 7 The third processing unit performs GSU multi-label intersection matching, and based on the matching results, selects the multi-labels with the top k similarities to form a second label sequence. The second label sequence is sent to the fourth unit, which determines the second to-be-weighted video feature corresponding to the historical video associated with the second label sequence.
[0142] S540 : Obtain a video preference feature by weighting the first video feature to be weighted and the second video feature to be weighted.
[0143] Specifically, the video preference feature corresponding to the target user is obtained by weighting the first video feature to be weighted and the second video feature to be weighted.
[0144] For object preference features, the single video tag can be replaced with a single object tag. Accordingly, the single tag sequence is a sequence of tags corresponding to the objects triggered by the target user within a preset historical duration. This same approach can be applied to multi-video, multi-tag, and multi-tag sequences to determine the object preference features corresponding to the target user.
[0145] Continue to see Figure 7After obtaining the video adaptation features, they can be fed into a deep neural network to output video matching attributes. These attributes represent the degree of match between the video to be pushed and the target user. Similarly, for the object adaptation features, the above steps are repeated to obtain the degree of match between the target object and the target user.
[0146] The technical solution provided by the embodiments of the present invention can extract the long-term video features of the target user relative to the video to be pushed, as well as the long-term object features of the target user relative to the target user. Based on the long-term video features, the long-term object features, the video attribute information of the video to be pushed, and the object attribute information of the target object, video preference features and object preference features are obtained through analysis and processing. The advantage of determining the video preference features and object preference features is that the preference attributes of the video to be pushed and the target object relative to the target user can be determined from multiple perspectives, thereby improving the accuracy of the subsequent determination of the target processing results corresponding to different task types.
[0147] Figure 8 A flowchart of a task processing method provided in an embodiment of the present invention can be further refined on the basis of the aforementioned embodiment: "Based on the feature processing module in the multi-task processing model, the first feature sequence, the second feature sequence, the first cross-feature, the second cross-feature, the third cross-feature, the video adaptation feature, the object adaptation feature, the video preference feature and the object preference feature are processed to output task processing results corresponding to different task types." The specific implementation method can be found in the detailed explanation of this embodiment, wherein the technical terms that are the same as or corresponding to the above-mentioned embodiment are not repeated in this embodiment.
[0148] like Figure 8 As shown, the method provided in this embodiment includes:
[0149] S610: When the features to be used include at least two of the features splicing sequence, the cross feature, the adaptation feature, and the preference feature, the features to be used are spliced to obtain a target splicing feature.
[0150] The target splicing feature is the feature obtained by splicing the features obtained by the above models. Optionally, the second feature sequence is spliced after the first feature sequence to obtain the first splicing feature. Based on the first splicing feature, the first cross feature, the second cross feature, the video adaptation feature, the object adaptation feature, the video preference attribute, and the object preference attribute are further spliced to obtain the target splicing feature.
[0151] Specifically, the first feature sequence, the second feature sequence, the first cross feature, the second cross feature, the third cross feature, the video adaptation feature, the object adaptation feature, the video preference attribute and the object preference attribute can be spliced in sequence according to a pre-set splicing order to obtain the target splicing feature.
[0152] For example, see Figure 2 , based on the feature splicing layer (contact layer), the features output by each sub-model are spliced to obtain the target splicing features.
[0153] In this embodiment, after obtaining the video preference features and the object preference features, the method further includes: processing the video preference features based on a deep neural network to obtain video preference attributes, and processing the object preference features based on a deep neural network to obtain object preference attributes.
[0154] Among them, the deep neural network can be any existing model that can analyze and process preference features. The input of the deep neural network can be a feature sequence, and the output can be the evaluation attribute value corresponding to the feature sequence. The higher the video preference attribute, the higher the compatibility between the video to be pushed and the target user, and the lower the video preference attribute, the lower the compatibility between the video to be pushed and the target user. The first model can be spliced with a deep neural network, and the second model can also be spliced with a deep neural network. The deep neural network spliced with the first model is mainly used to process video adaptation features. The deep neural network connected to the second model is mainly used to process object adaptation features. Accordingly, the higher the object preference attribute, the more compatible the target object is with the target user, that is, the higher the probability of the target user obtaining the target object. Conversely, the lower the object preference attribute, the less compatible the target object is with the target user, that is, the lower the probability of the target user obtaining the target object.
[0155] S620: Process the target splicing features based on the multi-task learning sub-model in the multi-task processing model, and output task processing results under different task types.
[0156] The number of multi-task learning sub-models can include one or more, and the specific number is related to the number of pre-set task types. The multi-task learning sub-model can be an existing expert network. For each expert network, it can learn the features processed by other expert networks and combine them with their corresponding features to determine the final task processing result.
[0157] This can be understood as determining a multi-task learning sub-model that matches the number of pre-set task types. Task types are pre-set task types to be processed. These optional task types include interactive behavior prediction, viewing duration prediction, and the probability of watching the next video after the current video. Users can set task types based on their actual needs.
[0158] The output of the interactive behavior prediction task type can include interactive behavior data, which can optionally include likes, comments, and other data. The output of the viewing duration prediction task type can include viewing probability information, based on which the viewing duration of the video to be pushed can be determined. The task type based on the probability of watching the next video from the current video can be based on the probability of sliding to watch the next video to be pushed.
[0159] It should be noted that its specific content may be related to the input data and constraint output results in the model training phase.
[0160] In this embodiment, after obtaining the target splicing features, the target splicing features can be input into the multi-task learning sub-model. The multi-task learning sub-model can analyze and process the target splicing features so that each task learning sub-model outputs a task processing result corresponding to its task type.
[0161] In this embodiment, after obtaining the task processing result, the method further includes: sending the task processing results of the to-be-pushed video under different task types to the target downstream system, so that the target downstream system uses the task processing result.
[0162] This means that after obtaining the task processing results corresponding to different task types for the video to be pushed, which has the target object attached, the task processing results corresponding to the different task types can be sent to the target downstream system. The target downstream system can then determine whether to push the video to the target user's system, or other systems that require the task processing results corresponding to different task types.
[0163] The technical solution provided by the embodiment of the present invention can process the target splicing features based on the multi-task learning sub-model, and can output the target processing results corresponding to different task types. At the same time, the target processing results can be sent to the target downstream system so that the target downstream system can improve the accuracy of downstream task processing based on more accurate target processing results.
[0164] The following is an embodiment of a task processing device provided by an embodiment of the present invention. The device and the task processing methods of the above embodiments belong to the same inventive concept. For details not fully described in the embodiment of the task processing device, please refer to the embodiment of the above task processing method.
[0165] Figure 9 This is a structural diagram of a task processing device provided by an embodiment of the present invention. The device specifically includes: an information acquisition module 710, a feature determination module 720 and a result output module 730.
[0166] The information acquisition module 710 is used to obtain the video attribute information of the video to be pushed with the target object mounted thereon, the object attribute information of the target object, and the user association information of the target user; wherein, the target user is the user to whom the video to be pushed is to be pushed; the feature determination module 720 is used to perform feature analysis processing on the video attribute information, the object attribute information, and the user association information based on the pre-trained multi-task processing model to obtain at least one feature to be used; wherein, the feature to be used includes a feature splicing sequence, a cross feature, an adaptation feature, or a preference feature, and the feature splicing sequence is used to characterize the video attribute information. , object attribute information and user association information feature extraction and splicing obtained, the cross feature is used to characterize the cross information between the video to be pushed, the target object and the target user, the adaptation feature is used to characterize the adaptation information of the target object and the video to be pushed relative to the target user, and the preference feature is used to characterize the target user's preference information for videos and objects; the result output module 730 is used to process the at least one feature to be used based on the task processing sub-model corresponding to different task types in the multi-task processing model, and output the task processing results corresponding to different task types.
[0167] On the basis of the above technical solutions, the device further includes:
[0168] A first information determination module is configured to determine, based on a first preset condition, first video information, first object information, and first user information for determining a feature sequence from the video attribute information, the object attribute information, and the user association information; and / or
[0169] a second information determining module, configured to obtain, based on a second preset condition, second video information, second object information, and second user information for determining the cross-feature from the video attribute information, the object attribute information, and the user association information; and / or
[0170] The third information determination module is used to obtain third video information, third object information, and third user information for determining adaptation features from the video attribute information and the user association information based on a third preset condition; wherein the third user information includes historical viewing behavior information of the target user within a historical preset time period, and historical interaction behavior information of the target user with the display object.
[0171] On the basis of the above technical solutions, the feature to be used corresponds to a feature splicing sequence, the feature splicing sequence includes a first feature sequence and / or a second feature sequence, and the feature determination module includes:
[0172] a first feature sequence determining unit, configured to normalize the first video information, the first object information, and the first user information based on a normalization module, and to concatenate the first video features, the first object features, and the first user features obtained after the normalization to obtain a first feature sequence; and / or
[0173] The second feature sequence determination unit is used to extract the first video information, the first object information and the first user information respectively based on the first feature extraction module to perform feature extraction, and splice the extracted second video features, second object features and second user features to obtain a second feature sequence.
[0174] On the basis of the above technical solutions, the features to be used correspond to cross features, and the cross features include a first cross feature between the target object and the video to be pushed, a second cross feature between the video to be pushed and the target object, and a third cross feature between the target object and the target user. The feature determination module is also used to
[0175] Based on the prediction sub-model in the multi-task processing model, cross-feature extraction is performed on the second video information and the second object information to obtain the first cross-feature, cross-feature extraction is performed on the second video information and the second user information to obtain the second cross-feature, and cross-feature extraction is performed on the second object information and the second user information to obtain the third cross-feature.
[0176] Based on the above technical solutions, the features to be used also include fusion features, and the multi-task processing model also includes a fusion sub-model, which is used to perform feature fusion on the second video information, the second object information and the second user information to obtain the fusion features.
[0177] Based on the above technical solutions, the features to be used correspond to adaptation features, the adaptation features include video adaptation features and object adaptation features, the multi-task processing model also includes a first model and a second model with the same model structure, and the feature determination module includes:
[0178] a video adaptation feature determination unit, configured to extract features of the third video information and the historical viewing behavior information in the third user information based on the first model, to obtain a video adaptation feature of the video to be pushed relative to the target user;
[0179] An object adaptation feature determination unit is used to extract features of the historical interaction behavior information in the third video information and the third user information based on the second model to obtain the object adaptation feature of the target object relative to the target user.
[0180] Based on the above technical solutions, the multi-task processing model includes a long-term preference extraction sub-model for determining the preference feature. The feature determination module includes:
[0181] A tag acquisition unit, configured to acquire a single video tag and multiple video tags from the video attribute information;
[0182] a label sequence determining unit, configured to process the video single label and the single label sequence corresponding to the video single label based on the first processing unit in the long-term preference extraction sub-model to determine a first label sequence, so as to determine, based on the first label sequence, a first to-be-weighted video feature of the corresponding historical video by the second processing unit; wherein the historical video is a video related to the first label sequence;
[0183] a feature determination unit configured to process the video multi-label and the multi-label sequence corresponding to the video multi-label based on the third processing unit in the long-term preference extraction sub-model to obtain a second label sequence, and determine a second to-be-weighted video feature of the corresponding historical video based on the second label sequence based on the fourth processing unit;
[0184] The preference determination unit is configured to obtain the video preference feature by weighting the first video feature to be weighted and the second video feature to be weighted.
[0185] Based on the above technical solutions, the feature determination module is further used to:
[0186] By processing the video preference feature and the object preference feature in the preference feature, a video preference attribute corresponding to the video preference feature and an object preference attribute corresponding to the object preference feature are obtained, so as to update the video preference feature based on the video preference attribute and update the object preference feature based on the object preference attribute.
[0187] Based on the above technical solutions, the result output module includes:
[0188] a feature splicing unit, configured to, when the features to be used include at least two of a feature splicing sequence, a cross feature, an adaptation feature, and a preference feature, splice the features to be used to obtain a target splicing feature;
[0189] A result output unit is used to process the target splicing features based on the multi-task learning sub-model in the multi-task processing model and output task processing results under different task types.
[0190] The technical solution provided by the embodiment of the present invention can input the acquired video attribute information of the video to be pushed with the target object mounted thereon, the object attribute information of the target object, and the user association information of the target user into a pre-trained multi-task processing model, and can determine multiple feature sequences, cross-features after information cross-processing, video adaptation features of the target user corresponding to the video to be pushed, object adaptation features of the target object relative to the target user, video preference features of the target user, and object preference features based on the multi-task processing model. Subsequently, the at least one feature is processed to obtain task processing results under different task types. Since the method provided by this embodiment can determine the features between the video and the user, the features between the user and the object, the features between the object and the video, the video adaptation features of the video relative to the user, the object adaptation features of the object relative to the user, and the target user's preference features and object preference features for the video, that is, the adaptation information of the video to be pushed and the target object relative to the target user is determined from multiple angles, thereby improving the accuracy and reliability of determining the target processing results, and solving the problem of low accuracy when determining the adaptability of the video or object relative to the target user from a single dimension in the prior art.
[0191] The task processing device provided by the embodiment of the present invention can execute the task processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the task processing method.
[0192] It is worth noting that in the embodiment of the above-mentioned task processing device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0193] Figure 10 A schematic diagram of the structure of a server provided in an embodiment of the present invention. Figure 10 A block diagram of an exemplary electronic device 12 suitable for implementing embodiments of the present invention is shown. Figure 10 The electronic device 12 shown is only an example and should not limit the functionality and scope of use of the embodiments of the present invention.
[0194] like Figure 10 As shown, electronic device 12 is implemented as a general-purpose computing device. Components of electronic device 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components (including system memory 28 and processing unit 16).
[0195] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0196] The electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0197] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be configured to read and write non-removable, non-volatile magnetic media ( Figure 10 Not shown, often called a "hard drive"). Although Figure 10Not shown, a magnetic disk drive for reading and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.
[0198] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods of the embodiments described herein.
[0199] The electronic device 12 can also communicate with one or more external devices 14 (e.g., a keyboard, pointing device, display 24, etc.), one or more devices that enable a user to interact with the electronic device 12, and / or any device that enables the electronic device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication can occur via an input / output (I / O) interface 22. Furthermore, the electronic device 12 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the electronic device 12 via a bus 18. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with the electronic device 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0200] The processing unit 16 executes various functional applications and task processing by running programs stored in the system memory 28, such as implementing the task processing method steps provided in the first embodiment of the present invention, which includes:
[0201] Obtaining video attribute information of a video to be pushed that is mounted with a target object, object attribute information of the target object, and user association information of a target user; wherein the target user is the user to whom the video to be pushed is to be pushed;
[0202] Based on a pre-trained multi-task processing model, feature analysis processing is performed on the video attribute information, the object attribute information, and the user association information to obtain at least one feature to be used; wherein, the feature to be used includes a feature splicing sequence, a cross feature, an adaptation feature, or a preference feature, the feature splicing sequence is used to characterize the features obtained by splicing the video attribute information, the object attribute information, and the user association information after feature extraction, the cross feature is used to characterize the cross information between any two of the video to be pushed, the target object, and the target user, the adaptation feature is used to characterize the adaptation information of the target object and the video to be pushed relative to the target user, and the preference feature is used to characterize the target user's preference information for the video and the object;
[0203] The at least one feature to be used is processed based on the task processing sub-models corresponding to different task types in the multi-task processing model, and task processing results corresponding to different task types are output.
[0204] Of course, those skilled in the art will appreciate that the processor may also implement the technical solution of the task processing method provided in any embodiment of the present invention.
[0205] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the task processing method provided in the aforementioned embodiment of the present invention are implemented. The method includes:
[0206] Obtaining video attribute information of a video to be pushed that is mounted with a target object, object attribute information of the target object, and user association information of a target user; wherein the target user is the user to whom the video to be pushed is to be pushed;
[0207] Based on a pre-trained multi-task processing model, feature analysis processing is performed on the video attribute information, the object attribute information, and the user association information to obtain at least one feature to be used; wherein, the feature to be used includes a feature splicing sequence, a cross feature, an adaptation feature, or a preference feature, the feature splicing sequence is used to characterize the features obtained by splicing the video attribute information, the object attribute information, and the user association information after feature extraction, the cross feature is used to characterize the cross information between any two of the video to be pushed, the target object, and the target user, the adaptation feature is used to characterize the adaptation information of the target object and the video to be pushed relative to the target user, and the preference feature is used to characterize the target user's preference information for the video and the object;
[0208] The at least one feature to be used is processed based on the task processing sub-models corresponding to different task types in the multi-task processing model, and task processing results corresponding to different task types are output.
[0209] The computer storage medium of the embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.
[0210] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0211] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0212] Computer program code for performing the operations of embodiments of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0213] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will appreciate that the present invention is not limited to the specific embodiments herein, and that various obvious changes, readjustments, and substitutions are possible for those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present invention. The scope of the present invention is determined by the scope of the appended claims.
Claims
1. A task processing method, characterized in that: include: Obtaining video attribute information of a video to be pushed that is mounted with a target object, object attribute information of the target object, and user association information of a target user; wherein the target user is the user to whom the video to be pushed is to be pushed; Based on a pre-trained multi-task processing model, feature analysis processing is performed on the video attribute information, the object attribute information, and the user association information to obtain at least one feature to be used; wherein, the feature to be used includes a feature splicing sequence, a cross feature, an adaptation feature, or a preference feature, the feature splicing sequence is used to characterize the features obtained by splicing the video attribute information, the object attribute information, and the user association information after feature extraction, the cross feature is used to characterize the cross information between any two of the video to be pushed, the target object, and the target user, the adaptation feature is used to characterize the adaptation information of the target object and the video to be pushed relative to the target user, and the preference feature is used to characterize the target user's preference information for the video and the object; The at least one feature to be used is processed based on the task processing sub-models corresponding to different task types in the multi-task processing model, and task processing results corresponding to different task types are output.
2. The method according to claim 1, characterized in that The method further comprises: Based on a first preset condition, determining first video information, first object information, and first user information for determining a feature sequence from the video attribute information, the object attribute information, and the user association information; and / or, Based on a second preset condition, obtaining second video information, second object information, and second user information for determining the cross feature from the video attribute information, the object attribute information, and the user association information; and / or, Based on a third preset condition, third video information, third object information, and third user information for determining adaptation features are obtained from the video attribute information and the user association information; wherein the third user information includes historical viewing behavior information of the target user within a historical preset time period, and historical interaction behavior information of the target user with the displayed object.
3. The method according to claim 2, characterized in that The feature to be used corresponds to a feature splicing sequence, the feature splicing sequence includes a first feature sequence and / or a second feature sequence, and the multi-task processing model based on pre-training performs feature analysis processing on the video attribute information, the object attribute information, and the user association information to obtain at least one feature to be used, including: Normalizing the first video information, the first object information, and the first user information based on a normalization module, and concatenating the first video features, the first object features, and the first user features obtained after the normalization to obtain a first feature sequence; and / or, Based on the first feature extraction module, the first video information, the first object information and the first user information are respectively extracted for feature extraction, and the extracted second video features, second object features and second user features are spliced to obtain a second feature sequence.
4. The method according to claim 2, characterized in that The feature to be used corresponds to an intersection feature, and the intersection feature includes a first intersection feature between the target object and the video to be pushed, a second intersection feature between the video to be pushed and the target object, and a third intersection feature between the target object and the target user. The multi-task processing model based on pre-training performs feature analysis processing on the video attribute information, the object attribute information, and the user association information to obtain at least one feature to be used, including: Based on the prediction sub-model in the multi-task processing model, cross-feature extraction is performed on the second video information and the second object information to obtain the first cross-feature, cross-feature extraction is performed on the second video information and the second user information to obtain the second cross-feature, and cross-feature extraction is performed on the second object information and the second user information to obtain the third cross-feature.
5. The method according to claim 4, characterized in that The features to be used also include fusion features, and the multi-task processing model also includes a fusion sub-model, which is used to perform feature fusion on the second video information, the second object information and the second user information to obtain the fusion features.
6. The method according to claim 2, characterized in that The feature to be used corresponds to an adaptation feature, the adaptation feature includes a video adaptation feature and an object adaptation feature, the multi-task processing model also includes a first model and a second model with the same model structure, and the multi-task processing model based on pre-training performs feature analysis processing on the video attribute information, the object attribute information, and the user association information to obtain at least one feature to be used, including: Performing feature extraction on the third video information and the historical viewing behavior information in the third user information based on the first model to obtain a video adaptation feature of the video to be pushed relative to the target user; Based on the second model, feature extraction is performed on the historical interaction behavior information in the third video information and the third user information to obtain an object adaptation feature of the target object relative to the target user.
7. The method according to claim 1, characterized in that The multi-task processing model includes a long-term preference extraction sub-model for determining the preference feature, and determining the preference feature among the features to be used includes: Obtaining the video single tag and video multi-tag in the video attribute information; The first processing unit in the long-term preference extraction sub-model processes the video single label and the single label sequence corresponding to the video single label to determine a first label sequence, so as to determine a first to-be-weighted video feature of the corresponding historical video based on the first label sequence by a second processing unit; wherein the historical video is a video related to the first label sequence; The third processing unit in the long-term preference extraction sub-model processes the video multi-label and the multi-label sequence corresponding to the video multi-label to obtain a second label sequence, and determines a second to-be-weighted video feature of the corresponding historical video based on the second label sequence by the fourth processing unit; The video preference feature is obtained by weighting the first video feature to be weighted and the second video feature to be weighted.
8. The method according to claim 7, characterized in that The method further comprises: By processing the video preference feature and the object preference feature in the preference feature, a video preference attribute corresponding to the video preference feature and an object preference attribute corresponding to the object preference feature are obtained, so as to update the video preference feature based on the video preference attribute and update the object preference feature based on the object preference attribute.
9. The method according to claim 6, characterized in that The processing of the at least one feature to be used based on the task processing sub-models corresponding to different task types in the multi-task processing model and outputting task processing results corresponding to different task types includes: In a case where the features to be used include at least two of a feature splicing sequence, a cross feature, an adaptation feature, and a preference feature, splicing the features to be used to obtain a target splicing feature; The target splicing features are processed based on the multi-task learning sub-model in the multi-task processing model, and task processing results under different task types are output.
10. A task processing device, characterized in that: include: An information acquisition module is used to obtain video attribute information of a video to be pushed that is mounted with a target object, object attribute information of the target object, and user association information of a target user; wherein the target user is the user to whom the video to be pushed is to be pushed; a feature determination module, configured to perform feature analysis on the video attribute information, the object attribute information, and the user association information based on a pre-trained multi-task processing model to obtain at least one feature to be used; wherein the feature to be used includes a feature splicing sequence, a cross feature, an adaptation feature, or a preference feature; the feature splicing sequence is used to characterize the features obtained by splicing the video attribute information, the object attribute information, and the user association information after feature extraction; the cross feature is used to characterize the cross information between any two of the video to be pushed, the target object, and the target user; the adaptation feature is used to characterize the adaptation information of the target object and the video to be pushed relative to the target user; and the preference feature is used to characterize the target user's preference information for the video and the object; The result output module is used to process the at least one feature to be used based on the task processing sub-models corresponding to different task types in the multi-task processing model, and output task processing results corresponding to different task types.
11. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the task processing method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the task processing method according to any one of claims 1 to 9 is implemented.
13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the task processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Information pushing method, device and system
CN107517393A
Video recommendation method based on multi-modal video content and multi-task learning
CN111246256A
Push processing method, related device and medium
CN117349791A
Video recommendation method, electronic device, and storage medium
US20250013691A1
Video recommendation method and apparatus, and electronic device, computer-readable storage medium and computer program product
WO2024113641A1