Interactive processing method and device suitable for multiple scenes and multiple tasks, medium and equipment
By using feature fusion and interaction prediction processing, the problem of resource recommendation in multiple scenarios and tasks is solved, improving user experience and recommendation performance.
Patent Information
- Application Number
- CN202410474718.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-19
- Publication Date
- 2025-10-24
AI Technical Summary
In internet applications, existing technologies struggle to optimize resource recommendation models for multiple scenarios and tasks simultaneously, leading to poor user experience and ecosystem health issues.
By acquiring feature data of the target object, candidate resources, and multiple scenarios, feature fusion and extraction are performed. Expert networks and tower networks are used to predict the interaction between the task layer and the scenario layer, predicting the probability of interaction events between the target object and candidate resources.
It enables resource recommendation across multiple scenarios and tasks, improving recommendation effectiveness and user experience, and optimizing satisfaction with resource recommendations.
Smart Images

Figure CN120832435A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to an interaction processing method and device suitable for multiple scenes and multiple tasks, a medium and equipment. BACKGROUND
[0002] Artificial intelligence (AI) is a comprehensive technology of computer science, which makes machines have the functions of perception, reasoning and decision-making by studying the design principles and implementation methods of various intelligent machines. Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, such as natural language processing, machine learning, deep learning, etc. With the development of technology, artificial intelligence technology has been widely applied in the field of resource recommendation and has played an increasingly important role.
[0003] There are many recommendation scenarios in Internet applications, such as the selected page and each channel page in the video application, and each page is divided into different regions, such as the strong purpose region and the purposeless region under the selected page. The preferences and behavior habits of users in different scenarios are different, resulting in a large difference in the distribution of interaction indicators in different scenarios. In order to better improve the user experience and ensure the ecological health of the application, the interaction indicators that the recommendation model needs to optimize are usually not only one, and how to model these numerous scenes and multiple optimization tasks at the same time is a challenge in the field of resource recommendation. SUMMARY
[0004] In order to realize resource recommendation in multiple scenes and multiple tasks, the present application provides an interaction processing method and device suitable for multiple scenes and multiple tasks, a medium and equipment. The technical solution is as follows:
[0005] In a first aspect, the present application provides an interaction processing method suitable for multiple scenes and multiple tasks, which comprises:
[0006] obtaining original object feature data of a target object, original resource feature data of a candidate resource, and original scene feature data of each scene in multiple scenes;
[0007] fusing the original object feature data and the original resource feature data with the original scene feature data of each scene respectively to obtain first object feature data corresponding to each scene and first resource feature data corresponding to each scene;
[0008] performing resource feature extraction processing according to the original scene feature data of each scene and the original resource feature data to obtain second resource feature data corresponding to each scene;
[0009] input the first output result corresponding to each task in each scene into a tower network corresponding to each scene and each task, perform scene layer interaction prediction processing, obtain a first interaction prediction result corresponding to each task in each scene, and the first interaction prediction result represents the possibility of an interaction event indicated by the corresponding task occurring between the target object and the candidate resource in the corresponding scene.
[0010] input the first output result corresponding to each task in each scene into a tower network corresponding to each scene and each task, perform scene layer interaction prediction processing, obtain a first interaction prediction result corresponding to each task in each scene, and the first interaction prediction result represents the possibility of an interaction event indicated by the corresponding task occurring between the target object and the candidate resource in the corresponding scene.
[0011] In a second aspect, the present application provides an interaction processing device suitable for multiple scenes and multiple tasks, and the device comprises:
[0012] an acquisition module configured to acquire original object feature data of a target object, original resource feature data of a candidate resource, and original scene feature data of each scene in multiple scenes;
[0013] a first feature representation module configured to fuse the original object feature data and the original resource feature data with the original scene feature data of each scene, respectively, to obtain first object feature data corresponding to each scene and first resource feature data corresponding to each scene;
[0014] a second feature representation module configured to perform resource feature extraction processing according to the original scene feature data of each scene and the original resource feature data, to obtain second resource feature data corresponding to each scene;
[0015] a first feature extraction module configured to input first feature data corresponding to each scene into an expert network corresponding to each task in multiple tasks, perform task layer feature extraction processing, and obtain a first output result corresponding to each task in each scene; the first feature data corresponding to each scene is obtained by stacking the first object feature data corresponding to each scene, the first resource feature data corresponding to each scene, and the second resource feature data corresponding to each scene;
[0016] The first interaction prediction module is configured to input the first output result corresponding to each task in each scene into the tower network corresponding to each scene and each task, perform scene-layer interaction prediction processing, and obtain a first interaction prediction result corresponding to each task in each scene, the first interaction prediction result representing a possibility of an interaction event corresponding to each task between the target object and the candidate resource in the corresponding scene.
[0017] In a third aspect, the present application provides a computer readable storage medium, the computer readable storage medium storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by a processor to implement the interaction processing method for multiple scenes and multiple tasks according to the first aspect.
[0018] In a fourth aspect, the present application provides a computer device, the computer device comprising a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the interaction processing method for multiple scenes and multiple tasks according to the first aspect.
[0019] In a fifth aspect, the present application provides a computer program product, the computer program product comprising computer instructions, the computer instructions being executed by a processor to implement the interaction processing method for multiple scenes and multiple tasks according to the first aspect.
[0020] The interaction processing method for multiple scenes and multiple tasks, the device, the medium and the equipment provided by the present application have the following technical effects:
[0021] In the scheme provided in the application, by obtaining the original object feature data of the target object, the original resource feature data of the candidate resource and the original scene feature data of each scene in the plurality of scenes, and fusing the original object feature data and the original resource feature data with the original scene feature data of each scene respectively, the first object feature data corresponding to each scene and the first resource feature data corresponding to each scene can be obtained, wherein the first object feature data and the first resource feature data both incorporate the original scene feature data of the corresponding scene, that is, the perception of the scene is improved in the feature representation stage, and the different feature performances of the target object and the candidate resource in different scenes can be embodied. In the scheme provided in the application, the resource feature extraction processing is performed according to the original scene feature data and the original resource feature data of each scene, and the second resource feature data corresponding to each scene is obtained, wherein the second resource feature data is extracted from the original resource feature data under the influence of the original scene feature data, that is, the perception of the scene is improved in the feature representation stage, and the preference degree of the target object to the same candidate resource in different scenes can be embodied. In the scheme provided in the application, the first feature data corresponding to each scene is input into the expert network corresponding to each task in the plurality of tasks for task layer feature extraction processing, and the first output result corresponding to each task in each scene is obtained. The first feature data corresponding to each scene is obtained by stacking the first object feature data corresponding to each scene, the first resource feature data corresponding to each scene and the second resource feature data corresponding to each scene, and through the expert network corresponding to each task, more effective feature extraction of the first feature data can be performed from the task dimension, and parallel processing of different tasks is realized. In the scheme provided in the application, the first output result corresponding to each task in each scene is input into the tower network corresponding to each task and each scene for scene layer interaction prediction processing, and the first interaction prediction result corresponding to each task in each scene is obtained, wherein the first interaction prediction result represents the possibility of the corresponding interaction event between the target object and the candidate resource in the corresponding scene indicated by the corresponding task, and through the tower network corresponding to each task and each scene, the interaction prediction processing corresponding to each task from the scene dimension can be further performed on the basis of the foregoing, and parallel processing of different scenes is realized.
[0022] The scheme provided in the application can simultaneously perform interaction prediction processing of a plurality of scenes and a plurality of tasks on the target object and the candidate resource, and the perception of the scene is strengthened in the feature representation stage, so that the effect of resource recommendation can be more comprehensively and accurately estimated, the satisfaction of resource recommendation is improved, and the use experience of the user is optimized.
[0023] Additional aspects and advantages of the application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, and the advantages thereof, the drawings needed to be used in the description of the embodiments or the prior art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0025] Figure 1 is an implementation environment schematic diagram of a multi-scene multi-task interaction processing method provided by an embodiment of the present application;
[0026] Figure 2 is a flow schematic diagram of a multi-scene multi-task interaction processing method provided by an embodiment of the present application;
[0027] Figure 3 is a structural schematic diagram of a feature embedding representation network, a target attention network and a plurality of expert networks in an artificial intelligence model provided by an embodiment of the present application;
[0028] Figure 4 is a structural schematic diagram of a plurality of tower networks in an artificial intelligence model provided by an embodiment of the present application;
[0029] Figure 5 is an application schematic diagram provided by an embodiment of the present application;
[0030] Figure 6 is a schematic diagram of an interaction processing device suitable for multi-scene multi-task provided by an embodiment of the present application;
[0031] Figure 7 is a hardware structure schematic diagram of a device for implementing a multi-scene multi-task interaction processing method provided by an embodiment of the present application. DETAILED DESCRIPTION
[0032] Artificial Intelligence (AI) is the theory, method, technology and application system that use digital computer or digital computer controlled machine to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence, and produces a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology is a comprehensive discipline, which involves a wide range of fields, including hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc.
[0033] The scheme provided by the embodiments of the present application relates to deep learning (DL) and other technologies of artificial intelligence.
[0034] Deep learning (DL) is a major research direction in the field of machine learning (ML), which is introduced into machine learning to make it closer to the original goal-artificial intelligence. Deep learning is to learn the internal rules and representation levels of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images and sounds. The ultimate goal of deep learning is to enable machines to have analysis and learning ability like people, and to recognize text, images and sound data. Deep learning is a complex machine learning algorithm, which has achieved much better results in speech and image recognition than previous related technologies. Deep learning has achieved many results in search technology, data mining, machine learning, machine translation, natural language processing, multimedia learning, speech, recommendation and personalization technology, and other related fields. Deep learning enables machines to imitate human activities such as vision and thinking, solves many complex pattern recognition problems, and makes great progress in artificial intelligence related technologies.
[0035] The scheme provided by the embodiments of the present application can be deployed in the cloud, which also relates to cloud technology and the like.
[0036] Cloud technology: refers to the series of resources such as hardware, software and network in the wide area network or local area network are unified, realize the data calculation, storage, processing and sharing of a kind of hosting technology, also can be understood as the network technology, information technology, integration technology, management platform technology and application technology based on cloud computing business model application, can constitute resource pool, use as needed, flexible and convenient. The background service of technology network system needs a lot of computing and storage resources, such as video website, picture website and more portal website, with the high development and application of internet industry, every item may have its own identification mark in the future, which needs to be transmitted to the background system for logical processing. Different levels of data will be processed separately, and various industry data need strong system support, so cloud technology needs to be supported by cloud computing. Cloud computing is a computing mode, which distributes computing tasks on a large number of computing resources to form a resource pool, so that various application systems can obtain computing power, storage space and information service according to needs. The network providing resources is called "cloud". The resources in the "cloud" can be infinitely expanded in the eyes of the user, and can be obtained at any time, used on demand, expanded at any time, and paid according to use. As a basic ability provider of cloud computing, cloud platform will be established, which is generally called infrastructure as a service (IaaS). A variety of types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing device (virtual machine including operating system), storage device and network device.
[0037] The embodiment of the application provides an interactive processing method and device suitable for multiple scenes and multiple tasks, a medium and equipment. The technical solutions in the embodiments of the application will be described clearly and completely in combination with the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor belong to the scope of protection of the application. The examples of the embodiments are shown in the drawings, wherein the same or similar reference signs represent the same or similar elements or elements with the same or similar functions throughout.
[0038] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and in the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0039] It can be understood that in the specific embodiments of the present application, data related to object feature data and the like is involved, and when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use and processing of related data needs to comply with relevant laws, regulations and standards of relevant countries and regions.
[0040] Please refer to Figure 1 , which is an implementation environment diagram of an interactive processing method suitable for multi-scene multi-task provided by an embodiment of the present application, as shown in Figure 1 , the implementation environment can at least include a client 01 and a server 02.
[0041] Specifically, the client 01 can include devices such as smart phones, desktop computers, tablet computers, notebook computers, vehicle-mounted terminals, digital assistants, smart wearable devices and voice interaction devices, and can also include software running in the devices, such as web pages provided by some service providers to users, and can also be applications provided by the service providers to the users. Specifically, the client 01 can be used to send a recommendation request for a resource to the server 02.
[0042] Specifically, the server 02 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The server 02 can include a network communication unit, a processor, a memory, and the like. The terminal and the server can be connected directly or indirectly through wired or wireless communication, which is not limited in the present application. Specifically, the server 02 can be configured to, in response to a recommendation request for a resource sent by the client 01, obtain original object feature data of a target object, original resource feature data of a candidate resource, and original scene feature data of each scene in a plurality of scenes, wherein the target object corresponds to an account logged in on the client 01; fuse the original object feature data and the original resource feature data with the original scene feature data of each scene, respectively, to obtain first object feature data corresponding to each scene and first resource feature data corresponding to each scene; perform resource feature extraction processing according to the original scene feature data and the original resource feature data of each scene to obtain second resource feature data corresponding to each scene; input the first feature data corresponding to each scene into an expert network corresponding to each task in a plurality of tasks to perform task layer feature extraction processing and obtain a first output result corresponding to each task in each scene; the first feature data corresponding to each scene is obtained by stacking the first object feature data corresponding to each scene, the first resource feature data corresponding to each scene, and the second resource feature data corresponding to each scene; input the first output result corresponding to each task in each scene into a tower network corresponding to each scene and each task to perform scene layer interaction prediction processing and obtain a first interaction prediction result corresponding to each task in each scene, wherein the first interaction prediction result represents the possibility of an interaction event corresponding to the target object and the candidate resource in the corresponding scene.
[0043] The embodiment of the application can also be implemented in combination with cloud technology. Cloud technology refers to a kind of hosting technology that unifies a series of resources such as hardware, software and network to realize the calculation, storage, processing and sharing of data in a wide area network or local area network, and can also be understood as a general term of network technology, information technology, integration technology, management platform technology and application technology based on the cloud computing business model application. Cloud technology needs to be supported by cloud computing. Cloud computing is a computing mode that distributes computing tasks on a resource pool composed of a large number of computers, so that various application systems can obtain computing power, storage space and information services according to needs. The network that provides resources is called "cloud". Specifically, the server 02 and the database are located in the cloud, and the server 02 can be a physical machine or a virtual machine.
[0044] The following introduces an interactive processing method suitable for multiple scenes and multiple tasks provided by the application. Figure 2 The embodiment of the application provides a flowchart of an interactive processing method suitable for multiple scenes and multiple tasks, and the application provides the method operation steps as described in the embodiment or the flowchart, but more or fewer operation steps can be included based on conventional or non-creative labor. The order of steps listed in the embodiment is only one of the many execution orders, and does not represent the only execution order. When the system or server product is executed in practice, it can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment) according to the method order shown in the embodiment or the drawing. Please refer to Figure 2 The interactive processing method suitable for multiple scenes and multiple tasks provided by the embodiment of the application can include the following steps:
[0045] S210: obtaining original object feature data of a target object, original resource feature data of a candidate resource, and original scene feature data of each scene in multiple scenes.
[0046] In the embodiment of the application, the original object feature data of the target object is an embedding representation of the object attribute information of the target object, and the object attribute information of the target object can include but is not limited to object identification information, object identity information, object preference information, etc.; the original resource feature data of the candidate resource is an embedding representation of the resource attribute information of the candidate resource, and the resource attribute information of the candidate resource can include but is not limited to resource identification information, resource type information, resource content information, resource popularity information, etc. In the recommendation field, the resource can be a video, an article, news information, an advertisement, etc., and the target object is an object to which the resource is recommended.
[0047] In the embodiment of the application, the scene is a recommendation scene that provides an interactive function for objects and resources. There are many recommendation scenes in Internet applications, for example, Figure 5As shown, the video application can be divided into mobile terminal video application, computer terminal video application, web terminal video application, digital television terminal (Over The Top, OTT) application and the like according to terminal, and each type of video application has a selected page and a channel page, and each page can be divided into different areas according to business needs, for example, the selected page can be divided into a strong purpose area and a non-purpose area, the strong purpose area is a content display area for stimulating users to generate more interactive behaviors or for business promotion, and the strong purpose area can include but is not limited to a header picture module, a drama chasing module, a heavy module, a recently watched module and the like, wherein the header picture module and the heavy module can be used for distribution of new hot content, the drama chasing module and the recently watched module can be used for distribution of continued watching of medium and long videos, and the strong purpose area can mainly distribute content in the form of graphic materials; the non-purpose area is a content display area for meeting user viewing interests, and the non-purpose area is used for displaying different heterogeneous content such as long videos, medium videos, film lists, lists, live broadcasts, short videos and the like, and the non-purpose area can mainly distribute content in the form of video materials. The recommendation scene in the video application can be divided into multiple levels according to the application end type, the page type, the area type or the content module type, and the preferences and behavior habits of users in different scenes are different, thereby causing a large difference in the distribution of interactive indexes in different scenes.
[0048] In the embodiment of the present application, the interaction possibility of the target object and the candidate resource in multiple scenes and multiple tasks is predicted to determine whether to recommend the candidate resource to the target object, wherein the task is a prediction task for the interaction index or an optimization task for the interaction index.
[0049] S220: The original object feature data and the original resource feature data are fused with the original scene feature data of each scene respectively to obtain first object feature data corresponding to each scene and first resource feature data corresponding to each scene.
[0050] In the embodiment of the present application, the original object feature data and the original resource feature data are further represented by features, including being fused with the original scene feature data of each scene, so that the obtained first object feature data is integrated with the scene feature of the corresponding scene, and the obtained first resource feature data is integrated with the scene feature of the corresponding scene. That is, the perception of the scene is improved in the feature representation stage, and the different feature performances of the target object and the candidate resource in different scenes can be embodied.
[0051] In an embodiment of the present application, step S220 can include the following steps:
[0052] S221: The original scene feature data of each scene is input into a first gating network for feature transformation processing to obtain first scene feature data of each scene.
[0053] The gated neural network (which can be referred to as Gate NN for short) is a shallow neural network module that can selectively allow or block the passage of data. The first gated neural network is used to transform the original scene feature data of each scene, so that the obtained first scene feature data is more accurate and effective.
[0054] S222: Perform Hadamard product-based feature fusion processing on the first scene feature data and the original object feature data of each scene to obtain first object feature data corresponding to each scene.
[0055] Specifically, the first scene feature data of each scene is A s = {a ij}, and the original object feature data is B = {b ij}, where s is a scene identifier, A s and B are both m x n matrices, i.e., i≤m, j≤n. The Hadamard product-based feature fusion processing is performed on the first scene feature data and the original object feature data of each scene, i.e., the operation shown in formula (1) is performed:
[0056]
[0057] where C s is the first object feature data corresponding to the scene s.
[0058] The Hadamard product-based feature fusion processing fuses the original scene feature data of the corresponding scene into the first object feature data corresponding to each scene, and the first object feature data corresponding to different scenes also indicates the importance of the original object feature data in different scenes.
[0059] In some possible implementations, the first scene feature data and the original object feature data of each scene can also be subjected to weighted sum-based feature fusion processing to obtain first object feature data corresponding to each scene. The embodiments of the present application do not limit the applicable feature fusion method, and the above is only an example.
[0060] S223: Perform Hadamard product-based feature fusion processing on the first scene feature data and the original resource feature data of each scene to obtain first resource feature data corresponding to each scene.
[0061] The feature fusion processing of the first scene feature data and the original resource feature data of each scene can be referred to the foregoing, and will not be repeated here.
[0062] The first resource feature data corresponding to each scene obtained through the feature fusion processing based on the Hadamard product fuses the original scene feature data of the corresponding scene, and the first resource feature data corresponding to different scenes also indicates the importance of the original resource feature data in different scenes.
[0063] In the above embodiment, the perception of the scene is improved in the feature representation stage, the original scene feature data of each scene is fused with the original object feature data of the target object, and the original scene feature data of each scene is fused with the original resource feature data of the candidate resource, which can reflect the different feature performances of the target object and the candidate resource in different scenes, and improve the control degree of the scene information on the whole interaction processing.
[0064] S230: performing resource feature extraction processing according to the original scene feature data and the original resource feature data of each scene to obtain second resource feature data corresponding to each scene.
[0065] In the embodiment of the present application, the more critical and core second resource feature data is obtained through the resource feature extraction processing on the original resource feature data, and the second resource feature data is extracted from the original resource feature data under the influence of the original scene feature data, so that the second resource feature data corresponding to each scene can reflect the preference degree of the target object to the same candidate resource in different scenes, and the perception of the scene is further improved in the feature representation stage.
[0066] In an embodiment of the present application, the above-mentioned resource feature extraction processing is implemented based on an attention network. Specifically, step S230 can include the following steps:
[0067] S231: determining first parameter information according to the original scene feature data and the original resource feature data of each scene.
[0068] Feasibly, the original scene feature data and the original resource feature data of each scene are spliced, and the obtained feature data is a query in the form of a vector in the initial attention network, and then the first parameter information Q can be obtained by combining the first weight information corresponding to the query.
[0069] S232: obtaining historical interaction resource sequence data corresponding to the target object, and determining second parameter information and third parameter information according to the historical interaction resource sequence data, the historical interaction resource sequence data indicating at least one historical resource interacting with the target object.
[0070] The historical interaction resource sequence data includes resource feature data corresponding to at least one historical resource interacting with the target object, and the historical interaction resource sequence data is used as at least one key and at least one value of the initial attention network. Then, according to the at least one key and second weight information corresponding to the at least one key, the second parameter information K can be determined, and according to the at least one value and third weight information corresponding to the at least one value, the third parameter information V can be determined.
[0071] S233: updating the third parameter information according to the similarity between the first parameter information and the second parameter information, to obtain fourth parameter information.
[0072] Feasibly, the similarity between the first parameter information Q and the second parameter information K is calculated, such as calculating the vector space distance, to obtain similarity data, and the similarity data is used as new weight information of the third parameter information V. Then, the fourth parameter information V' can be obtained by combining the third parameter information V and the new weight information. Since the original scene feature data is introduced in the query, the accuracy and scene adaptation of the similarity data can be improved.
[0073] S234: determining a target attention network based on the first parameter information, the second parameter information and the fourth parameter information.
[0074] S234: inputting the original resource feature data into the target attention network for feature extraction processing to obtain second resource feature data corresponding to each scene.
[0075] That is, the initial attention network can be updated to the target attention network corresponding to the current query according to the first parameter information Q, the second parameter information K and the fourth parameter information V', so as to better extract the key and core features of the original resource feature data in each scene, and to reflect the preference degree of the target object to the same candidate resource in different scenes through the second resource feature data.
[0076] The target attention network is mainly realized by an encoder-decoder network model to perform the above feature extraction processing.
[0077] In the above embodiment, the original resource feature data is extracted by the attention network to obtain more key and core second resource feature data, and the second resource feature data is extracted from the original resource feature data under the influence of the original scene feature data. The second resource feature data corresponding to each scene can reflect the preference degree of the target object to the same candidate resource in different scenes, further improving the perception of the scene in the feature representation stage and enhancing the control degree of the scene information on the whole interaction process.
[0078] S240: input the first feature data corresponding to each scene into the expert network corresponding to each task in the plurality of tasks to perform feature extraction processing at the task layer, and obtain a first output result corresponding to each task under each scene; the first feature data corresponding to each scene is obtained by stacking the first object feature data corresponding to each scene, the first resource feature data corresponding to each scene, and the second resource feature data corresponding to each scene.
[0079] In the embodiment of the present application, the interaction prediction processing of multiple tasks is realized through the expert network corresponding to each task and the tower network corresponding to each task, such as simultaneously performing prediction processing of three interaction indexes of whether to play a video, video playing time, and video playing times, so as to comprehensively improve the interaction effect of video playing.
[0080] In the embodiment of the present application, the expert network corresponding to each task includes a shared expert network module, a specific expert network module corresponding to each task, and a second gating network corresponding to each task. The shared expert network module in the expert network corresponding to each task is a same network module, and is used to learn the commonality between multiple tasks. The specific expert network module in the expert network corresponding to each task has different parameters, and is used to learn the characteristics of the corresponding task. In the case of N tasks, the entire expert network layer includes N+1 expert network modules in total, of which 1 is a shared expert network module and N are specific expert network modules. In addition, the entire expert network layer further includes N second gating networks, and the parameters of the N second gating networks can also be different. The parameters of the second gating network include the output weight of the shared expert network module in the same expert network and the output weight of the specific expert network module in the same expert network.
[0081] In an embodiment of the present application, step S240 can include the following steps:
[0082] S241: input the first feature data corresponding to each scene into the shared expert network module in the expert network corresponding to each task to perform feature extraction processing for the commonality of multiple tasks, and obtain a first extraction result corresponding to each task under each scene.
[0083] That is, based on the learning experience of the commonality between multiple tasks, the first feature data corresponding to each scene is subjected to feature extraction, and the obtained first extraction result contains feature data associated with the commonality of multiple tasks.
[0084] S242: input the first feature data corresponding to each scene into the specific expert network module in the expert network corresponding to each task to perform feature extraction processing for the characteristics of each task, and obtain a second extraction result corresponding to each task under each scene.
[0085] That is, based on the learning experience of the respective characteristics of each task, the first feature data corresponding to each scene is feature extracted, and the obtained second extraction result contains feature data associated with the respective characteristics of each task.
[0086] S243: input the first extraction result corresponding to each task in each scene and the second extraction result corresponding to each task in each scene into the second gating network corresponding to each task, perform weighted processing, and obtain the third extraction result corresponding to each task in each scene.
[0087] In the embodiments of the present application, the second gating network is used for weighted sum processing of multiple inputs, that is, the product of the first extraction result and the output weight corresponding to the shared expert network module is added to the product of the second extraction result and the output weight corresponding to the specific expert network module, and the third extraction result corresponding to each task in each scene can be obtained.
[0088] S244: add the third extraction result corresponding to each task in each scene and the first feature data corresponding to each scene to obtain the first output result corresponding to each task in each scene.
[0089] Adding the third extraction result corresponding to each task in each scene and the first feature data corresponding to each scene can reduce the fading and loss of scene information, thereby strengthening the perception of scene information and improving the control of scene information on the entire interaction prediction processing process.
[0090] In the above embodiments, the expert network layer is constructed based on the shared expert network module and the specific expert network module corresponding to the task, so that the first extraction result containing feature data associated with the commonality of multiple tasks and the second extraction result containing feature data associated with the respective characteristics of each task can be extracted from the first feature data corresponding to each scene. At the same time, the data obtained by fusing the first extraction result and the second extraction result is added to the first feature data, which further enhances the perception of scene information.
[0091] S250: input the first output result corresponding to each task in each scene into the tower network corresponding to each scene and each task to perform scene layer interaction prediction processing, and obtain the first interaction prediction result corresponding to each task in each scene. The first interaction prediction result represents the possibility of the corresponding interaction event between the target object and the candidate resource indicated by the corresponding task in the corresponding scene.
[0092] In the embodiments of the present application, a single task corresponds to a tower network set, the number of tower networks in a single tower network set is consistent with the number of scenes, a single tower network corresponds to a single scene in addition to a single task, an interaction index prediction layer is constructed based on the tower network corresponding to each scene and each task, and interaction prediction processing of different task indexes in different scenes can be implemented. For example, under the premise of two scenes and three tasks, the interaction possibility between a single target object and a single candidate resource is predicted, and six interaction index prediction values can be obtained, which can more comprehensively and comprehensively measure the interaction between the target object and the candidate resource.
[0093] In an embodiment of the present application, the tower network corresponding to each scene and each task includes a shared scene tower network module corresponding to each task and a specific scene tower network module corresponding to each task and each scene, that is, the shared scene tower network module in a single tower network set corresponding to a single task can be shared by the tower network corresponding to different scenes in the task, and the shared scene tower network module corresponding to different tasks is different. Step S250 can include the following steps:
[0094] S251: determining first weight parameter information of the shared scene tower network module corresponding to each task and first bias parameter information of the shared scene tower network module corresponding to each task.
[0095] The shared scene tower network module can be a three-layer fully connected neural network, the first weight parameter information includes first weight values of each layer, and the first bias parameter information includes first bias values of each layer.
[0096] S252: determining second weight parameter information of the specific scene tower network module corresponding to each task and each scene and second bias parameter information of the specific scene tower network module corresponding to each task and each scene.
[0097] Feasibly, the network architecture of the specific scene tower network module is the same as that of the shared scene tower network module, but the parameter information is different, and the parameter information of the specific scene tower network module corresponding to different tasks and different scenes is different.
[0098] S253: determining target weight parameter information corresponding to each scene and each task according to the first weight parameter information and the second weight parameter information.
[0099] Feasibly, the first weight value of each layer in the first weight parameter information and the second weight value of the same layer in the second weight parameter information can be added to obtain the target weight parameter information, and the target weight parameter information includes target weight values of each layer.
[0100] S254: determining target bias parameter information corresponding to each scene and each task according to the first bias parameter information and the second bias parameter information.
[0101] Feasibly, the first bias value of each layer in the first bias parameter information and the second bias value of the same layer in the second bias parameter information can be added to obtain the target bias parameter information, and the target bias parameter information includes a target bias value of each layer.
[0102] The shared scene tower network module corresponding to each task is used to learn the commonality of different scenes in the corresponding task, and the specific scene tower network module corresponding to each task and each scene is used to learn the characteristics of the corresponding scene in the corresponding task. In the embodiment of the present application, the interaction of the shared scene tower network module and the specific scene tower network module corresponding to each scene in the same task can be realized by adding the parameters of the shared scene tower network module and the parameters of the specific scene tower network module corresponding to each scene in the same task.
[0103] S255: performing nonlinear transformation processing on the first output result corresponding to each task in each scene based on the target weight parameter information corresponding to each scene and each task and the target bias parameter information corresponding to each scene and each task, to obtain the first interaction prediction result corresponding to each task in each scene.
[0104] Based on the target weight parameter information corresponding to each scene and each task and the target bias parameter information corresponding to each scene and each task, a target tower network corresponding to each task and each scene can be constructed, and the first output result corresponding to the scene and the task is processed by nonlinear transformation using the target tower network, which is equivalent to calculating using an activation function, and the first interaction prediction result corresponding to the scene and the task can be obtained.
[0105] In the above embodiment, by introducing the shared scene tower network module and the specific scene tower network module, the commonality and the characteristics of each scene in each task are learned, and the features of the scene information are further enhanced, and the control of the scene information on the entire interaction processing is improved.
[0106] In an embodiment of the present application, the method can also be implemented as:
[0107] S310: obtaining original object feature data of a target object, original resource feature data of a candidate resource, and original scene feature data of each scene in a plurality of scenes.
[0108] S320: Fuse the original object feature data and the original resource feature data with the original scene feature data of each scene respectively to obtain first object feature data corresponding to each scene and first resource feature data corresponding to each scene.
[0109] S330: Perform resource feature extraction processing according to the original scene feature data and the original resource feature data of each scene to obtain second resource feature data corresponding to each scene.
[0110] Steps S310, S320 and S330 can refer to the foregoing embodiments, and details are not described herein.
[0111] S340: Perform feature cross processing on the first feature data corresponding to each scene to obtain second feature data corresponding to each scene; the first feature data corresponding to each scene is obtained by stacking the first object feature data corresponding to each scene, the first resource feature data corresponding to each scene and the second resource feature data corresponding to each scene.
[0112] Feasibly, the first feature data corresponding to each scene can be input into a Cross Network for feature cross processing to obtain the second feature data corresponding to each scene. In a specific Cross Network, the principle of feature cross can be as follows:
[0113] X i+1 = X0X i T W i +b i +X i (2)
[0114] Wherein, X o is the first feature data input initially, X i is the result output by the i-th layer feature cross processing, W i and b i are parameters of the i-th layer, and X i+1 is the result output by the i+1-th layer feature cross processing.
[0115] S350: Splice the first feature data corresponding to each scene and the second feature data corresponding to each scene to obtain third feature data corresponding to each scene.
[0116] Feasibly, the first feature data corresponding to each scene and the second feature data corresponding to each scene are spliced by tensor (Concat) to obtain the third feature data of each scene.
[0117] S360: Input the first feature data corresponding to each scene and the third feature data corresponding to each scene into the expert network corresponding to each task, perform feature extraction processing at the task layer, and obtain a second output result corresponding to each task in each scene.
[0118] In step S360, the third feature data corresponding to each scenario is input into the shared expert network module and the dedicated expert network module in the expert network corresponding to each task, and feature extraction for the corresponding task is performed. The extracted result is added to the first feature data corresponding to the scenario to obtain the second output result corresponding to each task in each scenario.
[0119] In one embodiment of the present application, step S360 may include:
[0120] S361: Input the third feature data corresponding to each scene into the shared expert network module in the expert network corresponding to each task, perform feature extraction processing on common features of multiple tasks, and obtain the fourth extraction result corresponding to each task in each scene.
[0121] S362: Input the third feature data corresponding to each scenario into the dedicated expert network module in the expert network corresponding to each task, perform feature extraction processing on the characteristics of each task, and obtain the fifth extraction result corresponding to each task in each scenario.
[0122] S363: Input the fourth extraction result corresponding to each task in each scenario and the fifth extraction result corresponding to each task in each scenario into the second gating network corresponding to each task, perform weighted processing, and obtain the sixth extraction result corresponding to each task in each scenario.
[0123] S364: Add the sixth extraction result corresponding to each task in each scenario and the first feature data corresponding to each scenario to obtain a second output result corresponding to each task in each scenario.
[0124] Adding the sixth extraction result corresponding to each task in each scene and the first feature data corresponding to each scene can reduce the fading and loss of scene information, thereby enhancing the perception of scene information and improving the control of scene information over the entire interactive prediction processing process.
[0125] For steps S361 to S364, specific reference may be made to the embodiments provided in steps S241 to S244, which will not be described in detail here.
[0126] S370: input the second output result corresponding to each task in each scene into the tower network corresponding to each scene and each task, perform scene layer interaction prediction processing, and obtain the second interaction prediction result corresponding to each task in each scene. The second interaction prediction result represents the possibility of the interaction event between the target object and the candidate resource in the corresponding scene.
[0127] Step S370 can refer to step S250 in the foregoing embodiments, which will not be repeated here.
[0128] In the foregoing embodiments, by performing feature cross processing on the first feature data corresponding to each scene, the understanding of the interaction or relationship between the first object feature data corresponding to each scene, the first resource feature data corresponding to each scene, and the second resource feature data corresponding to each scene can be enhanced, and the expression ability of the feature data can be improved.
[0129] In an embodiment of the present application, an artificial intelligence model is constructed to perform the interaction processing method for multiple scenes and multiple tasks provided in the foregoing embodiments. The artificial intelligence model includes a first feature representation module, a second feature representation module, a plurality of expert networks corresponding to a plurality of tasks one by one, and a plurality of tower networks corresponding to a plurality of tasks and a plurality of scenes one by one.
[0130] Figure 3 is a structural diagram of a first feature representation module (Feature Embedding Module), a second feature representation module (Attention Module), a plurality of expert networks (Expert1 Module, Expert2 Module, Expert13 Module, etc.) in an artificial intelligence model provided by an embodiment of the present application. User features are original object feature data of a target object, item features are original resource feature data of a candidate resource, and domain features are original scene feature data of any one scene.
[0131] Figure 4 is a structural diagram of a plurality of tower networks in an artificial intelligence model provided by an embodiment of the present application. The plurality of tasks are respectively to estimate three targets of whether to play a candidate video (corresponding to the click conversion rate in Figure 4 ), the playing time, and the playing frequency, and the plurality of scenes for each task include a strong purpose area and a purposeless area; out1, out2, and out3 are respectively outputs of the expert networks corresponding to the three tasks.
[0132] Figure 3 and Figure 4 The model processing process shown in and can refer to the foregoing embodiments, which will not be repeated here.
[0133] In a specific application embodiment of the present application, video resources are pushed to users in different scenarios of a video application. The scenarios can be divided into multiple levels according to terminal types, page types, region types and content module types of the video application. The terminal types can include mobile terminals, computer terminals, digital television terminals and the like. The page types can include selected pages, live broadcast pages, channel pages and the like. The region types can include strong purpose regions and purposeless regions and the like. The content modules include a head picture module, a follow-up drama module, a heavy module, a recently watched module and the like in the strong purpose region, and a list module, a film list module, a live broadcast module and the like in the purposeless region. Based on the interactive processing method suitable for multiple scenarios and multiple tasks provided in the embodiments of the present application, the interactive effects corresponding to each terminal application, each page, each region and each content module are positively improved. For example, each interactive index under the content module dimension, each interactive index under the region dimension, each interactive index under the page dimension and the like are improved. In particular, the interactive indexes such as the per capita positive film play time in the strong purpose region, the per capita play times in the head picture module, the per capita positive film play time and play times in the follow-up drama module and the like are significantly improved. The global interactive effect of each type of terminal application is also effectively improved.
[0134] As can be known from the above embodiments, the application provides an interaction processing method suitable for multiple scenes and multiple tasks. The original object feature data of a target object, the original resource feature data of a candidate resource, and the original scene feature data of each scene in multiple scenes are obtained, and the original object feature data and the original resource feature data are fused with the original scene feature data of each scene respectively, so that the first object feature data corresponding to each scene and the first resource feature data corresponding to each scene can be obtained. The first object feature data and the first resource feature data both incorporate the original scene feature data of the corresponding scene, that is, the perception of the scene is improved in the feature representation stage, and the different feature performances of the target object and the candidate resource in different scenes can be reflected. In the scheme provided by the application, the resource feature extraction processing is performed according to the original scene feature data of each scene and the original resource feature data, so that the second resource feature data corresponding to each scene can be obtained. The second resource feature data is extracted from the original resource feature data under the influence of the original scene feature data, that is, the perception of the scene is improved in the feature representation stage, and the preference degree of the target object for the same candidate resource in different scenes can be reflected. In the scheme provided by the application, the first feature data corresponding to each scene is input into the expert network corresponding to each task in multiple tasks for task-level feature extraction processing, so that the first output result corresponding to each task in each scene can be obtained. The first feature data corresponding to each scene is obtained by stacking the first object feature data corresponding to each scene, the first resource feature data corresponding to each scene, and the second resource feature data corresponding to each scene. Through the expert network corresponding to each task, the first feature data can be more effectively extracted from the task dimension, and the parallel processing of different tasks can be realized. In the scheme provided by the application, the first output result corresponding to each task in each scene is input into the tower network corresponding to each scene and each task for scene-level interaction prediction processing, so that the first interaction prediction result corresponding to each task in each scene can be obtained. The first interaction prediction result represents the possibility of the occurrence of the interaction event indicated by the corresponding task between the target object and the candidate resource in the corresponding scene. Through the tower network corresponding to each task and each scene, the interaction prediction processing corresponding to each task from the scene dimension can be further performed on the basis of the foregoing, and the parallel processing of different scenes can be realized.
[0135] The scheme provided by the application can simultaneously perform the interaction prediction processing of multiple scenes and multiple tasks on the target object and the candidate resource, simultaneously strengthens the perception of the scene in the feature representation stage, can more comprehensively and accurately estimate the effect of resource recommendation, improves the satisfaction of resource recommendation, and optimizes the user experience.
[0136] The application embodiment further provides an interaction processing device 600 suitable for multiple scenes and multiple tasks, as shown in Figure 6As shown, the apparatus can include:
[0137] The acquisition module 610 is configured to acquire original object feature data of a target object, original resource feature data of a candidate resource, and original scene feature data of each scene in a plurality of scenes.
[0138] The first feature representation module 620 is configured to fuse the original object feature data and the original resource feature data with the original scene feature data of each scene respectively to obtain first object feature data corresponding to each scene and first resource feature data corresponding to each scene.
[0139] The second feature representation module 630 is configured to perform resource feature extraction processing according to the original scene feature data of each scene and the original resource feature data to obtain second resource feature data corresponding to each scene.
[0140] The first feature extraction module 640 is configured to input first feature data corresponding to each scene into an expert network corresponding to each task in a plurality of tasks to perform task-layer feature extraction processing and obtain first output results corresponding to each task in each scene under the each scene; the first feature data corresponding to each scene is obtained by stacking first object feature data corresponding to each scene, first resource feature data corresponding to each scene, and second resource feature data corresponding to each scene.
[0141] The first interaction prediction module 650 is configured to input the first output results corresponding to each task in each scene into a tower network corresponding to each scene and each task to perform scene-layer interaction prediction processing and obtain first interaction prediction results corresponding to each task in each scene under the each scene, the first interaction prediction results representing a possibility of an interaction event indicated by the corresponding task occurring between the target object and the candidate resource in the corresponding scene.
[0142] In an embodiment of the present application, the first feature representation module 620 can include:
[0143] The feature transformation unit is configured to input the original scene feature data of each scene into a first gating network to perform feature transformation processing and obtain first scene feature data of each scene.
[0144] The first feature fusion unit is configured to perform Hadamard product-based feature fusion processing on the first scene feature data of each scene and the original object feature data to obtain first object feature data corresponding to each scene.
[0145] The second feature fusion unit is configured to perform Hadamard product-based feature fusion processing on the first scene feature data of each scene and the original resource feature data, to obtain first resource feature data corresponding to each scene.
[0146] In an embodiment of the present application, the second feature representation module 630 can include:
[0147] The first parameter determination unit is configured to determine first parameter information according to the original scene feature data of each scene and the original resource feature data.
[0148] The second parameter determination unit is configured to obtain historical interactive resource sequence data corresponding to the target object, and determine second parameter information and third parameter information according to the historical interactive resource sequence data, the historical interactive resource sequence data indicating at least one historical resource that has interacted with the target object.
[0149] The third parameter determination unit is configured to update the third parameter information according to the similarity between the first parameter information and the second parameter information, to obtain fourth parameter information.
[0150] The attention network determination unit is configured to determine a target attention network based on the first parameter information, the second parameter information and the fourth parameter information.
[0151] The key feature extraction unit is configured to input the original resource feature data into the target attention network, to perform feature extraction processing, to obtain second resource feature data corresponding to each scene.
[0152] In an embodiment of the present application, the expert network corresponding to each task includes a shared expert network module, a specific expert network module corresponding to each task, and a second gating network corresponding to each task, and the first feature extraction module 640 can include:
[0153] The first extraction unit is configured to input the first feature data corresponding to each scene into the shared expert network module in the expert network corresponding to each task, to perform feature extraction processing for commonality of the multiple tasks, to obtain a first extraction result corresponding to each task in each scene.
[0154] The second extraction unit is configured to input the first feature data corresponding to each scene into the specific expert network module in the expert network corresponding to each task, to perform feature extraction processing for characteristics of each task, to obtain a second extraction result corresponding to each task in each scene.
[0155] a first weighting unit, configured to input the first extraction result corresponding to each task in each scene and the second extraction result corresponding to each task in each scene into the second gating network corresponding to each task for weighting processing to obtain a third extraction result corresponding to each task in each scene;
[0156] a first adding unit, configured to add the third extraction result corresponding to each task in each scene and the first feature data corresponding to each scene to obtain a first output result corresponding to each task in each scene.
[0157] In an embodiment of the present application, the tower network corresponding to each scene and each task includes a shared scene tower network module corresponding to each task and a specific scene tower network module corresponding to each task and each scene, and the first interaction prediction module 650 can include:
[0158] a fourth parameter information determination unit, configured to determine first weight parameter information of the shared scene tower network module corresponding to each task and first bias parameter information of the shared scene tower network module corresponding to each task;
[0159] a fifth parameter information determination unit, configured to determine second weight parameter information of the specific scene tower network module corresponding to each task and each scene and second bias parameter information of the specific scene tower network module corresponding to each task and each scene;
[0160] a sixth parameter information determination unit, configured to determine target weight parameter information corresponding to each scene and each task according to the first weight parameter information and the second weight parameter information;
[0161] a seventh parameter information determination unit, configured to determine target bias parameter information corresponding to each scene and each task according to the first bias parameter information and the second bias parameter information;
[0162] a first interaction prediction unit, configured to perform nonlinear transformation processing on the first output result corresponding to each task in each scene based on the target weight parameter information corresponding to each scene and each task and the target bias parameter information corresponding to each scene and each task to obtain a first interaction prediction result corresponding to each task in each scene.
[0163] In an embodiment of the present application, the device 600 can further include:
[0164] a feature cross module, configured to perform feature cross processing on the first feature data corresponding to each of the scenes to obtain second feature data corresponding to each of the scenes;
[0165] a splicing module, configured to splice the first feature data corresponding to each of the scenes and the second feature data corresponding to each of the scenes to obtain third feature data corresponding to each of the scenes;
[0166] a second feature extraction module, configured to input the first feature data corresponding to each of the scenes and the third feature data corresponding to each of the scenes into the expert network corresponding to each of the tasks to perform feature extraction processing at a task layer to obtain second output results corresponding to each of the tasks under each of the scenes;
[0167] a second interaction prediction module, configured to input the second output results corresponding to each of the tasks under each of the scenes into the tower network corresponding to each of the scenes and each of the tasks to perform interaction prediction processing at a scene layer to obtain second interaction prediction results corresponding to each of the tasks under each of the scenes, the second interaction prediction results representing a possibility of an interaction event corresponding to each of the tasks occurring between the target object and the candidate resource under the corresponding scene.
[0168] In an embodiment of the present application, the expert network corresponding to each of the tasks includes a shared expert network module, a specific expert network module corresponding to each of the tasks, and a second gating network corresponding to each of the tasks, and the second feature extraction module includes:
[0169] a third extraction unit, configured to input the third feature data corresponding to each of the scenes into the shared expert network module in the expert network corresponding to each of the tasks to perform feature extraction processing for commonality of the multiple tasks to obtain fourth extraction results corresponding to each of the tasks under each of the scenes;
[0170] a fourth extraction unit, configured to input the third feature data corresponding to each of the scenes into the specific expert network module in the expert network corresponding to each of the tasks to perform feature extraction processing for characteristics of each of the tasks to obtain fifth extraction results corresponding to each of the tasks under each of the scenes;
[0171] a second weighting unit, configured to input the fourth extraction results corresponding to each of the tasks under each of the scenes and the fifth extraction results corresponding to each of the tasks under each of the scenes into the second gating network corresponding to each of the tasks to perform weighting processing to obtain sixth extraction results corresponding to each of the tasks under each of the scenes;
[0172] A second adding unit is configured to add the sixth extraction result corresponding to each task in each scene and the first feature data corresponding to each scene to obtain a second output result corresponding to each task in each scene.
[0173] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.
[0174] It should be noted that the device provided by the above-mentioned embodiments, in realizing its functions, is only exemplified by the above-mentioned division of each functional module, and in actual application, the above-mentioned functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above-described functions. In addition, the device and method embodiments provided by the above-mentioned embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be described here.
[0175] The embodiments of the present application provide a computer device, which comprises a processor and a memory, and the memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement the interactive processing method for multiple scenes and multiple tasks provided by the above-mentioned method embodiments.
[0176] Figure 7 A hardware structure schematic diagram of a device for implementing the interactive processing method for multiple scenes and multiple tasks provided by the embodiments of the present application is shown, and the device can participate in constituting or containing the device or system provided by the embodiments of the present application. As shown in Figure 7 The device 10 can include one or more (shown as 1002a, 1002b, …, 1002n in the figure) processors 1002 (the processor 1002 can include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1004 for storing data, and a transmission device 1006 for communication function. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. Those skilled in the art can understand that, Figure 7The illustrated structure is merely a schematic and does not limit the structure of the electronic device described above. For example, the device 10 can further include more or less components than those shown, or have a different configuration of the components. Figure 7 Figure 7
[0177] It should be noted that the one or more processors 1002 and / or other data processing circuitry described above can be generally referred to herein as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part within any one of the other elements of the device 10 (or mobile device). As referred to in embodiments of the present application, the data processing circuitry serves as a processor to control, for example, the selection of the variable resistance terminal path connected to the interface.
[0178] The memory 1004 can be used to store software programs and modules of application software, and program instructions / data storage means corresponding to the method described in embodiments of the present application. The processor 1002 can execute various functional applications and data processing by running the software programs and modules stored in the memory 1004, i.e. implement the above-mentioned interactive processing method suitable for multiple scenarios and multiple tasks. The memory 1004 can include a high-speed random access memory, and can further include a non-volatile memory such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 1004 can further include a memory remotely arranged with respect to the processor 1002, which can be connected to the device 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0179] The transmission device 1006 is used to receive or send data via a network. Specific examples of the above-mentioned network can include a wireless network provided by a communication provider of the device 10. In one example, the transmission device 1006 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 1006 can be a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet in a wireless manner.
[0180] The display can be, for example, a touch screen type liquid crystal display (LCD), which can enable a user to interact with the user interface of the device 10 (or mobile device).
[0181] The embodiment of the present application further provides a computer readable storage medium, which can be arranged in a server to save at least one instruction or at least one program for implementing a multi-scene and multi-task interactive processing method.
[0182] Optionally, in the embodiment, the storage medium can be located in at least one of a plurality of network servers of a computer network. Optionally, in the embodiment, the storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various storage program code media.
[0183] The embodiment of the present application further provides a computer program product or a computer program, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the multi-scene and multi-task interactive processing method provided in the various optional embodiments.
[0184] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The above-mentioned specific embodiments of the present application are described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be executed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.
[0185] Each of the embodiments in the present application is described in a progressive manner, and the same or similar parts of each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the device, equipment and storage medium embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.
[0186] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or can be instructed to relevant hardware by program. The program can be stored in a computer readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0187] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An interactive processing method suitable for multi-scene multi-task, characterized in that, The method comprises: obtaining original object feature data of a target object, original resource feature data of a candidate resource, and original scene feature data of each scene in a plurality of scenes; fusing the original object feature data and the original resource feature data with the original scene feature data of each scene respectively to obtain first object feature data corresponding to each scene and first resource feature data corresponding to each scene; performing resource feature extraction processing according to the original scene feature data of each scene and the original resource feature data to obtain second resource feature data corresponding to each scene; inputting first feature data corresponding to each scene into an expert network corresponding to each task in a plurality of tasks to perform task layer feature extraction processing and obtain first output results corresponding to each task in each scene; the first feature data corresponding to each scene is obtained by stacking the first object feature data corresponding to each scene, the first resource feature data corresponding to each scene, and the second resource feature data corresponding to each scene; inputting the first output results corresponding to each task in each scene into a tower network corresponding to each scene and each task to perform scene layer interaction prediction processing and obtain first interaction prediction results corresponding to each task in each scene, the first interaction prediction results representing the possibility of an interaction event indicated by the corresponding task occurring between the target object and the candidate resource in the corresponding scene.
2. The method of claim 1, wherein, The fusing of the original object feature data and the original resource feature data with the original scene feature data of each scene respectively to obtain first object feature data corresponding to each scene and first resource feature data corresponding to each scene comprises: inputting the original scene feature data of each scene into a first gating network to perform feature transformation processing and obtain first scene feature data of each scene; performing Hadamard product-based feature fusion processing on the first scene feature data of each scene and the original object feature data to obtain first object feature data corresponding to each scene; performing Hadamard product-based feature fusion processing on the first scene feature data of each scene and the original resource feature data to obtain first resource feature data corresponding to each scene.
3. The method of claim 1, wherein, The performing of resource feature extraction processing according to the original scene feature data of each scene and the original resource feature data to obtain second resource feature data corresponding to each scene comprises: determining first parameter information according to the original scene feature data of each scene and the original resource feature data; obtaining historical interaction resource sequence data corresponding to the target object, and determining second parameter information and third parameter information according to the historical interaction resource sequence data, the historical interaction resource sequence data indicating at least one historical resource that has interacted with the target object; updating the third parameter information according to a similarity between the first parameter information and the second parameter information, to obtain fourth parameter information; determining a target attention network based on the first parameter information, the second parameter information, and the fourth parameter information; inputting the original resource feature data into the target attention network for feature extraction processing, to obtain second resource feature data corresponding to each scene.
4. The method of claim 1, wherein, The expert network corresponding to each task includes a shared expert network module, a specific expert network module corresponding to each task, and a second gating network corresponding to each task. The first feature data corresponding to each scene is input into the expert network corresponding to each task for task-level feature extraction processing to obtain a first output result corresponding to each task in each scene, including: The first feature data corresponding to each scene is input into the shared expert network module in the expert network corresponding to each task for feature extraction processing common to the multiple tasks to obtain a first extraction result corresponding to each task in each scene. The first feature data corresponding to each scene is input into the specific expert network module in the expert network corresponding to each task for feature extraction processing of the characteristics of each task to obtain a second extraction result corresponding to each task in each scene. The first extraction result corresponding to each task in each scene and the second extraction result corresponding to each task in each scene are input into the second gating network corresponding to each task for weighting processing to obtain a third extraction result corresponding to each task in each scene. The third extraction result corresponding to each task in each scene and the first feature data corresponding to each scene are added to obtain the first output result corresponding to each task in each scene.
5. The method of claim 1, wherein, The tower network corresponding to each scene and each task includes a shared scene tower network module corresponding to each task and a specific scene tower network module corresponding to each task and each scene. The first output result corresponding to each task in each scene is input into the tower network corresponding to each scene and each task for scene-level interaction prediction processing to obtain a first interaction prediction result corresponding to each task in each scene, including: determining first weight parameter information of the shared scene tower network module corresponding to each task and first bias parameter information of the shared scene tower network module corresponding to each task; determining second weight parameter information of the specific scene tower network module corresponding to each task and each scene and second bias parameter information of the specific scene tower network module corresponding to each task and each scene; determining target weight parameter information corresponding to each scene and each task according to the first weight parameter information and the second weight parameter information; and determining target bias parameter information corresponding to each scene and each task according to the first bias parameter information and the second bias parameter information. According to the first bias parameter information and the second bias parameter information, target bias parameter information corresponding to each scene and each task is determined; Based on the target weight parameter information corresponding to each scene and each task and the target bias parameter information corresponding to each scene and each task, a first output result corresponding to each task in each scene is subjected to nonlinear transformation processing to obtain a first interaction prediction result corresponding to each task in each scene.
6. The method of claim 1, wherein, The method further comprises: performing feature cross processing on the first feature data corresponding to each scene to obtain second feature data corresponding to each scene; splicing the first feature data corresponding to each scene and the second feature data corresponding to each scene to obtain third feature data corresponding to each scene; inputting the first feature data corresponding to each scene and the third feature data corresponding to each scene into the expert network corresponding to each task to perform feature extraction processing at the task layer to obtain a second output result corresponding to each task in each scene; inputting the second output result corresponding to each task in each scene into the tower network corresponding to each scene and each task to perform interaction prediction processing at the scene layer to obtain a second interaction prediction result corresponding to each task in each scene, which represents the possibility of an interaction event corresponding to the target object and the candidate resource in the corresponding scene.
7. The method of claim 6, wherein, The expert network corresponding to each task comprises a shared expert network module, a specific expert network module corresponding to each task, and a second gating network corresponding to each task, and the inputting of the first feature data corresponding to each scene and the third feature data corresponding to each scene into the expert network corresponding to each task to perform feature extraction processing at the task layer to obtain a second output result corresponding to each task in each scene comprises: inputting the third feature data corresponding to each scene into the shared expert network module in the expert network corresponding to each task to perform feature extraction processing for the commonality of the plurality of tasks to obtain a fourth extraction result corresponding to each task in each scene; inputting the third feature data corresponding to each scene into the specific expert network module in the expert network corresponding to each task to perform feature extraction processing for the characteristics of each task to obtain a fifth extraction result corresponding to each task in each scene; inputting the fourth extraction result corresponding to each task in each scene and the fifth extraction result corresponding to each task in each scene into the second gating network corresponding to each task to perform weighting processing to obtain a sixth extraction result corresponding to each task in each scene; Add the sixth extraction result corresponding to each task in each scene and the first feature data corresponding to each scene to obtain a second output result corresponding to each task in each scene.
8. An interactive processing device suitable for multi-scenario multi-task, characterized in that, The device comprises: An acquisition module is configured to acquire original object feature data of a target object, original resource feature data of a candidate resource, and original scene feature data of each scene in a plurality of scenes. A first feature representation module is configured to fuse the original object feature data and the original resource feature data with the original scene feature data of each scene to obtain first object feature data corresponding to each scene and first resource feature data corresponding to each scene. A second feature representation module is configured to perform resource feature extraction processing according to the original scene feature data of each scene and the original resource feature data to obtain second resource feature data corresponding to each scene. A first feature extraction module is configured to input the first feature data corresponding to each scene into an expert network corresponding to each task in a plurality of tasks to perform task-layer feature extraction processing and obtain a first output result corresponding to each task in each scene; the first feature data corresponding to each scene is obtained by stacking the first object feature data corresponding to each scene, the first resource feature data corresponding to each scene, and the second resource feature data corresponding to each scene. A first interaction prediction module is configured to input the first output result corresponding to each task in each scene into a tower network corresponding to each scene and each task to perform scene-layer interaction prediction processing and obtain a first interaction prediction result corresponding to each task in each scene, the first interaction prediction result representing a possibility of an interaction event corresponding to each task between the target object and the candidate resource in the corresponding scene.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the interaction processing method for multiple scenes and multiple tasks according to any one of claims 1 to 7.
10. A computer device, comprising: The computer device comprises a processor and a memory, and the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the interaction processing method for multiple scenes and multiple tasks according to any one of claims 1 to 7.
11. A computer program product comprising a computer program, characterized in that, The computer program is executed by the electronic device to implement the interaction processing method for multiple scenes and multiple tasks according to any one of claims 1 to 7.