Live room information processing method, live room information providing method and device

By clustering and associating data from live streaming rooms, rich sample data for information retrieval is generated. This data is then used to enhance the training of the information retrieval model, solving the problems of information processing and accuracy in live streaming rooms and improving the model's accuracy.

CN122160527APending Publication Date: 2026-06-05TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2024-12-04
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing live streaming information processing and delivery information retrieval models suffer from low accuracy due to the sparse positive and negative sample data, resulting in low accuracy in live streaming information processing.

Method used

By clustering multiple sample live streaming rooms, a set of target live streaming rooms with associated advertising information is selected. Live streaming rooms without associated advertising information are then associated with live streaming rooms with associated advertising information to generate live streaming room advertising sample data. Audience account operation data is used to generate information retrieval sample data to enhance the advertising information retrieval model.

Benefits of technology

By making full use of live streaming data in the natural traffic domain, the sample data for live streaming ad placement was enriched, and the accuracy of live streaming information processing and ad placement information recall model was improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122160527A_ABST
    Figure CN122160527A_ABST
Patent Text Reader

Abstract

The application discloses a live broadcast room information processing method, a live broadcast room information providing information recall method and device. The method comprises the following steps: clustering a plurality of sample live broadcast rooms as nodes to obtain a plurality of sample live broadcast room sets; in the plurality of sample live broadcast room sets, a target live broadcast room set to which a live broadcast room associated with information providing information belongs is screened; for each target live broadcast room set, a live broadcast room in the target live broadcast room set which is not associated with information providing information is associated with information providing information associated with the live broadcast room in the same target live broadcast room set to obtain live broadcast room information providing sample data; based on audience account operation data corresponding to the live broadcast room in the live broadcast room information providing sample data, information recall sample data is generated; and the information recall sample data is used for enhancing training of an information providing information recall model to obtain an enhanced information providing information recall model. The method can improve the accuracy of live broadcast room information processing and the accuracy of the information providing information recall model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method and apparatus for processing information in a live streaming room and for recalling information released in a live streaming room. Background Technology

[0002] With the development of internet technology and mobile devices, various software applications often need to recommend personalized resources to accounts. Currently, recommendation systems typically employ recall and ranking processes for layered filtering. In the recall phase, the recall model needs to quickly calculate and retrieve resources that the target account might be interested in. With the development of e-commerce technology, merchants utilize live streaming platforms to share their products. Furthermore, to achieve better promotional results, merchants often use internet advertising platforms to place ads in live streaming rooms to promote their products.

[0003] However, the number of live streams that ran ads was smaller than those that did not, and for those that did run ads, viewer feedback on the ads was also less frequent. Therefore, due to the sparse positive and negative sample data, the accuracy of the trained ad recall model was low, and the accuracy of live stream information processing was also low. Summary of the Invention

[0004] This application provides a method for processing live streaming room information, a method and apparatus for recalling live streaming room information, which can improve the accuracy of live streaming room information processing and the accuracy of the information recall model.

[0005] On the one hand, this application provides a method for processing live streaming room information, the method comprising:

[0006] Clustering is performed using multiple sample live streaming rooms as nodes to obtain a set of multiple sample live streaming rooms;

[0007] Among the multiple sample live streaming room sets, filter the target live streaming room set to which the live streaming room with associated advertising information belongs;

[0008] For each set of target live streaming rooms, the live streaming rooms in the target live streaming room set that are not associated with the advertising information are associated with the advertising information associated with the live streaming rooms in the same set of target live streaming rooms to obtain live streaming room advertising sample data.

[0009] Based on the viewer account operation data corresponding to the live broadcast room in the live broadcast room delivery sample data, information recall sample data is generated.

[0010] The enhanced information retrieval model is obtained by training the information retrieval sample data.

[0011] On the other hand, this application provides a method for recalling live streaming room advertising information, the method comprising:

[0012] Obtain the target audience's account attribute information and the corresponding live stream advertising candidate data;

[0013] The target account attribute information and the live broadcast room delivery candidate data are input into the enhanced delivery information recall model to perform delivery information operation data prediction processing and obtain a prediction result; the prediction result is used to characterize whether to recommend the candidate delivery information to the target audience account; the enhanced delivery information recall model is trained according to the live broadcast room information processing method described above.

[0014] On the other hand, this application provides a live streaming room information processing device, which includes:

[0015] The sample live streaming room clustering module is used to cluster multiple sample live streaming rooms as nodes to obtain a set of multiple sample live streaming rooms.

[0016] The target live streaming room set determination module is used to filter the target live streaming room set to which the live streaming room associated with the advertising information belongs from the multiple sample live streaming room sets;

[0017] The delivery information association module is used to associate the live rooms in the target live room set that are not associated with delivery information with the delivery information associated with the live rooms in the same target live room set for each target live room set, so as to obtain live room delivery sample data.

[0018] The information recall sample data generation module is used to generate information recall sample data based on the viewer account operation data corresponding to the live broadcast room in the live broadcast room delivery sample data.

[0019] The model training module is used to enhance the delivery information recall model using the information recall sample data to obtain an enhanced delivery information recall model.

[0020] On the other hand, this application provides a live streaming room ad delivery information recall device, the device comprising:

[0021] The acquisition module is used to acquire the target account attribute information and candidate data for live streaming room delivery corresponding to the target audience account and the candidate delivery information.

[0022] The processing module is used to input the target account attribute information and the live broadcast room delivery candidate data into the enhanced delivery information recall model, perform delivery information operation data prediction processing, and obtain a prediction result; the prediction result is used to characterize whether to recommend the candidate delivery information to the target audience account; the enhanced delivery information recall model is trained according to the live broadcast room information processing method described above.

[0023] On the other hand, an electronic device is provided, which includes a processor and a memory. The memory stores at least one instruction or at least one program. The at least one instruction or the at least one program is loaded and executed by the processor to implement the live room information processing method or the live room delivery information recall method as described above.

[0024] On the other hand, a computer storage medium is provided, which stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the live room information processing method or the live room information recall method as described above.

[0025] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the live-streaming information processing method or the live-streaming information recall method as described above.

[0026] The live streaming room information processing method, live streaming room ad delivery information recall method and device provided in this application have the following technical effects:

[0027] In this embodiment, multiple sample live-streaming rooms are clustered as nodes to obtain multiple sample live-streaming room sets. Within these sets, target live-streaming room sets are selected from those associated with advertising information. For each target live-streaming room set, live-streaming rooms without associated advertising information are associated with the advertising information associated with the live-streaming rooms in the same target live-streaming room set, resulting in live-streaming room advertising sample data. Clustering allows the data from live-streaming rooms without associated advertising information to be utilized. Based on the viewer account operation data corresponding to the live-streaming rooms in the advertising sample data, information retrieval sample data is generated. The advertising information retrieval model is enhanced and trained using this information retrieval sample data, resulting in an enhanced advertising information retrieval model. Compared to traditional solutions that only use live-streaming room data associated with advertising information for training, this application fully utilizes data from live-streaming rooms (live-streaming rooms without associated advertising information) in the natural traffic domain, making the live-streaming room advertising sample data richer, thereby improving the accuracy of live-streaming room information processing and the accuracy of the advertising information retrieval model. Attached Figure Description

[0028] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a schematic diagram of the application environment of a delivery information recall model provided in an embodiment of this application;

[0030] Figure 2 This is a flowchart illustrating a live-stream information processing method provided in an embodiment of this application;

[0031] Figure 3 This is a flowchart illustrating the process of determining cluster nodes for target samples provided in an embodiment of this application;

[0032] Figure 4 This is a schematic diagram of the process for generating information recall sample data provided in the embodiments of this application;

[0033] Figure 5 This is a flowchart illustrating the enhanced delivery information retrieval model obtained through training, as provided in the embodiments of this application.

[0034] Figure 6 This is a schematic diagram of the overall process of training the information retrieval model provided in the embodiments of this application;

[0035] Figure 7This is a schematic diagram of the structure of the live streaming room information processing device provided in the embodiments of this application;

[0036] Figure 8 This is a schematic diagram of the structure of the live broadcast room delivery information recall device provided in the embodiments of this application;

[0037] Figure 9 This is a hardware structure block diagram of a server for a live streaming information processing method provided in an embodiment of this application. Detailed Implementation

[0038] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0039] It is understood that in the specific implementation of this application, data related to account characteristics is involved. When the above embodiments of this application are applied to specific products or technologies, account permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0040] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0041] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0042] Please see Figure 1 , Figure 1 This is a schematic diagram of the application environment of a delivery information recall model provided in an embodiment of this application. The application environment may include at least a server 100 and a terminal 200.

[0043] In an optional embodiment, server 100 can be used to perform enhanced training on the delivery information retrieval model to obtain an enhanced delivery information retrieval model; server 100 can use the enhanced delivery information retrieval model to determine whether candidate delivery information can be recalled and recommended, and send it to terminal 100. Server 100 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0044] In an optional embodiment, terminal 200 can be used to provide audience account operation data required for the training process; terminal 200 can also be used to recommend and display the delivery information sent by server 100. Specifically, terminal 200 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, in-vehicle terminals, and smart TVs; it can also be software running on the above-mentioned electronic devices, such as applications and mini-programs. The operating system running on the electronic device in this embodiment can include, but is not limited to, Android, iOS, Linux, and Windows systems.

[0045] The following describes a method for processing live streaming room information according to this application. Figure 2 This is a flowchart illustrating a live-stream information processing method provided in this application. This specification provides the operational steps of the method as described in the embodiments or flowcharts, but based on conventional or non-inventive methods, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual system or server product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as shown... Figure 2 As shown, the method may include:

[0046] S201: Cluster multiple sample live stream rooms as nodes to obtain multiple sample live stream room sets. In one implementation, live stream rooms pushed to sample viewer accounts within a preset historical time period can be used as sample live stream rooms.

[0047] It should be noted that multiple sample live stream rooms can include those that do not contain advertising information in the advertising domain; these live stream rooms have organic traffic domain data. Advertising information can be understood as advertising. Multiple sample live stream rooms can also include those that do contain advertising information in the advertising domain; these live stream rooms have both organic traffic domain data and advertising domain data.

[0048] By using multiple sample live streaming rooms as nodes, sample live streaming rooms with similar natural traffic domain data characteristics can be aggregated into the same sample live streaming room set.

[0049] In one implementation, the clustering process can have fixed classification criteria, such as manually set criteria like products belonging to the same merchant or identical products in the product list. In another implementation, the clustering process can be based on a clustering model, using the model to uncover the inherent relationships between sample live streams, thereby clustering the sample live streams.

[0050] S203: Among multiple sample live streaming room sets, filter the target live streaming room set to which the live streaming room with associated advertising information belongs.

[0051] A live stream with associated advertising information can be understood as a live stream that has run ads. Based on whether or not there are associated live streams with advertising information, a target set of live streams is selected from multiple sample live stream sets to facilitate subsequent data migration. For the remaining sample live stream sets outside the target set, since none of them contain associated live streams with advertising information, it can be assumed that all sample live streams in this set cannot be associated with ads in the advertising domain; therefore, the relevant sample data for these live streams can be discarded.

[0052] S205: For each target live room set, associate the live rooms in the target live room set that are not associated with the advertising information with the advertising information associated with the live rooms in the same target live room set to obtain live room advertising sample data.

[0053] After the above clustering and filtering, the live streams in the target live stream set that have associated advertising information and those that do not have associated advertising information have similar natural traffic data characteristics, so data migration is possible. Therefore, live streams in the target live stream set that do not have associated advertising information can be associated with the advertising information associated with live streams in the same target live stream set, so that live streams that were originally not associated with advertising information are also associated with the advertising information in the advertising domain.

[0054] It should be noted that if there is only one live room in the target live room set that is associated with advertising information, then the advertising information associated with that live room can be directly associated with live rooms in the target live room set that are not associated with advertising information. If there are multiple live rooms in the target live room set that are associated with advertising information, then any one of the live rooms associated with advertising information can be randomly selected and the advertising information associated with that live room can be associated with live rooms in the target live room set that are not associated with advertising information.

[0055] Through the above process of associating advertising information, each live room in the live room set is associated with advertising information. Therefore, the sample data of live room advertising includes the natural traffic domain data corresponding to each live room in the target live room set as well as the advertising domain data of the advertising information associated with the live room.

[0056] S207: Generate information recall sample data based on the viewer account operation data corresponding to the live broadcast room in the live broadcast room delivery sample data.

[0057] Viewer account operation data can be understood as the actual operation data of sample viewer accounts on the advertising information associated with the live broadcast room. In one example, assuming that sample viewer account 1 clicks on advertising information a1 associated with live broadcast room A, the viewer account operation data corresponding to advertising information a1 can be set to 1; in another example, assuming that sample viewer account 2 does not click on advertising information b1 associated with live broadcast room B, the viewer account operation data corresponding to advertising information b1 can be set to 0.

[0058] It should be noted that the audience account operation data may also include the viewing time data and conversion data of the sample audience accounts for the advertising information associated with the live broadcast room.

[0059] By combining live stream delivery sample data with audience account operation data, a training sample is formed consisting of audience account-delivery information-audience account operation data. This results in information recall sample data, which can make full use of data from the natural traffic domain, making the information recall sample data richer.

[0060] S209: Enhance the information recall model by using information recall sample data to obtain an enhanced information recall model.

[0061] It should be noted that the recall process typically employs a dual-tower structure for information dissemination, thereby decoupling accounts and items. The dual-tower structure usually comprises an account tower and an item tower.

[0062] Account Tower is used to receive account-related features (such as account attribute information, historical behavior sequences, etc.), perform feature learning and representation learning through multi-layer neural networks, and finally output the semantic vector of the account.

[0063] The item tower is used to receive item-related features (such as basic item information, attribute information, etc.), and also performs feature learning and representation learning through a multi-layer neural network, ultimately outputting the semantic vector of the item.

[0064] Next, the semantic vectors of the account and the semantic vectors of the item are used to calculate the similarity between the account and the item, so that it can be used for tasks such as recall and ranking in the recommendation system.

[0065] During online prediction, the dual-tower structure can achieve offline calculation of the semantic vector of items, real-time online calculation of the semantic vector of accounts, and online similarity calculation to ensure real-time performance.

[0066] Therefore, in this embodiment of the application, the information delivery recall model can be a dual-tower model, including an account tower and an information delivery tower.

[0067] It should be noted that the information retrieval model in this application can be an untrained information retrieval model or a pre-trained information retrieval model. This application utilizes information retrieval sample data to perform augmentation training on the existing information retrieval model, thereby obtaining an enhanced information retrieval model.

[0068] In this embodiment, multiple sample live-streaming rooms are clustered as nodes to obtain multiple sample live-streaming room sets. Within these sets, target live-streaming room sets are selected from those associated with advertising information. For each target live-streaming room set, live-streaming rooms without associated advertising information are associated with the advertising information associated with the live-streaming rooms in the same target live-streaming room set, resulting in live-streaming room advertising sample data. Clustering allows the data from live-streaming rooms without associated advertising information to be utilized. Based on the viewer account operation data corresponding to the live-streaming rooms in the advertising sample data, information retrieval sample data is generated. The advertising information retrieval model is enhanced and trained using this information retrieval sample data, resulting in an enhanced advertising information retrieval model. Compared to traditional solutions that only use live-streaming room data associated with advertising information for training, this application fully utilizes data from live-streaming rooms (live-streaming rooms without associated advertising information) in the natural traffic domain, making the live-streaming room advertising sample data richer, thereby improving the accuracy of live-streaming room information processing and the accuracy of the advertising information retrieval model.

[0069] In one example, associating live streams without associated advertising information within the target live stream set with the advertising information associated with live streams within the same target live stream set to obtain live stream advertising sample data can include: obtaining a first advertising feature of the advertising information associated with live streams in the target live stream set; determining the first advertising feature as a second advertising feature corresponding to live streams without associated advertising information within the target live stream set; and determining live stream advertising sample data based on the first and second advertising features. The live stream advertising sample data includes both data from live streams already associated with advertising information and the advertising features of their associated advertising information, as well as data from live streams subsequently associated with advertising information and the advertising features of their associated advertising information. By reusing the advertising features of advertising information, the association of advertising information can be achieved efficiently.

[0070] In one embodiment, clustering multiple sample live streaming rooms as nodes to obtain a set of multiple sample live streaming rooms may include:

[0071] Obtain the sample attribute features corresponding to each of the multiple sample live streaming rooms to obtain a sample attribute feature set. Each sample live streaming room can correspond to sample attribute features, which can be obtained by concatenating the features corresponding to the sample live streaming room in the natural traffic domain. In addition to the unique identifier (ID) of the live streaming room, the features corresponding to the sample live streaming room in the natural traffic domain can also include the name of the product being shared in the live streaming room, the product list in the live streaming room, the name of the host, etc.

[0072] Based on the sample attribute feature set, an initial sample topology graph is constructed; the initial sample topology graph includes multiple initial sample nodes and multiple initial sample connection edges; each initial sample node represents a sample live broadcast room; each initial sample connection edge represents the association relationship between two connected initial sample nodes.

[0073] It should be noted that the initial sample topology graph can be constructed using the K-Nearest Neighbor (KNN) algorithm. First, all sample live streams are considered as initial sample nodes in the initial sample topology graph. For each initial sample node, the distance between it and the other initial sample nodes is measured, and they are sorted from nearest to farthest. The K nearest nodes are then designated as the neighbors of that initial sample node, and initial sample connection edges are established between the initial sample node and its neighbors. The distance between nodes can be calculated based on the sample attribute features of the corresponding sample live stream.

[0074] Next, based on the initial sample topology graph, node aggregation is performed on multiple initial sample nodes to obtain target sample cluster nodes. Each target sample cluster node is obtained by aggregating multiple initial sample nodes with similar sample attribute features. Then, based on the correspondence between target sample cluster nodes and initial sample nodes, and the correspondence between initial sample nodes and sample live streaming rooms, the set of sample live streaming rooms is determined. By tracing back which initial sample nodes were aggregated to form the target sample cluster nodes, the initial sample nodes can be grouped. Since there is a one-to-one correspondence between initial sample nodes and sample live streaming rooms, the grouping of sample live streaming rooms can be determined, resulting in multiple sets of sample live streaming rooms.

[0075] In this embodiment of the application, node clustering is performed by constructing an initial sample topology graph, which can improve the accuracy and convenience of clustering, so as to facilitate the subsequent association of live streaming information in the same sample live streaming room set.

[0076] It should be noted that, based on the initial sample topology graph, node aggregation processing of multiple initial sample nodes can be performed based on a preset clustering model. Figure 3 This is a flowchart illustrating the process of determining target sample cluster nodes according to an embodiment of this application. In one embodiment, based on the initial sample topology graph, performing node aggregation processing on multiple initial sample nodes to obtain target sample cluster nodes may include:

[0077] S301: Obtain a preset clustering model, which includes multiple clustering modules arranged in sequence.

[0078] The preset clustering model can be a pre-trained clustering model, which can be directly applied in the embodiments of this application. In one example, the preset clustering model can be a pre-trained Hi-LANDER model, which is a hierarchical clustering model based on a graph neural network (GNN). Each level corresponds to a clustering module.

[0079] S303: Determine the initial sample topology as the current sample topology, and determine the first clustering module among multiple sequentially arranged clustering modules as the current clustering module.

[0080] The preset clustering model consists of multiple clustering modules arranged sequentially, which can be denoted as the clustering module corresponding to level 1, the clustering module corresponding to level 2, and so on. It should be noted that the multiple sequentially arranged clustering modules have the same structure and parameters.

[0081] S305: Perform graph clustering on the current sample topology graph based on the current clustering module to obtain multiple intermediate sample clustering nodes.

[0082] In one embodiment, step S305 may include: inputting the current sample topology graph into the current clustering module for edge prediction processing to obtain a set of predicted connection edges; the edge prediction processing of the current clustering module can be regarded as using a function The current sample topology graph G and the node features H corresponding to the nodes in the current sample topology graph are processed to obtain the predicted connection edge set E', and the predicted connection edge set E' is a subset of the connection edge set E in the current sample topology graph.

[0083] Next, the nodes in the current sample topology graph are connected according to the predicted connection edges contained in the predicted connection edge set to obtain multiple connected components; each connected component includes member nodes and predicted connection edges between member nodes; in one example, assuming there are five nodes A, B, C, D and E in the current sample topology graph, the predicted connection edge set obtained in the above steps includes the connection edge between A and B, and the connection edge between B and C, then nodes A, B and C can form a connected component 1, node D forms a connected component 2 on its own; node E also forms a connected component 3 on its own.

[0084] Finally, for each connected component, the nodes contained in the connected component are aggregated to obtain multiple intermediate cluster nodes of samples.

[0085] Using the example above, if connected component 1 contains three member nodes A, B, and C, then nodes A, B, and C can be aggregated according to a preset node aggregation strategy to obtain the intermediate cluster node A'. The preset node aggregation strategy can be to take the average value of the node features corresponding to the member nodes and use this average value as the node feature of the intermediate cluster node. Since connected components 2 and 3 each have only one member node, this member node can be directly determined as the intermediate cluster node.

[0086] In another example, assuming the current clustering module is a level 1 clustering module, that is, the first clustering module, the node features of the nodes are the sample attribute features of the sample live room. Then, after obtaining the connected components, the average value of the sample attribute features corresponding to the nodes in the connected components can be taken, and the average value can be determined as the node features of the intermediate clustering nodes of the formed samples, thus serving as the input of the level 2 clustering module.

[0087] In this embodiment, the current clustering module performs edge prediction on the current sample topology graph to form connected components to aggregate the nodes in the current sample topology graph, thereby obtaining the intermediate clustering nodes of the sample and improving the accuracy of clustering.

[0088] S307: Construct an intermediate sample topology graph based on multiple intermediate sample clustering nodes.

[0089] The process of constructing an intermediate sample topology graph based on intermediate cluster nodes is similar to the process of constructing an initial sample topology graph based on initial sample nodes, and will not be described in detail here.

[0090] S309: Use the intermediate sample topology graph as the current sample topology graph again; and use the next clustering module after the current clustering module as the current clustering module again. Repeat steps S305 and S307 until the clustering termination condition is met.

[0091] The clustering termination condition can be that, during step S305, the set of predicted connecting edges obtained by inputting the current sample topology graph into the current clustering module for edge prediction processing is empty, meaning no edges are predicted in the current sample topology graph. Hierarchical clustering is achieved by continuously using the sample topology graph formed by the intermediate clustering nodes output by the level i clustering module as input to the level i+1 clustering module.

[0092] S311: The intermediate sample clustering node output by the clustering module at the end of the clustering process is determined as the target sample clustering node.

[0093] When clustering ends, the last clustering module, i.e. the last clustering module, can no longer generate connection edges for the intermediate sample clustering nodes it outputs. At this time, the intermediate sample clustering nodes output by the last clustering module are determined as the target sample clustering nodes.

[0094] In this embodiment, hierarchical clustering is performed using multiple sequentially arranged clustering modules, ensuring the accuracy of clustering and facilitating subsequent association of delivery information and model training based on the clustering nodes of the target samples.

[0095] In one embodiment, inputting the current sample topology graph into the current clustering module for edge prediction processing to obtain the predicted connection edge set may include: performing node encoding processing on the nodes contained in the current sample topology graph to obtain the encoded features corresponding to each node in the current sample topology graph; in one implementation, a Graph Attention Network (GAT) can be used for node encoding; in another implementation, a Graph Convolution Network (GCN) can also be used for node encoding; additionally, a Graph Sample and Aggregate (GraphSage) network can also be used for node encoding. This application does not limit the specific method of node encoding, and it can be set according to the limitations of model memory usage and the requirements for detection accuracy in actual applications.

[0096] Traverse the nodes contained in the current sample topology graph, and determine the node density of each traversed current node and the correlation coefficient corresponding to the connection edge between the current node and its neighboring nodes based on the encoding features of each node in the current sample topology graph; wherein, the correlation coefficient represents the correlation degree between the current node and its neighboring nodes; the neighboring node is the node in the current sample topology graph that has a connection edge with the current node.

[0097] If the correlation coefficient is greater than or equal to the preset coefficient threshold, the connection edge between the current node and its neighboring nodes is determined as a candidate edge; the preset coefficient threshold is a hyperparameter that remains fixed after the preset clustering model is trained.

[0098] Based on the correlation coefficient corresponding to each candidate edge, predictable connecting edges are determined from the candidate edges, resulting in a set of predictable connecting edges.

[0099] In another embodiment, nodes with a density greater than or equal to a preset density threshold can be selected from the nodes in the current sample topology graph, and the remaining nodes can be discarded. Node density can be used as a confidence level for node clustering; the higher the node density, the more accurate the clustering.

[0100] In this embodiment, the predicted connection edge set corresponding to the current sample topology graph is determined by two dimensions: node density and correlation coefficient. This allows the predicted connection set to accurately connect nodes belonging to the same cluster, thereby improving the accuracy of clustering.

[0101] It should be noted that, in one embodiment, determining the node density of each traversed node and the correlation coefficients of the connection edges between the current node and its neighbors based on the encoding features corresponding to each node in the current sample topology graph may include:

[0102] Obtain the first encoded feature corresponding to the current node and the second encoded feature corresponding to the neighboring nodes;

[0103] The first encoded feature corresponding to the current node is the feature obtained after the current node has undergone node encoding. Let h be the feature of the current node with index i. i After node encoding, the first encoded feature h′ is obtained. i The neighbor node of index j has the characteristic h. j After node encoding, the second encoded feature h is obtained. j '.

[0104] The first encoded feature and the second encoded feature are concatenated, and the connection probability is determined based on the concatenated encoded feature; where the connection probability represents the probability that the current node and its neighboring nodes belong to the same cluster;

[0105] In one implementation, the concatenated encoded features can be further processed through a Multi Layer Perceptron (MLP) followed by a softmax transform layer, thereby converting the vector output by the MLP into probabilities and outputting connection probabilities.

[0106] Determine the correlation coefficients of the connection edges between the current node and its neighboring nodes based on the connection probability.

[0107] In one embodiment, the correlation coefficient can be determined by the following formula:

[0108]

[0109] in, P(y) represents the association coefficient between the current node at index i and its neighbor node at index j; i =y j P(y) represents the probability that the current node at index i and its neighbor node at index j belong to the same cluster, i.e., the connection probability; i ≠y j The probability that the current node at index i and its neighbor node at index j do not belong to the same cluster is represented by .

[0110] The node density of the current node is determined based on the node similarity between the current node and its neighboring nodes in the current sample topology graph, as well as the correlation coefficients corresponding to the connecting edges between the current node and its neighboring nodes.

[0111] The node similarity between the current node and its neighbors in the current sample topology graph can be determined by the following formula:

[0112] a ij = <h i ,h j >

[0113] Among them, a ij h represents the inner product between the current node at index i and its neighbor node at index j, i.e., the node similarity. i h represents the feature of the current node with index i. j This represents the characteristics of the neighbor node with index j.

[0114] The node density of the current node can be determined by the following formula:

[0115]

[0116] in, This represents the node density of the current node with index i. a represents the correlation coefficient. ijThis represents the node similarity, where k is the number of neighboring nodes of the current node at index i.

[0117] In this embodiment, by determining the association coefficient based on the connection probability and the node density based on the association coefficient and similarity, the association coefficient and node density can be unified within a single prediction framework, thereby improving the prediction speed and clustering efficiency of the preset clustering model. It should be noted that in this application, after obtaining the sample account attribute information corresponding to the live stream in the live stream delivery sample data, information recall sample data can be generated based on the sample account attribute information and the live stream delivery sample data; wherein, the sample account attribute information corresponds to the same sample viewer account as the viewer account operation data. In this implementation, since the live stream delivery sample data is obtained after associating delivery information, the sample is expanded. Directly using the sample account attribute information and the live stream delivery sample data as input to the delivery information recall model can result in a more accurate delivery information recall model.

[0118] In one embodiment, Figure 4 This is a schematic diagram of the process for generating information recall sample data provided in this application embodiment. Besides expanding the training samples by associating delivery information in step S205, in order to more comprehensively utilize the data corresponding to the sample live stream in the natural traffic domain, the generation of information recall sample data in this application embodiment includes:

[0119] S401: Determine the intermediate sample clustering node corresponding to the live room in the live room delivery sample data as the sample augmentation node.

[0120] S403: Based on the sample augmentation nodes and the initial sample nodes corresponding to the live rooms in the live room delivery sample data, determine the live room augmentation sample features corresponding to the live rooms in the live room delivery sample data. Figure 3In the clustering process shown, each clustering module can output the node features of each intermediate sample cluster node. The intermediate sample cluster node corresponding to the live room in the live room delivery sample data is determined as the sample augmentation node corresponding to that live room. Then, the sample attribute features corresponding to the initial sample node corresponding to the live room in the live room delivery sample data are numerically averaged with the node features corresponding to the sample augmentation node to obtain the live room augmentation sample features corresponding to the live room in the live room delivery sample data. In one example, suppose that for live room A in the live room delivery sample data, the node feature corresponding to its initial sample node is f1. After the first clustering module, intermediate sample cluster node 1_1 is obtained; after the second clustering module, intermediate sample cluster node 2_1 is obtained; and the clustering ends after these two clustering modules. The node feature of intermediate sample cluster node 1_1 is f2, and the node feature of intermediate sample cluster node 2_1 is f3. Then, the live room augmentation sample features of live room A can be determined by the following formula:

[0121] s = mean(f1,f2,f3)

[0122] Where s represents the enhanced sample feature of the live room corresponding to the live room in the live room delivery sample data; f1 represents the node feature of the initial sample node corresponding to the live room in the live room delivery sample data; f2 and f3 represent the node features of the intermediate sample cluster nodes at different levels.

[0123] It should be understood that in practical applications, the number of node features after f1 in the above formula is determined based on the number of clustering modules involved in the actual clustering process.

[0124] S405: Obtain the sample account attribute information corresponding to the live room in the live room delivery sample data.

[0125] Among them, the sample account attribute information and the audience account operation data correspond to the same sample audience account.

[0126] S407: Generate information recall sample data based on sample account attribute information, live broadcast room delivery sample data, and live broadcast room enhanced sample features.

[0127] Among them, the information recall sample data is labeled with audience account operation data tags. The audience account operation data tags represent the actual operation data of the sample audience accounts on the sample delivery information. The sample delivery information is the delivery information associated with the live room in the live room delivery sample data.

[0128] In this embodiment, by combining the sample attribute features of the initial sample nodes with the node features corresponding to the intermediate sample cluster nodes, the enhanced sample features of the live streaming room can take into account both the sample attribute features of the live streaming room itself and the sample attribute features of all live streaming rooms within the same set of live streaming rooms to which the live streaming room belongs. Then, based on the sample account attribute information, live streaming room delivery sample data, and enhanced sample features of the live streaming room, information retrieval sample data is generated. This makes fuller use of the natural traffic domain data of the live streaming room, improves the accuracy of live streaming room information processing, and enhances the training of the delivery information retrieval model, thereby improving the model's accuracy.

[0129] In one embodiment, based on Figure 4 The information recall sample data obtained from the illustrated embodiment can be used to enhance the training of the information recall model. Figure 5 This is a flowchart illustrating the enhanced delivery information recall model trained according to an embodiment of this application. Since the delivery information recall model may include an account information representation network (account tower) and a delivery information representation network (item tower), the enhanced delivery information recall model, obtained by enhancing the model with delivery information recall sample data, may include:

[0130] S501: Input the sample account attribute information into the account information representation network to perform account feature extraction processing and obtain the sample account representation vector.

[0131] Since the information recall sample data includes the sample account attribute information corresponding to the live room in the live room delivery sample data, the account information representation network can be used to output the sample account representation vector.

[0132] S503: Input the live broadcast room delivery sample data into the delivery information representation network, and fuse the output of the delivery information representation network with the live broadcast room enhanced sample features to obtain the sample delivery information representation vector.

[0133] The sample data of the live broadcast room is input into the broadcast information representation network. The output of the second to last layer of the broadcast information representation network is concatenated with the enhanced sample features of the live broadcast room and then fed into the last layer of the broadcast information representation network to obtain the sample broadcast information representation vector.

[0134] In this embodiment, by using a residual connection (shortcut) between the enhanced live-streaming room features and the live-streaming room delivery sample data, the features of the shallow enhanced live-streaming room features and the features of the deep live-streaming room delivery sample data can be directly fused. On the one hand, this can alleviate the vanishing and exploding gradient problems, and also mitigate the degradation problem. Furthermore, the residual connection also facilitates faster convergence of the recall model, improving training efficiency. On the other hand, compared to directly inputting the live-streaming room delivery sample data into the delivery information representation network to obtain the output result, this application utilizes the enhanced live-streaming room features, enabling the sample delivery information representation vector to accurately represent the delivery information associated with the live-streaming room, and also allowing for full utilization of natural traffic domain data.

[0135] S505: The vector similarity between the sample account representation vector and the sample delivery information representation vector is determined as the prediction operation data of the sample audience account for the sample delivery information.

[0136] In one embodiment, the vector similarity between the sample account representation vector and the sample delivery information representation vector can be achieved by calculating their inner product. The implementation process of this step is existing technology and will not be described in detail here.

[0137] S507: Based on the difference between the predicted operation data and the audience account operation data tags, enhance the information delivery recall model until the training termination condition is met.

[0138] S509: The delivery information recall model at the end of training is determined as the enhanced delivery information recall model.

[0139] Based on the difference between the predicted operation data and the audience account operation data tags, the loss value can be determined first. Then, backpropagation is performed based on the loss value to update the parameters of the account information representation network and the delivery information representation network. Iterative training is carried out until the training termination condition is met, and the delivery information retrieval model at the end of training is determined as the enhanced delivery information retrieval model. The training termination condition can be reaching a preset number of iterations or the loss value being less than a preset loss threshold, etc. This application does not limit the training termination condition.

[0140] In this embodiment, a dual-tower model is constructed to decouple sample audience accounts from sample delivery information, enabling offline calculation of the representation vector corresponding to the delivery information during application. Furthermore, by introducing enhanced live-stream sample features into the training process of the delivery information representation network, this application ensures that natural traffic domain data is fully utilized during training, thereby improving the accuracy of model training.

[0141] Figure 6 This is a schematic diagram of the overall process of training the information retrieval model provided in the embodiments of this application.

[0142] like Figure 6 As shown, an initial KNN graph is first constructed from multiple sample live streams. This initial KNN graph is then input into a pre-defined clustering model to obtain the target sample cluster nodes. It should be noted that the pre-defined clustering model is a pre-trained Hi-LANDER model. Specifically, the Hi-LANDER model includes multiple clustering modules, which share parameters. After processing the initial KNN graph through the first clustering module, a set of predicted edges is output. The nodes in the initial KNN graph are then connected according to the predicted edges in this set, resulting in multiple connected components. Next, for each connected component, the nodes are aggregated into an intermediate sample cluster node. A new KNN graph is then constructed based on these intermediate sample cluster nodes and used as input for the next level of clustering modules; this process continues until the target sample cluster nodes are finally obtained through this hierarchical clustering method.

[0143] Next, based on the clustering nodes of the target samples, it can be traced which sample live streams belong to the same set, thus determining multiple sample live stream sets. Then, live streams without associated advertising information are associated with the advertising information associated with live streams in the same sample live stream set, forming new training samples. These new training samples are combined with the training samples corresponding to the live streams originally associated with advertising information to obtain the live stream advertising sample data. Additionally, enhanced live stream sample features can be generated based on the intermediate sample clustering nodes during the clustering process. See [link to process details] for details. Figure 4 The embodiments shown are not described in detail here.

[0144] It should be noted that the training of the Hi-LANDER model may include: obtaining multiple training samples, which are live streaming rooms, where training samples belonging to the same advertiser (merchant) are set with the same label; then inputting the training samples into the initial Hi-LANDER model for cluster prediction, and training the model based on the difference between the cluster prediction results and the labels, thereby obtaining the pre-trained Hi-LANDER model, i.e. the preset clustering model.

[0145] In this embodiment, the ad delivery information recall model can be a dual-tower model, including an account tower and an ad delivery information tower. Sample account attribute information is input into the account tower for account feature extraction, resulting in a sample account representation vector. Live stream ad delivery sample data is input into the ad delivery information tower, and the output of the ad delivery information tower is fused with the enhanced live stream sample features using a residual connection, resulting in a sample ad delivery information representation vector. It should be noted that this feature fusion process can involve concatenating the output of the ad delivery information tower with the enhanced live stream sample features before feeding it into a fully connected (Dense) layer. Then, the dot product between the sample account representation vector and the sample ad delivery information representation vector is calculated to obtain the sample audience account's prediction operation data for the sample ad delivery information, which is used to train the ad delivery information recall model, resulting in the enhanced ad delivery information recall model. It should be noted that the organic traffic domain and the ad domain essentially belong to the same traffic but different scenarios, exhibiting high homogeneity in account interests and conversion intentions. Therefore, data from the organic domain can improve the model capabilities of the ad domain. However, the data from the organic traffic domain and the ad domain are heterogeneous. Typically, data from organic traffic differs significantly from that from advertising data in terms of product and feature systems. Furthermore, advertising data possesses unique characteristics not present in organic traffic data, such as ad creatives and bidding strategies. This application addresses this by using clustering to reuse and associate ads within the advertising domain and integrating enhanced features from live streaming rooms, thus fully utilizing organic traffic data to improve the accuracy of the advertising domain's ad delivery information retrieval model.

[0146] This application embodiment also provides a method for recalling live broadcast room delivery information, the method comprising:

[0147] Obtain the target account attribute information of the target audience account and the live room delivery candidate data corresponding to the candidate delivery information; among them, the live room delivery candidate data corresponding to the candidate delivery information can be obtained by querying the live room delivery sample data corresponding to the sample delivery information saved during the training process.

[0148] The target account attribute information and live broadcast room delivery candidate data are input into the enhanced delivery information recall model to perform delivery information operation data prediction processing and obtain the prediction result; the prediction result is used to characterize whether to recommend the candidate delivery information to the target audience account; wherein, the enhanced delivery information recall model is trained according to the live broadcast room information processing method of the above embodiment.

[0149] By using the enhanced delivery information recall model trained with the live broadcast room information processing method described above, the accuracy of delivery information recall and recommendation is improved, which is conducive to improving the click-through rate, conversion rate and other indicators of the delivery information associated with the live broadcast room for the target audience account.

[0150] In one embodiment, candidate live-streaming data corresponding to candidate delivery information can be input offline into the delivery information representation network of the enhanced delivery information retrieval model to obtain candidate delivery information representation vectors, which are then stored in the delivery information representation vector library. When online delivery information retrieval is performed for target audience accounts, the target account attribute information of the target audience account is input into the account information representation network of the enhanced delivery information retrieval model for account feature extraction processing to obtain the target account representation vector. Then, vector indexing and matching are performed from the delivery information representation vector library to obtain the target delivery information representation vector corresponding to the target account representation vector, thereby determining the target delivery information. This saves the step of vector representation of candidate delivery information and effectively improves the retrieval processing efficiency.

[0151] In one embodiment, enhanced candidate features for live streaming rooms can also be introduced for ad delivery information retrieval. Specifically, enhanced candidate features for live streaming rooms can be obtained by clustering candidate live streaming rooms corresponding to ad delivery information as initial candidate nodes to obtain intermediate candidate cluster nodes; the enhanced candidate features are then determined based on the initial candidate nodes and intermediate candidate cluster nodes. The live streaming room ad delivery candidate data is then input into the ad delivery information representation network of the enhanced ad delivery information retrieval model, and the output of the ad delivery information representation network is fused with the enhanced candidate features for live streaming rooms to obtain candidate ad delivery information representation vectors, which are stored in the ad delivery information representation vector library for subsequent vector indexing and matching.

[0152] Figure 7 This is a schematic diagram of the structure of the live streaming room information processing device provided in the embodiments of this application, as shown below. Figure 7 As shown, the device 700 may include:

[0153] The sample live streaming room clustering module 701 is used to cluster multiple sample live streaming rooms as nodes to obtain a set of multiple sample live streaming rooms.

[0154] The target live streaming room set determination module 702 is used to filter the target live streaming room set to which the live streaming room associated with the advertising information belongs from the plurality of sample live streaming room sets.

[0155] The delivery information association module 703 is used to associate the live rooms in the target live room set that are not associated with delivery information with the delivery information associated with the live rooms in the same target live room set for each target live room set, so as to obtain live room delivery sample data.

[0156] The information recall sample data generation module 704 is used to generate information recall sample data based on the viewer account operation data corresponding to the live broadcast room in the live broadcast room delivery sample data.

[0157] The model training module 705 is used to enhance the delivery information recall model through the information recall sample data to obtain the enhanced delivery information recall model.

[0158] In some embodiments, the sample live streaming room clustering module may include:

[0159] The sample attribute feature set acquisition submodule is used to acquire the sample attribute features corresponding to each of the multiple sample live rooms to obtain the sample attribute feature set.

[0160] The initial sample topology graph construction submodule is used to construct an initial sample topology graph based on the sample attribute feature set; the initial sample topology graph includes multiple initial sample nodes and multiple initial sample connection edges; each initial sample node represents a sample live broadcast room; each initial sample connection edge represents the association relationship between two connected initial sample nodes;

[0161] The target sample clustering node determination submodule is used to perform node aggregation processing on multiple initial sample nodes according to the initial sample topology graph to obtain the target sample clustering nodes;

[0162] The sample live room set determination submodule is used to determine the sample live room set based on the correspondence between the target sample clustering node and the initial sample node, and the correspondence between the initial sample node and the sample live room.

[0163] In some embodiments, the target sample clustering node determination submodule may include:

[0164] A clustering model acquisition unit is used to acquire a preset clustering model, wherein the preset clustering model contains multiple clustering modules arranged in sequence;

[0165] The current clustering module determination unit is used to take the initial sample topology map as the current sample topology map and to determine the first clustering module among the plurality of sequentially arranged clustering modules as the current clustering module;

[0166] The intermediate sample clustering node determination unit is used to perform graph clustering processing on the current sample topology graph based on the current clustering module to obtain multiple intermediate sample clustering nodes;

[0167] An intermediate sample topology graph construction unit is used to construct an intermediate sample topology graph based on the plurality of intermediate sample clustering nodes;

[0168] The repeating unit is used to re-use the intermediate sample topology graph as the current sample topology graph; and to re-use the next clustering module arranged after the current clustering module as the current clustering module; repeatedly perform graph clustering processing on the current sample topology graph based on the current clustering module to obtain multiple intermediate sample clustering nodes; and construct an intermediate sample topology graph based on the multiple intermediate sample clustering nodes until the clustering termination condition is met;

[0169] The target sample clustering node determination unit is used to determine the intermediate sample clustering node output by the clustering module at the end of the clustering process as the target sample clustering node.

[0170] In some embodiments, the information recall sample data generation module may include:

[0171] The sample augmentation node determination submodule is used to determine the intermediate sample clustering node corresponding to the live room in the live room delivery sample data as the sample augmentation node.

[0172] The live room enhancement sample feature determination submodule is used to determine the live room enhancement sample features corresponding to the live room in the live room delivery sample data based on the sample enhancement node and the initial sample node corresponding to the live room in the live room delivery sample data.

[0173] The sample account attribute information acquisition submodule is used to acquire the sample account attribute information corresponding to the live room in the live room delivery sample data; the sample account attribute information corresponds to the same sample viewer account as the viewer account operation data;

[0174] The information recall sample data generation submodule is used to generate the information recall sample data based on the sample account attribute information, the live broadcast room delivery sample data, and the live broadcast room enhanced sample features. The information recall sample data is labeled with audience account operation data tags, which represent the actual operation data of the sample audience account on the sample delivery information. The sample delivery information is the delivery information associated with the live broadcast room in the live broadcast room delivery sample data.

[0175] In some embodiments, the delivery information recall model includes an account information representation network and a delivery information representation network; the model training module may include:

[0176] The account feature extraction submodule is used to input the sample account attribute information into the account information representation network for account feature extraction processing to obtain the sample account representation vector.

[0177] The delivery information feature enhancement submodule is used to input the live broadcast room delivery sample data into the delivery information representation network, and to fuse the output of the delivery information representation network with the live broadcast room enhanced sample features to obtain the sample delivery information representation vector.

[0178] The prediction operation data determination submodule is used to determine the vector similarity between the sample account representation vector and the sample delivery information representation vector as the prediction operation data of the sample audience account for the sample delivery information.

[0179] The enhanced training submodule is used to enhance the delivery information recall model based on the difference between the predicted operation data and the audience account operation data tags until the training termination condition is met.

[0180] The enhanced delivery information recall model determination submodule is used to determine the delivery information recall model at the end of training as the enhanced delivery information recall model.

[0181] In some embodiments, the intermediate sample clustering node determination unit may include:

[0182] The edge prediction subunit is used to input the current sample topology graph into the current clustering module for edge prediction processing to obtain a set of predicted connection edges.

[0183] A connectivity component determination unit is used to connect the nodes in the current sample topology graph according to the predicted connection edges contained in the predicted connection edge set to obtain multiple connectivity components.

[0184] The node aggregation unit is used to perform node aggregation processing on the nodes contained in each connected component to obtain the intermediate clustering nodes of the multiple samples.

[0185] In some embodiments, the edge prediction subunit may include:

[0186] The node encoding subunit is used to perform node encoding processing on the nodes contained in the current sample topology graph to obtain the encoding features corresponding to each node in the current sample topology graph.

[0187] The parameter determination subunit is used to traverse the nodes contained in the current sample topology graph, and determine the node density of each traversed current node and the correlation coefficient corresponding to the connection edge between the current node and its neighboring nodes based on the encoding features corresponding to each node in the current sample topology graph; the correlation coefficient represents the correlation degree between the current node and its neighboring nodes; the neighboring nodes are the nodes in the current sample topology graph that have a connection edge with the current node.

[0188] The candidate edge determination sub-unit is used to determine the connection edge between the current node and the neighboring node as a candidate edge if the correlation coefficient is greater than or equal to a preset coefficient threshold.

[0189] The predictive connection edge determination subunit is used to determine the predicted connection edge from the candidate edges based on the correlation coefficient corresponding to each candidate edge, thereby obtaining the predicted connection edge set.

[0190] In some embodiments, the parameter determination subunit may include:

[0191] The encoding feature acquisition subunit is used to acquire the first encoding feature corresponding to the current node and the second encoding feature corresponding to the neighboring node;

[0192] The connection probability determination subunit is used to concatenate the first encoded feature and the second encoded feature, and determine the connection probability based on the concatenated encoded feature; the connection probability represents the probability that the current node and the neighboring node belong to the same cluster;

[0193] The correlation coefficient determination subunit is used to determine the correlation coefficient corresponding to the connection edge between the current node and the neighboring node based on the connection probability.

[0194] The node density determination subunit is used to determine the node density of the current node based on the node similarity between the current node and the neighboring nodes in the current sample topology graph and the correlation coefficient corresponding to the connecting edge between the current node and the neighboring nodes.

[0195] In some embodiments, the delivery information association module may include:

[0196] The first delivery feature acquisition submodule is used to acquire the first delivery feature of the delivery information associated with the live rooms in the target live room set;

[0197] The feature reuse submodule is used to determine the first feature as the second feature corresponding to the live room in the target live room set that has no associated feature information;

[0198] The live streaming room delivery sample data determination submodule is used to determine the live streaming room delivery sample data based on the first delivery feature and the second delivery feature.

[0199] In some embodiments, the information recall sample data generation module may further include:

[0200] The account attribute information acquisition submodule is used to acquire the sample account attribute information corresponding to the live room in the live room delivery sample data; the sample account attribute information corresponds to the same sample viewer account as the viewer account operation data;

[0201] The information recall sample data determination submodule is used to generate the information recall sample data based on the sample account attribute information and the live broadcast room delivery sample data; the information recall sample data is labeled with audience account operation data tags, the audience account operation data tags represent the actual operation data of the sample audience account on the sample delivery information, and the sample delivery information is the delivery information associated with the live broadcast room in the live broadcast room delivery sample data.

[0202] Figure 8 This is a schematic diagram of the live streaming room information recall device provided in an embodiment of this application. Figure 8 As shown, the device 800 may include:

[0203] Module 801 is used to obtain the target account attribute information and live broadcast room delivery candidate data corresponding to the candidate delivery information of the target audience account.

[0204] The processing module 802 is used to input the target account attribute information and the live broadcast room delivery candidate data into the enhanced delivery information recall model, perform delivery information operation data prediction processing, and obtain a prediction result; the prediction result is used to characterize whether to recommend the candidate delivery information to the target audience account; the enhanced delivery information recall model is trained according to the live broadcast room information processing method described above.

[0205] The apparatus and method embodiments described herein are based on the same inventive concept.

[0206] This application provides an electronic device, which includes a processor and a memory. The memory stores at least one instruction or at least one program. The at least one instruction or at least one program is loaded and executed by the processor to implement the live room information processing method or the live room delivery information recall method provided in the above method embodiments.

[0207] The embodiments of this application also provide a computer storage medium, which can be disposed in a terminal to store at least one instruction or at least one program related to implementing a live room information processing method or a live room information recall method in the method embodiments. The at least one instruction or at least one program is loaded and executed by the processor to implement the live room information processing method or the live room information recall method provided in the above method embodiments.

[0208] Embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the live-streaming information processing method or the live-streaming information recall method provided in the above-described method embodiments.

[0209] Optionally, in this embodiment, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0210] The memory described in this application embodiment can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for the functions, etc.; the data storage area may store data created according to the use of the device, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory may also include a memory controller to provide the processor with access to the memory.

[0211] The live streaming information processing method provided in this application can be executed on a mobile terminal, computer terminal, server, or similar computing device. Taking running on a server as an example... Figure 9 This is a hardware structure block diagram of a server for a live streaming information processing method provided in an embodiment of this application. For example... Figure 9As shown, the server 900 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 910 (CPUs 910 may include, but are not limited to, microprocessors (MCUs) or programmable logic devices (FPGAs), a memory 930 for storing data, and one or more storage media 920 (e.g., one or more mass storage devices) for storing application programs 923 or data 922. The memory 930 and storage media 920 may be temporary or persistent storage. The program stored in the storage media 920 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the CPU 910 may be configured to communicate with the storage media 920 and execute the series of instruction operations stored in the storage media 920 on the server 900. Server 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input / output interfaces 940, and / or one or more operating systems 921, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0212] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 900. In one example, the input / output interface 940 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 940 may be a radio frequency (RF) module for wireless communication with the Internet.

[0213] Those skilled in the art will understand that Figure 9 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 900 may also include... Figure 9 The more or fewer components shown, or having the same Figure 9 The different configurations shown.

[0214] As can be seen from the embodiments of the live room information processing method, live room delivery information recall method, device, equipment, or storage medium provided in this application, this application obtains multiple sample live room sets by clustering multiple sample live rooms as nodes; in the multiple sample live room sets, the target live room set to which the live rooms associated with delivery information belong are selected; for each target live room set, the live rooms in the target live room set that are not associated with delivery information are associated with the delivery information associated with the live rooms in the same target live room set to obtain live room delivery sample data; the clustering method enables the data of live rooms that are not associated with delivery information to also be utilized. Based on the viewer account operation data corresponding to the live room in the live room delivery sample data, information recall sample data is generated. The delivery information recall model is enhanced and trained using the information recall sample data to obtain an enhanced delivery information recall model. Compared with the traditional solution that only uses the live room data associated with the delivery information to train the delivery information recall model, this application makes full use of the data of live rooms (live rooms not associated with delivery information) in the natural traffic domain, making the live room delivery sample data richer, thereby improving the accuracy of live room information processing and the accuracy of the delivery information recall model.

[0215] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0216] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0217] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer storage medium, such as a read-only memory, a disk, or an optical disk.

[0218] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for processing information in a live streaming room, characterized in that, The method includes: Clustering is performed using multiple sample live streaming rooms as nodes to obtain a set of multiple sample live streaming rooms; Among the multiple sample live streaming room sets, filter the target live streaming room set to which the live streaming room with associated advertising information belongs; For each set of target live streaming rooms, the live streaming rooms in the target live streaming room set that are not associated with the advertising information are associated with the advertising information associated with the live streaming rooms in the same set of target live streaming rooms to obtain live streaming room advertising sample data. Based on the viewer account operation data corresponding to the live broadcast room in the live broadcast room delivery sample data, information recall sample data is generated. The enhanced information retrieval model is obtained by training the information retrieval sample data.

2. The method according to claim 1, characterized in that, The process of clustering multiple sample live streaming rooms as nodes to obtain a set of multiple sample live streaming rooms includes: Obtain the sample attribute features corresponding to each of the multiple sample live streaming rooms to obtain the sample attribute feature set; Based on the sample attribute feature set, an initial sample topology graph is constructed; the initial sample topology graph includes multiple initial sample nodes and multiple initial sample connection edges; each initial sample node represents a sample live broadcast room; each initial sample connection edge represents the association relationship between two connected initial sample nodes; Based on the initial sample topology graph, node aggregation processing is performed on multiple initial sample nodes to obtain target sample clustering nodes; The set of sample live streaming rooms is determined based on the correspondence between the target sample clustering nodes and the initial sample nodes, as well as the correspondence between the initial sample nodes and the sample live streaming rooms.

3. The method according to claim 2, characterized in that, The step of performing node aggregation processing on multiple initial sample nodes based on the initial sample topology graph to obtain target sample clustering nodes includes: Obtain a preset clustering model, which includes multiple clustering modules arranged sequentially; The initial sample topology graph is used as the current sample topology graph, and the first clustering module among the multiple sequentially arranged clustering modules is determined as the current clustering module; Based on the current clustering module, graph clustering processing is performed on the current sample topology graph to obtain multiple intermediate sample clustering nodes; Construct an intermediate sample topology graph based on the multiple intermediate sample clustering nodes; The intermediate sample topology graph is used as the current sample topology graph again; and the next clustering module after the current clustering module is used as the current clustering module again; the graph clustering process of the current sample topology graph based on the current clustering module is repeated to obtain multiple intermediate sample clustering nodes; the intermediate sample topology graph is constructed based on the multiple intermediate sample clustering nodes until the clustering termination condition is met; The intermediate sample clustering node output by the clustering module at the end of the clustering process is determined as the target sample clustering node.

4. The method according to claim 3, characterized in that, The step of generating information recall sample data based on the viewer account operation data corresponding to the live broadcast room in the live broadcast room delivery sample data includes: The intermediate sample clustering nodes corresponding to the live room in the live room delivery sample data are determined as sample augmentation nodes; Based on the sample enhancement node and the initial sample node corresponding to the live room in the live room delivery sample data, determine the live room enhancement sample features corresponding to the live room in the live room delivery sample data. Obtain the sample account attribute information corresponding to the live room in the live room delivery sample data; the sample account attribute information corresponds to the same sample audience account as the audience account operation data; Based on the sample account attribute information, the live broadcast room delivery sample data, and the live broadcast room enhanced sample features, the information recall sample data is generated; the information recall sample data is labeled with audience account operation data tags, the audience account operation data tags represent the actual operation data of the sample audience account on the sample delivery information, and the sample delivery information is the delivery information associated with the live broadcast room in the live broadcast room delivery sample data.

5. The method according to claim 4, characterized in that, The delivery information recall model includes an account information representation network and a delivery information representation network; The step of enhancing the information retrieval model by using the information retrieval sample data to obtain the enhanced information retrieval model includes: The sample account attribute information is input into the account information representation network for account feature extraction processing to obtain the sample account representation vector. The live broadcast room delivery sample data is input into the delivery information representation network, and the output of the delivery information representation network is fused with the live broadcast room enhanced sample features to obtain the sample delivery information representation vector. The vector similarity between the sample account representation vector and the sample delivery information representation vector is determined as the prediction operation data of the sample audience account for the sample delivery information. Based on the difference between the predicted operation data and the audience account operation data tags, the delivery information recall model is enhanced and trained until the training termination condition is met. The delivery information recall model at the end of training is determined as the enhanced delivery information recall model.

6. The method according to claim 3, characterized in that, The process of performing graph clustering on the current sample topology graph based on the current clustering module to obtain multiple intermediate sample clustering nodes includes: The current sample topology graph is input into the current clustering module for edge prediction processing to obtain a set of predicted connection edges. The nodes in the current sample topology graph are connected according to the predicted connection edges contained in the predicted connection edge set to obtain multiple connected components; For each connected component, the nodes contained in the connected component are subjected to node aggregation processing to obtain the intermediate cluster nodes of the multiple samples.

7. The method according to claim 6, characterized in that, The step of inputting the current sample topology graph into the current clustering module for edge prediction processing to obtain a set of predicted connection edges includes: Node encoding processing is performed on the nodes contained in the current sample topology graph to obtain the encoding features corresponding to each node in the current sample topology graph; The nodes in the current sample topology graph are traversed. Based on the encoding features corresponding to each node in the current sample topology graph, the node density of each traversed current node and the correlation coefficient corresponding to the connection edge between the current node and its neighboring nodes are determined. The correlation coefficient represents the correlation degree between the current node and its neighboring nodes. The neighboring nodes are the nodes in the current sample topology graph that have a connection edge with the current node. If the correlation coefficient is greater than or equal to a preset coefficient threshold, the connection edge between the current node and the neighboring node is determined as a candidate edge; Based on the correlation coefficient corresponding to each candidate edge, the predicted connection edge is determined from the candidate edges to obtain the set of predicted connection edges.

8. The method according to claim 7, characterized in that, The process of determining the node density of each traversed current node and the correlation coefficient of the connection edges between the current node and its neighbors based on the encoding features corresponding to each node in the current sample topology graph includes: Obtain the first encoded feature corresponding to the current node and the second encoded feature corresponding to the neighboring node; The first encoded feature and the second encoded feature are concatenated, and the connection probability is determined based on the concatenated encoded feature; the connection probability represents the probability that the current node and the neighboring node belong to the same cluster; The association coefficient corresponding to the connection edge between the current node and the neighboring node is determined based on the connection probability; The node density of the current node is determined based on the node similarity between the current node and its neighboring nodes in the current sample topology graph and the correlation coefficient corresponding to the connecting edges between the current node and its neighboring nodes.

9. The method according to claim 1, characterized in that, The step of associating live streams without associated advertising information in the target live stream set with the advertising information associated with live streams in the same target live stream set to obtain live stream advertising sample data includes: Obtain the first delivery feature of the delivery information associated with the live rooms in the target live room set; The first delivery feature is determined as the second delivery feature corresponding to the live room in the target live room set that has no associated delivery information; The live streaming room delivery sample data is determined based on the first delivery feature and the second delivery feature.

10. The method according to claim 1, characterized in that, The step of generating information recall sample data based on the viewer account operation data corresponding to the live broadcast room in the live broadcast room delivery sample data includes: Obtain the sample account attribute information corresponding to the live room in the live room delivery sample data; the sample account attribute information corresponds to the same sample audience account as the audience account operation data; The information recall sample data is generated based on the sample account attribute information and the live broadcast room delivery sample data; the information recall sample data is labeled with audience account operation data tags, the audience account operation data tags represent the actual operation data of the sample audience account on the sample delivery information, and the sample delivery information is the delivery information associated with the live broadcast room in the live broadcast room delivery sample data.

11. A method for recalling information delivered in a live streaming room, characterized in that, The method includes: Obtain the target audience's account attribute information and the corresponding live stream advertising candidate data; The target account attribute information and the live broadcast room delivery candidate data are input into the enhanced delivery information recall model to perform delivery information operation data prediction processing to obtain a prediction result; the prediction result is used to characterize whether to recommend the candidate delivery information to the target audience account; the enhanced delivery information recall model is trained according to the live broadcast room information processing method according to any one of claims 1-10.

12. A live streaming room information processing device, characterized in that, The device includes: The sample live streaming room clustering module is used to cluster multiple sample live streaming rooms as nodes to obtain a set of multiple sample live streaming rooms. The target live streaming room set determination module is used to filter the target live streaming room set to which the live streaming room associated with the advertising information belongs from the multiple sample live streaming room sets; The delivery information association module is used to associate the live rooms in the target live room set that are not associated with delivery information with the delivery information associated with the live rooms in the same target live room set for each target live room set, so as to obtain live room delivery sample data. The information recall sample data generation module is used to generate information recall sample data based on the viewer account operation data corresponding to the live broadcast room in the live broadcast room delivery sample data. The model training module is used to enhance the delivery information recall model using the information recall sample data to obtain an enhanced delivery information recall model.

13. A live-streaming room information recall device, characterized in that, The device includes: The acquisition module is used to acquire the target account attribute information and candidate data for live streaming room delivery corresponding to the target audience account and the candidate delivery information. The processing module is used to input the target account attribute information and the live broadcast room delivery candidate data into the enhanced delivery information recall model, perform delivery information operation data prediction processing, and obtain a prediction result; the prediction result is used to characterize whether to recommend the candidate delivery information to the target audience account; the enhanced delivery information recall model is trained according to any one of the live broadcast room information processing methods according to claims 1-10.

14. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the live room information processing method as described in any one of claims 1-10, or the live room delivery information recall method as described in claim 11.

15. A computer storage medium, characterized in that, The computer storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the live room information processing method as described in any one of claims 1-10, or the live room delivery information recall method as described in claim 11.