Data processing method and device, equipment and storage medium

By determining the combination of data to be processed based on object clusters, content clusters, and auxiliary information in the data processing device, calling the data matching model to generate the target data combination and sending personalized content, the problem that existing technologies cannot meet personalized needs is solved, thereby improving click-through rate and conversion rate.

CN116644224BActive Publication Date: 2026-05-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-02-14
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing data processing equipment cannot meet personalized needs, resulting in low click-through rates or conversion rates, and the accuracy of existing matching based on historical data is not high.

Method used

Based on object clusters, content clusters, and auxiliary information, the combination of data to be processed is determined, the data matching model is called to process the feature representation information, the target data combination is generated, and personalized content is sent.

Benefits of technology

Personalized content delivery improved the accuracy of intelligent data delivery, leading to higher click-through rates and conversion rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116644224B_ABST
    Figure CN116644224B_ABST
Patent Text Reader

Abstract

The application discloses a data processing method, device and equipment and a storage medium. The method comprises the following steps: determining a plurality of to-be-processed data combinations based on an object cluster, a content cluster and auxiliary information, and acquiring feature representation information of each to-be-processed data combination in the plurality of to-be-processed data combinations; calling a data matching model to process the feature representation information of each to-be-processed data combination, and obtaining a target data combination of each object in at least one object, wherein the target data combination comprises target auxiliary information and target content; and sending the target content to each object based on the target auxiliary information. Through the embodiment of the application, the click rate or conversion rate of the object to the content can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer software technology, and in particular to a data processing method, apparatus, device and storage medium. Background Technology

[0002] With the development of mobile internet technology, the automation and intelligence of data matching have also made significant progress. In existing solutions, data processing devices can cluster objects into multiple object groups and send specified content to each object within each group. However, in this case, the data processing device sends the same content to every object within a group, failing to meet personalized needs. Alternatively, the data processing device can match corresponding content for each object based on historical data, such as matching based on object features and content features. This method has low accuracy, leading to low click-through rates or conversion rates. Summary of the Invention

[0003] This application provides a data processing method, apparatus, device, and storage medium that can effectively improve the click-through rate or conversion rate of content while meeting personalized needs.

[0004] On one hand, embodiments of this application provide a data processing method, which includes:

[0005] Multiple data combinations to be processed are determined based on object clusters, content clusters, and auxiliary information. The object cluster includes at least one object, the content cluster includes at least one piece of content, and the auxiliary information includes one or two of time and channel.

[0006] Obtain the feature representation information of each data combination to be processed from multiple data combinations to be processed;

[0007] The data matching model is invoked to process the feature representation information of each data combination to be processed, so as to obtain the target data combination of each object in at least one object. The target data combination includes target auxiliary information and target content.

[0008] Send target content to each object based on target auxiliary information.

[0009] On the other hand, embodiments of this application provide a data processing apparatus, which includes:

[0010] The determining unit is used to determine multiple combinations of data to be processed based on object clusters, content clusters, and auxiliary information. The object cluster includes at least one object, the content cluster includes at least one piece of content, and the auxiliary information includes one or two of time and channel.

[0011] The acquisition unit is used to acquire the feature representation information of each data combination to be processed from multiple data combinations to be processed.

[0012] The processing unit is used to call the data matching model to process the feature representation information of each data combination to be processed, and to obtain the target data combination of each object in at least one object. The target data combination includes target auxiliary information and target content.

[0013] The sending unit is used to send target content to each object based on target auxiliary information.

[0014] Furthermore, embodiments of this application provide a data processing device, which includes an input interface and an output interface, and further includes:

[0015] A processor, adapted to implement one or more instructions; and,

[0016] A computer storage medium that stores one or more instructions, which are adapted to be loaded by a processor and executed as follows:

[0017] Multiple data combinations to be processed are determined based on object clusters, content clusters, and auxiliary information. The object cluster includes at least one object, the content cluster includes at least one piece of content, and the auxiliary information includes one or two of time and channel.

[0018] Obtain the feature representation information of each data combination to be processed from multiple data combinations to be processed;

[0019] The data matching model is invoked to process the feature representation information of each data combination to be processed, so as to obtain the target data combination of each object in at least one object. The target data combination includes target auxiliary information and target content.

[0020] Send target content to each object based on target auxiliary information.

[0021] Furthermore, this application provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the following steps:

[0022] Multiple data combinations to be processed are determined based on object clusters, content clusters, and auxiliary information. The object cluster includes at least one object, the content cluster includes at least one piece of content, and the auxiliary information includes one or two of time and channel.

[0023] Obtain the feature representation information of each data combination to be processed from multiple data combinations to be processed;

[0024] The data matching model is invoked to process the feature representation information of each data combination to be processed, so as to obtain the target data combination of each object in at least one object. The target data combination includes target auxiliary information and target content.

[0025] Send target content to each object based on target auxiliary information.

[0026] In another aspect, embodiments of this application provide a computer storage medium storing one or more instructions, which are adapted to be loaded by a processor and executed as follows:

[0027] Multiple data combinations to be processed are determined based on object clusters, content clusters, and auxiliary information. The object cluster includes at least one object, the content cluster includes at least one piece of content, and the auxiliary information includes one or two of time and channel.

[0028] Obtain the feature representation information of each data combination to be processed from multiple data combinations to be processed;

[0029] The data matching model is invoked to process the feature representation information of each data combination to be processed, so as to obtain the target data combination of each object in at least one object. The target data combination includes target auxiliary information and target content.

[0030] Send target content to each object based on target auxiliary information.

[0031] In this embodiment, the data processing device can determine multiple data combinations to be processed based on at least one object in an object cluster, at least one piece of content in a content cluster, and at least one time and at least one channel in auxiliary information. It then calls a data matching model to process the feature representation information of each data combination to obtain a matching score for each data combination. Based on the matching score of each data combination, a target data combination for each object can be determined, allowing target content to be sent to each object based on the target auxiliary information included in the target data combination. Since the data combinations to be processed include auxiliary information, the target data combinations also correspondingly include target auxiliary information. Unlike existing schemes that randomly send target content to each object, this embodiment's data processing device effectively improves the accuracy of intelligent data delivery by sending target content to each object based on the target auxiliary information included in the target data combination, thereby effectively increasing the click-through rate or conversion rate of the content. Furthermore, this embodiment determines the target data combination for each object based on the matching score of each data combination to be processed, achieving human-level data processing. Unlike existing schemes that send the same content to each object in an object group, this allows for the determination of different target content for different objects, fulfilling diverse and personalized needs. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is a schematic diagram of the architecture of a data processing system provided in an embodiment of this application;

[0034] Figure 2 This is a flowchart illustrating a data processing method provided in an embodiment of this application;

[0035] Figure 3 This is a flowchart illustrating another data processing method provided in an embodiment of this application;

[0036] Figure 4 This is a schematic diagram of the architecture of a data processing method provided in an embodiment of this application;

[0037] Figure 5 This is a schematic diagram of the training process of a data matching model provided in an embodiment of this application;

[0038] Figure 6 This is a schematic diagram of the architecture of another data processing system provided in an embodiment of this application;

[0039] Figure 7 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;

[0040] Figure 8 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Detailed Implementation

[0041] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0042] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning. Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance.

[0043] With the research and advancement of artificial intelligence (AI) technology, AI has been applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, autonomous driving, drones, robots, smart healthcare, and smart customer service. Beyond these applications, AI can also be used in other areas, such as machine learning to send intelligent data. This application proposes a machine learning-based data processing method. This method allows a data processing device to build a data matching model based on machine learning algorithms. When multiple data combinations to be processed are determined based on object clusters, content clusters, and auxiliary information, the data matching model can be invoked to process the feature representation information of each data combination, obtaining a target data combination for each object. This allows for the intelligent sending of data based on the target data combination for each object. Matching different target data combinations to different objects can meet the personalized needs of different objects. Furthermore, the target data combination includes auxiliary information; sending corresponding content to each object based on this auxiliary information can effectively improve the click-through rate or conversion rate of the intelligent data.

[0044] In one embodiment, the data processing method can be applied to, for example... Figure 1 In the data processing system shown, such as Figure 1 As shown, the data processing system may include at least: a data processing device 100 and a terminal device 200. The data processing device 100 can be any device with data processing capabilities; for example, the data processing device 100 may be... Figure 1The server shown can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, content delivery networks (CDN), middleware services, domain name services, security services, and big data and artificial intelligence platforms, etc. Each terminal device 200 is associated with an object in the object cluster, and the data processing device 100 can send target content to the terminal device 200 associated with each object based on the target auxiliary information included in the target data combination. The terminal device can be, but is not limited to, smart terminals such as smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, and wearable devices.

[0045] Specifically, the data processing device 100 can determine multiple data combinations to be processed based on at least one object included in the object cluster, at least one piece of content included in the content cluster, and at least one time and at least one channel included in the auxiliary information. It then calls a data matching model to process the feature representation information of each data combination to obtain a matching score for each data combination. Based on the matching score of each data combination, it can determine the target data combination for each object, so that target content can be sent to the terminal device 200 corresponding to each object based on the target auxiliary information included in the target data combination. Since the data combinations to be processed include auxiliary information, the target data combinations also correspondingly include target auxiliary information. Unlike existing schemes that randomly send target content to each object, the data processing device 100 in this embodiment can effectively improve the accuracy of intelligent data transmission by sending target content to each object based on the target auxiliary information included in the target data combination, thereby effectively improving click-through rates or conversion rates. Furthermore, this embodiment determines the target data combination for each object based on the matching score of each data combination to be processed, achieving human-level data processing. Unlike existing schemes that send the same content to each object in a group of objects, it can determine different target content for different objects, achieving diversified and personalized needs.

[0046] Please see Figure 2 , Figure 2 This is a schematic flowchart of a data processing method provided in an embodiment of this application. This data processing method can be executed by the data processing device mentioned above. Figure 2 As shown, the data processing method may include S201-S204:

[0047] S201: Determine multiple combinations of data to be processed based on object clusters, content clusters, and auxiliary information. The object cluster includes at least one object, the content cluster includes at least one piece of content, and the auxiliary information includes one or both of time and channel.

[0048] The object cluster can include at least one object, the number of which can be determined by the actual application scenario. The at least one object in the object cluster can be any object. Optionally, the data processing device can determine at least one object based on registered user accounts. For example, the data processing device can determine at least one object based on user accounts registered by the client within a certain time period. Alternatively, the data processing device can determine at least one object based on all user accounts registered within the client. Optionally, the data processing device can also determine at least one object based on targeted object tags. Specifically, the data processing device can obtain data information generated by user accounts registered by the client, extract user tags from the user account data information, obtain preset targeted object tags, which are tags possessed by objects that meet the targeted feature requirements, and extract at least one object matching the targeted object tags from the registered user accounts based on the user tags. The preset targeted object tags are screening criteria; therefore, different screening criteria will result in different targeted object tags. For example, targeted object features can be set as maternal and infant characteristics, game characteristics, or automotive characteristics, etc. Optionally, the data processing device can also analyze the data information generated by registered user accounts based on a machine learning model to determine at least one object. The machine learning model may include, but is not limited to, one or more of the following: potential customer mining models, multi-gate mixture-of-experts (MMOE) models, and light gradient boosting machine (LightGBM) models. The data information may include, but is not limited to, one or more of the following: material data, order data, and member data.

[0049] The content cluster may include at least one piece of content, the number of which may be determined by the actual application scenario. At least one piece of content in the content cluster may have any type of presentation, including but not limited to one or more of text, images, audio, and video.

[0050] The auxiliary information may include one or more of time and channels. That is, the auxiliary information may include time, channels, or both. Time can be any one of a set of times, which may include different types of time. For example, it can be divided into different types of time based on weekdays and weekends. It can also be divided into different types of time based on one-hour units, and so on. Channels can be any one of a set of channels, which may include different types of channels. For example, based on the display platform, it may include one or more of instant messaging platforms (such as SMS, email, etc.), video playback platforms, and social interaction platforms. It should be noted that the auxiliary information in this embodiment is not limited to one or more of time and channels. As the business expands, the auxiliary information may also include other types, such as geographical location. No limitation is made.

[0051] In one embodiment, the data processing device can process object clusters, content clusters, time clusters, and channel clusters based on a combination strategy to generate multiple data combinations to be processed. Specifically, it can acquire object clusters and content clusters, determine any object from at least one object included in the object cluster, determine any content from at least one content included in the content cluster, determine any auxiliary information based on one or two of the time set and the channel set, wherein the auxiliary information includes any time in the time set and one or two of the channels in the channel set, and generate a corresponding data combination to be processed based on any object, any content, and any auxiliary information.

[0052] Assume that the object cluster includes at least one object, object 1, and object 2; the content cluster includes at least one content, content 1, and content 2; the time set includes time 1 and time 2; and the channel set includes channel 1 and channel 2. When determining auxiliary information based on the time set, any object can be determined from object 1 and object 2 in the object cluster, any content can be determined from content 1 and content 2 in the content cluster, and any time can be determined from time 1 and time 2 in the time set. A corresponding combination of data to be processed can then be generated based on any object, any content, and any time. In this case, the number of multiple combinations of data to be processed equals the number of objects in the object cluster * the number of content in the content cluster * the number of times in the time cluster. For example, if the object cluster includes Z objects (users), each object is represented as Z... i Then Z objects can be represented as: object = {object0, object1, ..., object...} Z-1 The content cluster consists of C content items, each represented as Ccontent. i Then the C contents can be represented as: content = {content0, content1, ..., content...}C-1 The time cluster contains T times, each represented as T. i Then, T time intervals can be represented as: Time = {T0, T1, ... T} T-1 The number of multiple data combinations to be processed = Z * C * T. Continuing with the example above, the number of multiple data combinations to be processed = 2 * 2 * 2 = 8. These 8 data combinations to be processed include: {object1, content1, time1}, {object2, content1, time1}, {object1, content2, time1}, {object1, content1, time2}, {object2, content2, time1}, {object2, content1, time2}, {object1, content2, time2}, {object2, content2, time2}.

[0053] When determining auxiliary information based on a channel set, any object can be determined from Object 1 and Object 2 included in the object cluster, any content can be determined from Content 1 and Content 2 included in the content cluster, and any channel can be determined from Channel 1 and Channel 2 included in the channel set. A corresponding combination of data to be processed is then generated based on this object, content, and channel. The number of multiple combinations of data to be processed equals the number of objects in the object cluster * the number of contents in the content cluster * the number of channels in the channel cluster. For example, if the object cluster includes Z objects (users), each object is represented as Z... i Then Z objects can be represented as: object = {object0, object1, ..., object...} Z-1 The content cluster consists of C content items, each represented as Ccontent. i Then the C contents can be represented as: content = {content0, content1, ..., content...} C-1 The channel cluster consists of V channels, each represented by CH. i Then V channels can be represented as: Channel = {CH0, CH1, ... CH2} V-1 The number of multiple data combinations to be processed = Z * C * V. Continuing the example above, the number of multiple data combinations to be processed = 2 * 2 * 2 = 8. These 8 data combinations to be processed can include: {Object1, Content1, Channel1}, {Object2, Content1, Channel1}, {Object1, Content2, Channel1}, {Object1, Content1, Channel2}, {Object2, Content2, Channel1}, {Object2, Content1, Channel2}, {Object1, Content2, Channel2}, {Object2, Content2, Channel2}.

[0054] When determining auxiliary information based on time and channel sets, any object can be determined from object 1 and object 2 in the object cluster, any content can be determined from content 1 and content 2 in the content cluster, any time can be determined from time 1 and time 2 in the time set, and any channel can be determined from channel 1 and channel 2 in the channel set. A corresponding combination of data to be processed is then generated based on any object, any content, any time, and any channel. The number of multiple combinations of data to be processed = the number of objects in the object cluster * the number of contents in the content cluster * the number of times in the time cluster * the number of channels in the channel cluster. For example, if the object cluster includes Z objects (users), each object is represented as Z... i Then Z objects can be represented as: object = {object0, object1, ..., object...} Z-1 The content cluster consists of C content items, each represented as Ccontent. i Then the C contents can be represented as: content = {content0, content1, ..., content...} C-1 The time cluster contains T times, each represented as T. i Then, T time intervals can be represented as: Time = {T0, T1, ... T} T-1 The channel cluster consists of V channels, each represented by CH. i Then V channels can be represented as: Channel = {CH0, CH1, ... CH2} V-1 The number of multiple data combinations to be processed = Z * C * T * V. Continuing the example above, the number of multiple data combinations to be processed = 2 * 2 * 2 * 2 = 16. These 16 data combinations can include: {Object1, Content1, Time1, Channel1}, {Object2, Content1, Time1, Channel1}, {Object1, Content2, Time1, Channel1}, {Object1, Content1, Time2, Channel1}, {Object2, Content2, Time1, Channel1}, {Object2, Content1, Time2, Channel1}, {Object1, Content2, Time2, Channel2}, {Object1, Content2, Time2, Channel2}. {Object 1, Content 2, Time 2, Channel 1}, {Object 1, Content 1, Time 1, Channel 2}, {Object 2, Content 1, Time 1, Channel 2}, {Object 1, Content 2, Time 1, Channel 2}, {Object 1, Content 1, Time 2, Channel 2}, {Object 2, Content 2, Time 1, Channel 2}, {Object 2, Content 1, Time 2, Channel 2}, {Object 1, Content 2, Time 2, Channel 2}, {Object 2, Content 2, Time 2, Channel 2}.

[0055] S202: Obtain the feature representation information of each data combination to be processed from multiple data combinations to be processed.

[0056] In one embodiment, the data processing device can acquire a first object, a first content, and first auxiliary information included in each of a plurality of data combinations to be processed, acquire feature representation information of the first object, feature representation information of the first content, and feature representation information of the first auxiliary information, and generate feature representation information for each data combination to be processed based on the feature representation information of the first object, the feature representation information of the first content, and the feature representation information of the first auxiliary information.

[0057] The feature representation information may include, but is not limited to, one or more of the following: feature expression, feature vector, and feature matrix.

[0058] In this context, each object can correspond to one object feature. An object's object feature can include one or more of the following: one-dimensional object features and multi-dimensional object features. For example, multi-dimensional object features can include object features across multiple business dimensions, such as object features related to the automotive business, the beauty business, and the savings business, etc. For the first object, in one embodiment, when the object feature of the first object is a one-dimensional object feature, a vector transformation can be directly performed on the one-dimensional object feature to obtain the feature representation information of the first object. In another embodiment, optionally, when the object feature of the first object is a multi-dimensional object feature, vector transformation can be performed on each dimension of the multi-dimensional object feature of the first object to obtain the feature representation information of each dimension, and the feature representation information of the first object can be generated based on the feature representation information of each dimension. For example, the feature representation information of each dimension can be concatenated to generate the feature representation information of the first object. In this case, the length of the feature representation information of the first object can be the sum of the lengths of the feature representation information of each dimension. That is,

[0059]

[0060] Where n is the number of dimensions of the multidimensional object's features, and n is a positive integer, X i Let be the length of the feature representation information for the i-th dimension, where i is a positive integer. It should be noted that the lengths of the feature representation information for different dimensions can be different.

[0061] For example, the data processing device can also perform weighted processing on the feature representation information of each dimension to generate the feature representation information of the first object. Specifically, it can obtain the weights of the feature representation information of each dimension, and perform weighted processing on the feature representation information of each dimension based on the weights of the feature representation information of each dimension to obtain the feature representation information of the first object.

[0062] The data processing device can use an encoding method to perform vector transformation on the object features to obtain the corresponding feature representation information. This application does not limit the encoding method, which can be one-hot encoding, embedded encoding, hard encoding (Label Encoding), target variable encoding, etc.

[0063] Each piece of content can correspond to a single content feature. Optionally, the data processing device can employ an encoding model from the field of Natural Language Processing (NLP) to perform vector transformation on each content feature, thereby obtaining the feature representation information of the content. That is, the data processing device can use an encoding model from the NLP field to perform vector transformation on the content features of the first content, thereby obtaining the feature representation information of the first content. This encoding model includes, but is not limited to, one or more of the following: keyword matching model, Word2vec model, or pre-trained BERT model.

[0064] The keyword matching model may include, but is not limited to: a keyword matching model trained based on the TF-IDF keyword extraction algorithm, a keyword matching model trained based on the TextRank keyword extraction algorithm, and a keyword matching model trained based on the LDA topic model keyword extraction algorithm, etc.

[0065] The Word2Vec model is a natural language processing model that vectorizes vocabulary; its full name is Word to Vector. A key feature of the Word2Vec model is its ability to vectorize all words, allowing for quantitative measurement of relationships between words and uncovering connections between them. The trained Word2Vec model is stored as a Word2Vec model dictionary. Based on this, in one embodiment, content features can be segmented into multiple words; the word vectors corresponding to each segment can be obtained from the Word2Vec model dictionary, and these word vectors can be concatenated to obtain the feature representation information of the content.

[0066] Among them, the BERT model is a novel language model. BERT stands for Bidirectional Encoder Representations from Transformers, and it pre-trains deep representations (Embeddings) by jointly adjusting the Transformer. The BERT model uses the Encoder part of the Transformer to learn the information of each word segment in the content features, thereby obtaining the feature representation information of the content.

[0067] In this system, each piece of auxiliary information corresponds to one auxiliary feature. Optionally, when the auxiliary information includes time, the auxiliary feature is determined by the time feature, and correspondingly, the feature representation information of the auxiliary information is determined by the time feature representation information. The data processing device can use an encoding method to perform vector transformation on the time feature to obtain the time feature representation information, and use this time feature representation information as the feature representation information of the auxiliary information. This encoding method can be One-Hot encoding, embedding encoding, Label Encoding, Target Encoding, etc. For example, when using One-Hot encoding for vector transformation, One-Hot encoding can use N bits to encode N states, each state having an independent register bit, and at any given time, only one bit is valid. When time is divided into one-hour units, the time feature representation information is 24 bits. If the time feature is 20, then only the 20th bit of the 24 register bits is valid, and the time feature representation information can be represented as "000000000000000000010000".

[0068] Optionally, when the auxiliary information includes a channel, the auxiliary features are determined by the channel features, and correspondingly, the feature representation information of the auxiliary information is determined by the feature representation information of the channel. The data processing device can use an encoding method to perform vector transformation on the channel features to obtain the feature representation information of the channel, and use this feature representation information as the feature representation information of the auxiliary information. Similar to the vector transformation of time features, One-Hot encoding can also be used to perform vector transformation on the channel features. When the channel includes instant messaging platforms, video playback platforms, and social interaction platforms, the length of the channel's feature representation information is 3 bits. If the channel feature is a video playback platform, then the feature representation information of that channel can be represented as "010".

[0069] Optionally, when the auxiliary information includes time and channel, the auxiliary features are determined by the time features and channel features. Correspondingly, the feature representation information of the auxiliary information is determined by the time feature representation information and the channel feature representation information. In one embodiment, when the data processing device obtains the time feature representation information and the channel feature representation information, it can concatenate the time feature representation information and the channel feature representation information to obtain the feature representation information of the auxiliary information. For example, when the time feature representation information is represented as "000000000000000000010000" and the channel feature representation information is represented as "010", the concatenated feature representation information of the auxiliary information can be represented as "0000000000000000000010000010". In another embodiment, when the data processing device obtains the time feature representation information and the channel feature representation information, it can perform weighted processing on the time feature representation information and the channel feature representation information to generate the feature representation information of the auxiliary information. Specifically, the weights of the feature representation information of time and the feature representation information of channels can be obtained, and the feature representation information of time and channels can be weighted based on the weights of the feature representation information of time and the feature representation information of channels to generate feature representation information of auxiliary information.

[0070] Therefore, the data processing device can obtain the feature representation information of the first auxiliary information based on one or more of the time features of the first time and the channel features of the first channel in the first auxiliary information.

[0071] Furthermore, when the feature representation information of the first object, the feature representation information of the first content, and the feature representation information of the first auxiliary information are obtained, feature representation information for each data combination to be processed can be generated based on these features. Optionally, the data processing device can concatenate the feature representation information of the first object, the feature representation information of the first content, and the feature representation information of the first auxiliary information to obtain the feature representation information for each data combination to be processed. Optionally, the data processing device can also perform weighted processing on the feature representation information of the first object, the feature representation information of the first content, and the feature representation information of the first auxiliary information to obtain the feature representation information for each data combination to be processed. Specifically, the weights of the feature representation information of the first object, the feature representation information of the first content, and the feature representation information of the first auxiliary information can be obtained separately, and weighted processing can be performed on the feature representation information of the first object, the feature representation information of the first content, and the feature representation information of the first auxiliary information based on these weights to obtain the feature representation information for each data combination to be processed.

[0072] It should be noted that, in the embodiments of this application, when splicing feature representation information, the splicing order of feature representation information can be set according to experience or business needs, or it can be random. This application does not limit this.

[0073] S203: Call the data matching model to process the feature representation information of each data combination to be processed, and obtain the target data combination of each object in at least one object, the target data combination including target auxiliary information and target content.

[0074] In one embodiment, the data processing device may invoke a data matching model to process the feature representation information of each data combination to be processed, obtain a matching score for each data combination to be processed, determine each data combination to be processed associated with each of at least one object from a plurality of data combinations to be processed, and determine the target data combination for each object from each data combination to be processed associated with each object based on the matching score of each data combination to be processed.

[0075] Each object's target data combination includes target content and target auxiliary information. This auxiliary information includes one or more of the following: target time and target channel. Specifically, when the auxiliary information includes time, the target data combination includes target time. When the auxiliary information includes channel, the target data combination includes target channel. When the auxiliary information includes both time and channel, the target data combination includes both target time and target channel.

[0076] Continuing from the previous example, when multiple data combinations to be processed include: {object1, content1, time1}, {object2, content1, time1}, {object1, content2, time1}, {object1, content1, time2}, {object2, content2, time1}, {object2, content1, time2}, {object1, content2, time2}, {object2, content2, time2}, the data matching model can be invoked to process the feature representation information of each data combination to obtain the matching score of each data combination, i.e., the matching score of the data combination "{object1, content1, time1}". The matching scores for the data combination "{object2, content1, time1}", the matching scores for the data combination "{object1, content2, time1}", the matching scores for the data combination "{object1, content1, time2}", the matching scores for the data combination "{object1, content1, time2}", the matching scores for the data combination "{object2, content2, time1}", the matching scores for the data combination "{object2, content1, time2}", the matching scores for the data combination "{object1, content2, time2}", and the matching scores for the data combination "{object2, content2, time2}". And from multiple pairs of data to be processed, we determine the data pairs related to object 1 (i.e., {object 1, content 1, time 1}, {object 1, content 2, time 1}, {object 1, content 1, time 2}, {object 1, content 2, time 2}) and the data pairs related to object 2 (i.e., {object 2, content 1, time 1}, {object 2, content 2, time 1}, {object 2, content 1, time 2}, {object 2, content 2, time 2}). Based on the matching scores of the data combination to be processed "{object1, content1, time1}", the matching scores of the data combination to be processed "{object2, content1, time1}", the matching scores of the data combination to be processed "{object1, content2, time1}", the matching scores of the data combination to be processed "{object1, content1, time2}", the matching scores of the data combination to be processed "{object2, content1, time2}", the matching scores of the data combination to be processed "{object1, content2, time2}", the matching scores of the data combination to be processed "{object1, content2, time2}", the matching scores of the data combination to be processed "{object1, content2, time2}", the matching scores of the data combination to be processed "{object1, content2, time2}", the matching scores of the data combination to be processed "{object2, content1 ... The matching score of the data combination “{object2, content2, time2}” determines the target data combination of object1 from the data combinations to be processed related to object1 (i.e., {object1, content1, time1}, {object1, content2, time1}, {object1, content1, time2}, {object1, content2, time2}), and determines the target data combination of object2 from the data combinations to be processed related to object2 (i.e., {object2, content1, time1}, {object2, content2, time1}, {object2, content1, time2}, {object2, content2, time2}).

[0077] Furthermore, determining the target data combination for each object from the various data combinations related to each object based on the matching scores of each data combination to be processed includes: for an object, finding the matching scores of various data combinations related to this object from the matching scores of various data combinations to be processed, sorting the various data combinations related to this object in descending order based on the matching scores of each data combination to be processed, and determining a preset number of data combinations to be processed at the top of the sort as the target data combination for this object; wherein, the preset number is an integer greater than or equal to 1.

[0078] The data matching model can be trained based on machine learning algorithms. These machine learning algorithms include, but are not limited to, one or more of the following: Decision Tree (DT), Rocchio, Extreme Gradient Boosting (XGBooste), Naive Bayes (NB), Linear Discriminant Analysis (LDA), Support Vector Machine (SVM), Random Forest (RF), Logistic Regression (LR), Multi-Task Learning (MTL), and LightGBM.

[0079] In one embodiment, when the data matching model is trained based on the LightGBM and LR algorithms, the data matching model may include LightGBM units and LR units. Each LightGBM unit comprises multiple subtrees. The data processing device can input the feature representation information of each data combination to be processed into each subtree of the LightGBM unit to obtain the matching score of each subtree. Based on the matching scores of each subtree, the device obtains the matching score of the LightGBM unit. Then, the device uses the LR unit to perform logistic regression processing on the matching scores of the LightGBM unit to obtain the matching score for each data combination to be processed.

[0080] Optionally, obtaining the matching score of a LightGBM unit based on the matching score of each subtree includes performing addition on the matching scores of each subtree to obtain the matching score of the LightGBM unit. That is,

[0081]

[0082] Where K is the number of subtrees, K is a positive integer, and Y jLet be the matching score of the j-th subtree, where j is a positive integer.

[0083] Please see Figure 3 , Figure 3 A flowchart of a data processing method is shown. For each data combination to be processed, including a first object, first content, and first auxiliary information: 301. Obtain object feature 1, object feature 2...object feature M corresponding to the first object, content feature 1 corresponding to the first content, and time feature 1 and channel feature 1 corresponding to the first auxiliary information. 302. Obtain the feature representation information L of the first object based on object feature 1, object feature 2...object feature M corresponding to the first object. 303. Obtain the feature representation information P of the first content based on content feature 1 corresponding to the first content. 304. Obtain the feature representation information Q of the first auxiliary information based on time feature 1 and channel feature 1 corresponding to the first auxiliary information. 305. Concatenate the feature representation information L of the first object, the feature representation information P of the first content, and the feature representation information Q of the first auxiliary information to obtain the feature representation information of each data combination to be processed. 306. Call each subtree of the LightGBM unit to process the feature representation information of each data combination to be processed, obtain the matching score of each subtree, and obtain the matching score of the LightGBM unit based on the matching score of each subtree. 307. Call the LR unit to perform logistic regression processing on the matching scores of the LightGBM unit, and output the matching score of each data combination to be processed.

[0084] S204: Send target content to each object based on target auxiliary information.

[0085] In one embodiment, when the target auxiliary information includes a target time, the data processing device sends the target content to the terminal device associated with each object when the target time arrives. When the target auxiliary information includes a target channel, the data processing device sends the target content to the terminal device associated with each object through the target channel. When the target auxiliary information includes both target time and target channel, the data processing device sends the target content to the terminal device associated with each object through the target channel when the target time arrives.

[0086] Furthermore, when sending target content to each object based on target auxiliary information, sending conditions can be set. When the target auxiliary information includes a target time, it is determined whether the matching score of the target data combination is greater than or equal to a preset score threshold. If so, the target content is sent to the terminal device associated with each object when the target time arrives. When the target auxiliary information includes a target channel, it is determined whether the matching score of the target data combination is greater than or equal to a preset score threshold. If so, the target content is sent to the terminal device associated with each object through the target channel. When the target auxiliary information includes both target time and target channel, it is determined whether the matching score of the target data combination is greater than or equal to a preset score threshold. If so, the target content is sent to the terminal device associated with each object through the target channel when the target time arrives.

[0087] Please see Figure 4 As shown, Figure 4 A schematic diagram of the architecture of a data processing method is shown. (For example...) Figure 4 As shown, the architecture diagram of this data processing method includes an overall framework diagram and a diagram of the method implementation process. Figure 4 As shown in the left figure, the overall framework diagram includes an information acquisition unit 401, a data management unit 402, a data processing unit 403, and a sending unit 404.

[0088] The information collection unit 401 is used to collect various information, including but not limited to data information of registered user accounts.

[0089] The data management unit 402 may include an object cluster bar, a content cluster bar, a time cluster bar, a channel cluster bar, and a reference bar. The object cluster bar indicates at least one object included in the object cluster. The content cluster bar indicates at least one piece of content included in the content cluster. The time cluster bar indicates at least one time included in the time cluster. The channel cluster bar indicates at least one channel included in the channel cluster. The reference bar indicates reference labels for the data combinations to be processed.

[0090] Among them, such as Figure 4As shown in the right-hand figure, the data processing unit 403 can execute the data processing method of this application embodiment. Specifically, 4031a: Tag delineation is performed based on targeted object tags to determine at least one object included in the object cluster. 4031b: At least one object included in the object cluster is determined based on the potential customer mining model, the MMOE model, and the LightGBM model. 4031a and 4031b are parallel steps; only 4031a can be executed, only 4031b can be executed, or both can be executed simultaneously to determine at least one object included in the object cluster. 4032: Multiple data combinations to be processed are determined based on the object cluster including at least one object, the content cluster including at least one piece of content, the time set including at least one time, and the channel set including at least one channel. 4033: Feature representation information of each data combination to be processed is obtained from the multiple data combinations to be processed. 4034: The data matching model is invoked to process the feature representation information of each data combination to be processed. The data matching model includes, but is not limited to, one or more of the following: classification models (such as DT models trained based on the DT algorithm) and MTL models (the MTL models are trained by the MTL algorithm). 4035. Output the matching score for each data combination to be processed. 4036. Determine the individual data combinations to be processed that are associated with each object in at least one object from the multiple data combinations to be processed, and determine the target data combination for each object from the individual data combinations to be processed associated with each object based on the matching scores of each data combination to be processed.

[0091] The sending unit 404 can be used to send target content to each object through the target channel when the target time in the target data combination arrives.

[0092] In this embodiment, the data processing device can determine multiple data combinations to be processed based on at least one object in an object cluster, at least one piece of content in a content cluster, and at least one time and at least one channel in auxiliary information. It then calls a data matching model to process the feature representation information of each data combination to obtain a matching score for each data combination. Based on the matching score of each data combination, a target data combination for each object can be determined, allowing target content to be sent to each object based on the target auxiliary information included in the target data combination. Since the data combinations to be processed include auxiliary information, the target data combinations also correspondingly include target auxiliary information. Unlike existing schemes that randomly send target content to each object, this embodiment's data processing device effectively improves the accuracy of intelligent data delivery by sending target content to each object based on the target auxiliary information included in the target data combination, thereby effectively increasing click-through rates or conversion rates. Furthermore, this embodiment determines the target data combination for each object based on the matching score of each data combination to be processed, achieving human-level data processing. Unlike existing schemes that send the same content to each object in an object group, this allows for the determination of different target content for different objects, fulfilling diverse and personalized needs.

[0093] See above Figure 2 As can be seen from the relevant description of the method embodiments shown, Figure 2 The data processing method shown can invoke a data matching model to process the feature representation information of each data combination to be processed, thereby obtaining the target data combination for each object. See also Figure 5 As shown in the figure, this application provides a schematic diagram of the training process for a data matching model. Figure 5 As shown, the training process diagram of this data matching model includes S501-S504:

[0094] S501: Obtain training samples, which include a combination of sample data and a reference label for the combination of sample data; the combination of sample data includes sample objects, sample content and corresponding auxiliary information; the reference label is used to indicate the response data of the sample objects to the sample content.

[0095] In one embodiment, the data processing device can acquire historical interaction data of the sample object with respect to the sample content, generate corresponding auxiliary information based on the historical interaction data, generate a sample data combination based on the sample object, the sample content and the corresponding auxiliary information, generate a reference label for the sample data combination based on the historical interaction data, and use the sample data combination and the reference label for the sample data combination as training samples.

[0096] Optionally, the historical interaction data may include the time and channel through which sample content was sent to the sample object, as well as the sample object's response data to the sample content. The data processing device can then determine corresponding auxiliary information based on one or both of the time and channel, generate a sample data combination based on the sample object, sample content, and corresponding auxiliary information, and generate reference labels for the sample data combination based on the response data. The sample data combination and its reference labels are then used as training samples.

[0097] For example, suppose historical interaction data includes the time (C) and channel (D) when sample object A sends sample content B, and the response data of sample object A to sample content B. When the data processing device determines the corresponding auxiliary information based on time, it can generate a sample data combination {sample object A, sample content B, time C} based on sample object A, sample content B, and time C, and generate reference labels for the sample data combination {sample object A, sample content B, time C} based on the response data. The sample data combination {sample object A, sample content B, time C} and its reference labels are used as training samples. When the data processing device determines the corresponding auxiliary information based on channel, it can generate a sample data combination {sample object A, sample content B, channel D} based on sample object A, sample content B, and time C, and generate reference labels for the sample data combination {sample object A, sample content B, channel D} based on the response data. The sample data combination {sample object A, sample content B, channel D} and its reference labels are used as training samples. When the data processing device determines the corresponding auxiliary information based on time and channel, it can generate a sample data combination {sample object A, sample content B, time C, channel D} based on sample object A, sample content B, time C, and channel D. It can also generate reference labels for the sample data combination {sample object A, sample content B, time C, channel D} based on the response data, and use the sample data combination {sample object A, sample content B, time C, channel D} and its reference labels as training samples.

[0098] The response data can include click data or conversion data. Click data indicates whether the intelligently sent content has been clicked. Conversion data indicates whether the intelligently sent content has been converted, i.e., whether further action was taken after the click, such as a purchase or registration. The response data can be obtained through... Figure 4 The reference bar indicates this.

[0099] Optionally, when the response data is click data, the click data includes two types: clicked and not clicked. If the click data of sample object A for sample content B is clicked, the click probability of the sample data combination can be determined to be 1, and the reference label of the sample data combination is 1. If the click data of sample object A for sample content B is not clicked, the click probability of the sample data combination can be determined to be 0, and the reference label of the sample data combination is 0. Optionally, when the response data is conversion data, the conversion data includes two types: converted and not converted. If the conversion data of sample object A for sample content B is converted, the conversion probability of the sample data combination can be determined to be 1, and the reference label of the sample data combination is 1. If the conversion data of sample object A for sample content B is not converted, the conversion probability of the sample data combination can be determined to be 0, and the reference label of the sample data combination is 0.

[0100] S502: The initial neural network is used to process the sample data combination to obtain the predicted matching score of the sample data combination.

[0101] Specifically, the data processing device can acquire the feature representation information of the sample data combination and call the initial neural network to process the feature representation information of the sample data combination to obtain the predicted matching score of the sample data combination. This predicted matching score can also be referred to as the predicted click probability or the predicted conversion probability.

[0102] Among them, the specific implementation method of data processing equipment acquiring feature representation information of sample data combinations and Figure 2 The implementation method of S202 to obtain the feature representation information of each data combination to be processed in multiple data combinations to be processed is similar, and will not be described in detail here.

[0103] S503: Determine the loss value based on the predicted matching score of the sample data combination and the reference label.

[0104] Specifically, the data processing device can obtain the loss function and substitute the predicted matching score of the sample data combination and the reference label into the loss function to obtain the corresponding loss value.

[0105] Optionally, the loss function in this embodiment can be the cross-entropy loss function. The expression for this cross-entropy loss function is: -(y*log(y p )+(y-1)*(log(1-y p )))

[0106] Where y is the reference label for the sample data combination, y p The predicted matching score is the combination of sample data.

[0107] S504: Adjust the parameters of the initial neural network based on the loss value to obtain the data matching model.

[0108] In one embodiment, the parameters of the initial neural network can be adjusted based on the loss value until the loss value is less than a preset value or the number of parameter adjustments reaches a preset number, at which point training is terminated and a data matching model is obtained.

[0109] like Figure 6 As shown, Figure 6 A schematic diagram of another data processing system architecture is shown. This data processing system may include an offline modeling unit and an online prediction unit. The offline modeling unit is used for: 601. Obtaining sample objects, sample content, and corresponding auxiliary information (which includes one or two of time and channel) from historical interaction data. 602. Generating sample data combinations based on sample objects, sample content, and corresponding auxiliary information. 603. Obtaining feature representation information of sample objects, sample content, and auxiliary information in the sample data combinations. 604. Generating feature representation information for each data combination to be processed based on the feature representation information of sample objects, sample content, and auxiliary information. 605. Processing the feature representation information of the sample data combinations using an initial neural network to obtain a predicted matching score for the sample data combinations, determining a loss value based on the predicted matching score and reference labels, and adjusting the parameters of the initial neural network based on the loss value to obtain a data matching model. The online prediction unit is used for: 606. Obtaining object clusters, content clusters, and auxiliary information. The object cluster includes at least one object, the content cluster includes at least one piece of content, and the auxiliary information includes one or two of time and channel. 607. Based on a combination strategy, process object clusters, content clusters, and auxiliary information to determine multiple data combinations to be processed. 608. Obtain the first object, first content, and first auxiliary information included in each data combination to be processed, and obtain the feature representation information of the first object, the first content, and the first auxiliary information. 609. Based on the feature representation information of the first object, the first content, and the first auxiliary information, generate feature representation information for each data combination to be processed. 610. Call a data matching model to process the feature representation information of each data combination to be processed, and obtain a matching score for each data combination to be processed. 611. Determine each data combination to be processed that is related to each object in at least one object from the multiple data combinations to be processed, and determine the target data combination for each object from the data combinations to be processed related to each object based on the matching scores of each data combination to be processed.

[0110] In this embodiment, the data processing device can process historical interaction data to obtain sample data combinations and reference labels for these combinations. By training the sample data combinations and reference labels using machine learning algorithms, a data matching model can be obtained. When the accuracy of the data matching model is high enough, it can determine the target data combination for each object in the object cluster without human intervention. Based on the target auxiliary information of the target data combination, it can send the target content of the target data combination to each object, thus achieving intelligent data transmission.

[0111] Based on the description of the above data processing method embodiments, this application also discloses a data processing apparatus. The data processing apparatus can be a computer program (including program code) running on the aforementioned data processing device. This data processing apparatus can execute... Figure 2 or Figure 5 The method shown.

[0112] Please see Figure 7 The data processing device can operate the following units:

[0113] The determining unit 701 is used to determine multiple combinations of data to be processed based on object clusters, content clusters, and auxiliary information. The object cluster includes at least one object, the content cluster includes at least one piece of content, and the auxiliary information includes one or two of time and channel.

[0114] The acquisition unit 702 is used to acquire feature representation information of each data combination to be processed in multiple data combinations to be processed.

[0115] The processing unit 703 is used to call the data matching model to process the feature representation information of each data combination to be processed, and to obtain the target data combination of each object in at least one object. The target data combination includes target auxiliary information and target content.

[0116] The sending unit 704 is used to send target content to each object based on target auxiliary information.

[0117] In one implementation, the determining unit 701 is used to determine multiple combinations of data to be processed based on object clusters, content clusters, and auxiliary information, including:

[0118] Retrieve object clusters and content clusters;

[0119] Determine any object from at least one object included in an object cluster, and determine any content from at least one content included in a content cluster;

[0120] Any auxiliary information is determined based on one or both of the time set and the channel set, wherein any auxiliary information includes any time in the time set and one or both of the channels in the channel set;

[0121] Generate a corresponding combination of data to be processed based on any object, any content, and any auxiliary information.

[0122] In another embodiment, the acquisition unit 702 is used to acquire feature representation information of each data combination to be processed in a plurality of data combinations to be processed, including:

[0123] Obtain the first object, first content, and first auxiliary information for each of the multiple data combinations to be processed;

[0124] Obtain feature representation information of the first object, feature representation information of the first content, and feature representation information of the first auxiliary information;

[0125] Based on the feature representation information of the first object, the feature representation information of the first content, and the feature representation information of the first auxiliary information, feature representation information for each combination of data to be processed is generated.

[0126] In another embodiment, the processing unit 703 is used to invoke a data matching model to process the feature representation information of each data combination to be processed, to obtain a target data combination for each object in at least one object, including:

[0127] The data matching model is invoked to process the feature representation information of each data combination to obtain a matching score for each data combination.

[0128] Identify, from multiple combinations of data to be processed, each combination of data to be processed that is associated with each of at least one object;

[0129] Based on the matching scores of each data combination to be processed, the target data combination for each object is determined from the data combinations to be processed related to each object.

[0130] In another embodiment, the target auxiliary information includes the target time and the target channel, and the sending unit 704 is used to send target content to each object based on the target auxiliary information, including:

[0131] Determine whether the matching score of the target data combination is greater than or equal to a preset score threshold;

[0132] If so, the target content will be sent to each object through the target channel when the target time arrives.

[0133] In another embodiment, before the processing unit 703 calls the data matching model to process the feature representation information of each data combination to be processed, and obtains the target data combination of each object in at least one object, the processing unit 703 is further configured to:

[0134] Acquire training samples, which include combinations of sample data and reference labels for those combinations. Each combination of sample data includes a sample object, sample content, and corresponding auxiliary information. The reference labels are used to indicate the response data of the sample object to the sample content.

[0135] The sample data combination is processed using an initial neural network to obtain the predicted matching score of the sample data combination;

[0136] The loss value is determined based on the predicted matching score of the sample data combination and the reference label;

[0137] The parameters of the initial neural network are adjusted based on the loss value to obtain a data matching model.

[0138] In another embodiment, the processing unit 703 is used to acquire training samples, including:

[0139] Obtain the time and channel through which sample content was sent to the sample object, as well as the sample object's response data to the sample content. The response data includes click data or conversion data.

[0140] The corresponding auxiliary information is determined based on one or two of the time and channels;

[0141] Generate sample data combinations based on sample objects, sample content, and corresponding auxiliary information;

[0142] Reference labels for sample data combinations are generated based on the response data, and the sample data combinations and their reference labels are used as training samples.

[0143] According to one embodiment of this application, Figure 2 or Figure 5 Each step involved in the method shown can be performed by... Figure 7 The data processing apparatus shown is executed by each unit. For example, Figure 2 S201 shown is composed of Figure 7 The determination unit 701 shown is used to execute S202. Figure 7 The acquisition unit 702 shown is used to execute S203. Figure 7 The processing unit 703 shown is responsible for executing S204. Figure 7 The sending unit 704 shown is used to perform this. For example, Figure 5 S501, S502, S503, and S504 are composed of Figure 7The processing unit 703 shown is used to execute this.

[0144] According to another embodiment of this application, Figure 7 The data processing apparatus shown can be composed of individual or combined units into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above-mentioned units are divided based on logical functions. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the data processing apparatus may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.

[0145] According to another embodiment of this application, processing elements and storage elements, such as a central processing unit (CPU), random access storage medium (RAM), and read-only storage medium (ROM), can be used. For example, a general-purpose computing device such as a computer can run on a device capable of performing tasks such as... Figure 2 or Figure 5 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 7 The data processing apparatus shown herein, and the data processing method for implementing the embodiments of this application, are described. A computer program may be recorded on, for example, a computer-readable recording medium, loaded onto the aforementioned data processing apparatus via the computer-readable recording medium, and executed therein.

[0146] In this embodiment, the data processing device can determine multiple data combinations to be processed based on at least one object in an object cluster, at least one content in a content cluster, and at least one time and at least one channel in auxiliary information. It then calls a data matching model to process the feature representation information of each data combination to obtain a matching score for each data combination. Based on the matching score of each data combination, a target data combination for each object can be determined, allowing target content to be sent to each object based on the target auxiliary information included in the target data combination. Since the data combinations to be processed include auxiliary information, the target data combinations also correspondingly include target auxiliary information. Unlike existing schemes that randomly send target content to each object, this embodiment's data processing device effectively improves the accuracy of intelligent data delivery by sending target content to each object based on the target auxiliary information included in the target data combination, thereby effectively increasing click-through rates or conversion rates. Furthermore, this embodiment determines the target data combination for each object based on the matching score of each data combination to be processed, achieving human-level data processing. Unlike existing schemes that send the same content to each object in an object group, this allows for the determination of different target content for different objects, fulfilling diverse and personalized needs.

[0147] Based on the description of the above data processing method embodiments, this application also discloses a data processing device. Please refer to... Figure 8 The data processing device includes at least a processor 801, an input interface 802, an output interface 803, and a computer storage medium 804, which can be connected via a bus or other means.

[0148] The computer storage medium 804 is a memory device in the data processing device, used to store programs and data. It is understood that the computer storage medium 804 can include the built-in storage medium of the data processing device, or it can include extended storage media supported by the data processing device. The computer storage medium 804 provides storage space that stores the operating system of the data processing device. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by the processor 801. These instructions can be one or more computer programs (including program code). It should be noted that the computer storage medium can be a high-speed RAM memory; optionally, it can also be at least one computer storage medium located remotely from the aforementioned processor. The processor can be called a Central Processing Unit (CPU), which is the core and control center of the data processing device, suitable for implementing one or more instructions, specifically loading and executing one or more instructions to achieve the corresponding method flow or function.

[0149] In one embodiment, processor 801 may load and execute one or more instructions stored in computer storage medium 804 to perform, for example, Figure 2 or Figure 5 In the specific implementation of the corresponding methods shown, one or more instructions in the computer storage medium 804 are loaded by the processor 801 and executed in the following steps:

[0150] Multiple data combinations to be processed are determined based on object clusters, content clusters, and auxiliary information. The object cluster includes at least one object, the content cluster includes at least one piece of content, and the auxiliary information includes one or two of time and channel.

[0151] Obtain the feature representation information of each data combination to be processed from multiple data combinations to be processed;

[0152] The data matching model is invoked to process the feature representation information of each data combination to be processed, so as to obtain the target data combination of each object in at least one object. The target data combination includes target auxiliary information and target content.

[0153] Send target content to each object based on target auxiliary information.

[0154] In one implementation, the processor 801 is configured to determine multiple combinations of data to be processed based on object clusters, content clusters, and auxiliary information, including:

[0155] Retrieve object clusters and content clusters;

[0156] Determine any object from at least one object included in an object cluster, and determine any content from at least one content included in a content cluster;

[0157] Any auxiliary information is determined based on one or both of the time set and the channel set, wherein any auxiliary information includes any time in the time set and one or both of the channels in the channel set;

[0158] Generate a corresponding combination of data to be processed based on any object, any content, and any auxiliary information.

[0159] In another embodiment, the processor 801 is used to acquire feature representation information for each of the multiple data combinations to be processed, including:

[0160] Obtain the first object, first content, and first auxiliary information for each of the multiple data combinations to be processed;

[0161] Obtain feature representation information of the first object, feature representation information of the first content, and feature representation information of the first auxiliary information;

[0162] Based on the feature representation information of the first object, the feature representation information of the first content, and the feature representation information of the first auxiliary information, feature representation information for each combination of data to be processed is generated.

[0163] In another embodiment, the processor 801 is used to invoke a data matching model to process the feature representation information of each data combination to be processed, to obtain a target data combination for each object in at least one object, including:

[0164] The data matching model is invoked to process the feature representation information of each data combination to obtain a matching score for each data combination.

[0165] Identify, from multiple combinations of data to be processed, each combination of data to be processed that is associated with each of at least one object;

[0166] Based on the matching scores of each data combination to be processed, the target data combination for each object is determined from the data combinations to be processed related to each object.

[0167] In another embodiment, the target auxiliary information includes target time and target channel, and the processor 801 is used to send target content to each object based on the target auxiliary information, including:

[0168] Determine whether the matching score of the target data combination is greater than or equal to a preset score threshold;

[0169] If so, the target content will be sent to each object through the target channel when the target time arrives.

[0170] In another embodiment, before the processor 801 calls the data matching model to process the feature representation information of each data combination to be processed, and obtains the target data combination of each object in at least one object, the processor 801 is further configured to:

[0171] Acquire training samples, which include combinations of sample data and reference labels for those combinations. Each combination of sample data includes a sample object, sample content, and corresponding auxiliary information. The reference labels are used to indicate the response data of the sample object to the sample content.

[0172] The sample data combination is processed using an initial neural network to obtain the predicted matching score of the sample data combination;

[0173] The loss value is determined based on the predicted matching score of the sample data combination and the reference label;

[0174] The parameters of the initial neural network are adjusted based on the loss value to obtain a data matching model.

[0175] In another embodiment, the processor 801 is used to acquire training samples, including:

[0176] Obtain the time and channel through which sample content was sent to the sample object, as well as the sample object's response data to the sample content. The response data includes click data or conversion data.

[0177] The corresponding auxiliary information is determined based on one or two of the time and channels;

[0178] Generate sample data combinations based on sample objects, sample content, and corresponding auxiliary information;

[0179] Reference labels for sample data combinations are generated based on the response data, and the sample data combinations and their reference labels are used as training samples.

[0180] In this embodiment, the data processing device can determine multiple data combinations to be processed based on at least one object in an object cluster, at least one piece of content in a content cluster, and at least one time and at least one channel in auxiliary information. It then calls a data matching model to process the feature representation information of each data combination to obtain a matching score for each data combination. Based on the matching score of each data combination, a target data combination for each object can be determined, allowing target content to be sent to each object based on the target auxiliary information included in the target data combination. Since the data combinations to be processed include auxiliary information, the target data combinations also correspondingly include target auxiliary information. Unlike existing schemes that randomly send target content to each object, this embodiment's data processing device effectively improves the accuracy of intelligent data delivery by sending target content to each object based on the target auxiliary information included in the target data combination, thereby effectively increasing click-through rates or conversion rates. Furthermore, this embodiment determines the target data combination for each object based on the matching score of each data combination to be processed, achieving human-level data processing. Unlike existing schemes that send the same content to each object in an object group, this allows for the determination of different target content for different objects, fulfilling diverse and personalized needs.

[0181] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the data processing method embodiment described above. Figure 2 or Figure 5 The steps performed in the process.

[0182] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions can perform the data processing method embodiments described above. Figure 2 or Figure 5 The steps performed in the process.

[0183] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0184] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Those skilled in the art will understand that all or part of the processes for implementing the above embodiments and equivalent variations made in accordance with the claims of this application are still within the scope of this application.

Claims

1. A data processing method, characterized in that, The method includes: Multiple data combinations to be processed are determined based on object clusters, content clusters, and auxiliary information. The object cluster includes at least one object, the content cluster includes at least one piece of content, and the auxiliary information includes time and channel. The channel is used to indicate the way the content is sent. When classified according to the display platform, the channel includes one or more of instant messaging platforms, video playback platforms, and social interaction platforms. Obtain feature representation information for each of the plurality of data combinations to be processed, wherein each data combination to be processed includes auxiliary information, an object in the object cluster, and a piece of content in the content cluster; The data matching model is invoked to process the feature representation information of each data combination to be processed, so as to obtain the target data combination of each object in the at least one object. The target data combination includes target auxiliary information and target content. The target auxiliary information includes target time and target channel. When the target time arrives, the target content is sent to each of the objects through the target channel.

2. The method as described in claim 1, characterized in that, The process of determining multiple combinations of data to be processed based on object clusters, content clusters, and auxiliary information includes: Retrieve object clusters and content clusters; Determine any object from at least one object included in the object cluster, and determine any content from at least one content included in the content cluster; Any auxiliary information is determined based on one or both of the time set and the channel set, wherein the auxiliary information includes any time in the time set and one or both of the channels in the channel set; Generate a corresponding combination of data to be processed based on any of the objects, any of the content, and any of the auxiliary information.

3. The method as described in claim 1 or 2, characterized in that, The step of obtaining the feature representation information of each of the plurality of data combinations to be processed includes: Obtain the first object, first content, and first auxiliary information included in each of the plurality of data combinations to be processed; Obtain feature representation information of the first object, feature representation information of the first content, and feature representation information of the first auxiliary information; Based on the feature representation information of the first object, the feature representation information of the first content, and the feature representation information of the first auxiliary information, feature representation information of each data combination to be processed is generated.

4. The method as described in claim 1 or 2, characterized in that, The invoked data matching model processes the feature representation information of each data combination to be processed, obtaining the target data combination for each object in the at least one object, including: The data matching model is invoked to process the feature representation information of each data combination to be processed, and a matching score is obtained for each data combination to be processed. From the plurality of data combinations to be processed, determine each data combination to be processed that is associated with each of the at least one object; Based on the matching scores of each combination of data to be processed, the target data combination for each object is determined from the combinations of data to be processed associated with each object.

5. The method as described in claim 4, characterized in that, When the target time arrives, sending the target content to each object through the target channel includes: Determine whether the matching score of the target data combination is greater than or equal to a preset score threshold; If so, when the target time arrives, the target content is sent to each object through the target channel.

6. The method as described in claim 1, characterized in that, Before the method calls the data matching model to process the feature representation information of each data combination to be processed to obtain the target data combination of each object in the at least one object, the method further includes: Acquire training samples, which include combinations of sample data and reference labels for the combinations of sample data; the combinations of sample data include sample objects, sample content, and corresponding auxiliary information; the reference labels are used to indicate the response data of the sample objects to the sample content; The sample data combination is processed using an initial neural network to obtain the predicted matching score of the sample data combination; The loss value is determined based on the predicted matching score of the sample data combination and the reference label; The parameters of the initial neural network are adjusted based on the loss value to obtain a data matching model.

7. The method as described in claim 6, characterized in that, The acquisition of training samples includes: The time and channel through which the sample content was sent to the sample object are obtained, as well as the response data of the sample object to the sample content, the response data including click data or conversion data; The corresponding auxiliary information is determined based on one or two of the time and the channels; A sample data combination is generated based on the sample object, the sample content, and the corresponding auxiliary information; Based on the response data, a reference label is generated for the sample data combination, and the sample data combination and the reference label of the sample data combination are used as training samples.

8. A data processing apparatus, characterized in that, The device includes: The determining unit is used to determine multiple combinations of data to be processed based on object clusters, content clusters, and auxiliary information. The object clusters include at least one object, the content clusters include at least one piece of content, and the auxiliary information includes time and channel. The channel is used to indicate the way the content is sent. When divided according to the display platform, the channel includes one or more of instant messaging platforms, video playback platforms, and social interaction platforms. The acquisition unit is used to acquire feature representation information of each data combination to be processed in the plurality of data combinations to be processed, wherein each data combination to be processed includes auxiliary information, an object in the object cluster, and a piece of content in the content cluster; The processing unit is used to call the data matching model to process the feature representation information of each data combination to be processed, and obtain the target data combination of each object in the at least one object. The target data combination includes target auxiliary information and target content. The target auxiliary information includes target time and target channel. A sending unit is used to send the target content to each object through the target channel when the target time arrives.

9. The apparatus as claimed in claim 8, characterized in that, The determining unit is used to determine multiple combinations of data to be processed based on object clusters, content clusters, and auxiliary information, including: Retrieve object clusters and content clusters; Determine any object from at least one object included in the object cluster, and determine any content from at least one content included in the content cluster; Any auxiliary information is determined based on one or both of the time set and the channel set, wherein the auxiliary information includes any time in the time set and one or both of the channels in the channel set; Generate a corresponding combination of data to be processed based on any of the objects, any of the content, and any of the auxiliary information.

10. The apparatus as claimed in claim 8 or 9, characterized in that, The acquisition unit is used to acquire feature representation information of each of the plurality of data combinations to be processed, including: Obtain the first object, first content, and first auxiliary information included in each of the plurality of data combinations to be processed; Obtain feature representation information of the first object, feature representation information of the first content, and feature representation information of the first auxiliary information; Based on the feature representation information of the first object, the feature representation information of the first content, and the feature representation information of the first auxiliary information, feature representation information of each data combination to be processed is generated.

11. The apparatus as claimed in claim 8 or 9, characterized in that, The processing unit is used to invoke a data matching model to process the feature representation information of each data combination to be processed, to obtain the target data combination of each object in the at least one object, including: The data matching model is invoked to process the feature representation information of each data combination to be processed, and a matching score is obtained for each data combination to be processed. From the plurality of data combinations to be processed, determine each data combination to be processed that is associated with each of the at least one object; Based on the matching scores of each combination of data to be processed, the target data combination for each object is determined from the combinations of data to be processed associated with each object.

12. The apparatus as claimed in claim 11, characterized in that, The sending unit is used to send the target content to each object through the target channel when the target time arrives, including: Determine whether the matching score of the target data combination is greater than or equal to a preset score threshold; If so, when the target time arrives, the target content is sent to each object through the target channel.

13. The apparatus as claimed in claim 8, characterized in that, Before the processing unit calls the data matching model to process the feature representation information of each data combination to be processed, and obtains the target data combination of each object in the at least one object, it is further configured to: Acquire training samples, which include combinations of sample data and reference labels for the combinations of sample data; the combinations of sample data include sample objects, sample content, and corresponding auxiliary information; the reference labels are used to indicate the response data of the sample objects to the sample content; The sample data combination is processed using an initial neural network to obtain the predicted matching score of the sample data combination; The loss value is determined based on the predicted matching score of the sample data combination and the reference label; The parameters of the initial neural network are adjusted based on the loss value to obtain a data matching model.

14. The apparatus as claimed in claim 13, characterized in that, The processing unit is used to acquire training samples, including: The time and channel through which the sample content was sent to the sample object are obtained, as well as the response data of the sample object to the sample content, the response data including click data or conversion data; The corresponding auxiliary information is determined based on one or two of the time and the channels; A sample data combination is generated based on the sample object, the sample content, and the corresponding auxiliary information; Based on the response data, a reference label is generated for the sample data combination, and the sample data combination and the reference label of the sample data combination are used as training samples.

15. A data processing device, comprising an input interface and an output interface, characterized in that, Also includes: A processor, adapted to implement one or more instructions; as well as, A computer storage medium storing one or more instructions, said one or more instructions being adapted to be loaded by said processor and executed as described in any one of claims 1-7.

16. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions, which are adapted to be loaded by a processor and executed by the data processing method as described in any one of claims 1-7.

17. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a processor, implement the data processing method according to any one of claims 1-7.