Data commodity recommendation method and device and computer equipment

By extracting and analyzing the characteristics of data products and users, and using the user tower and commodity tower models to determine recommended data products, the problem of low recommendation accuracy caused by using only fixed features in the prior art is solved, and a higher recommendation accuracy is achieved.

CN120146950APending Publication Date: 2025-06-13CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213856.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, since recommendations are only based on the fixed characteristics of the data elements themselves, the accuracy of data product recommendation is low.

Method used

By obtaining the original data, including the historical transaction records of the data products, the data products to be traded and the user information to be traded, the feature extraction is carried out to obtain the characteristics of the data products to be traded and the user characteristics to be traded. These features include category features, numerical features and sequence features, and are input into the user tower and commodity tower in the data product recommendation model for analysis, and obtain the user tower embedding vector and commodity tower embedding vector, and determine the recommended data products of the users to be traded based on these vectors.

Benefits of technology

By combining the fixed attribute characteristics and dynamic changes of data products and users, the accuracy of data products recommendations is improved, and the problem of low recommendation accuracy is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146950A_ABST
    Figure CN120146950A_ABST
Patent Text Reader

Abstract

The invention discloses a data commodity recommendation method and device and computer equipment. The method comprises the steps that original data are acquired, and the original data comprise historical transaction records of data commodities, to-be-transacted data commodities and to-be-transacted user information; feature extraction is carried out on the original data to obtain to-be-transacted data commodity features and to-be-transacted user features, and the to-be-transacted data commodity features and to-be-transacted user features comprise category features, numerical value features and sequence features; respectively inputting the to-be-transacted data commodity features and the to-be-transacted user features into a user tower and a commodity tower in a data commodity recommendation model for analysis to obtain a user tower embedded vector and a commodity tower embedded vector; and based on the user tower embedded vector and the commodity tower embedded vector, determining a recommended data commodity corresponding to the to-be-transacted user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology. Specifically, it relates to a data commodity recommendation method, apparatus, and computer device. Background Art

[0002] Data, as a new type of production factor, is the new quality productivity in the digital economy era. The data factor service platform can help data factors achieve circulation and at the same time empower the digital transformation of enterprises. As the most core trading object of the platform, data factors have characteristics such as high timeliness, ambiguity, and richness. This poses a major problem for how to quickly select suitable data services for oneself from the complex data factors. Currently, most solutions mainly rely on finding through one's own business experience, or quickly establishing a recommendation method through methods such as popular searches and feature matching. However, the methods in the related technologies all perform feature matching using the original fixed features of the data factors, resulting in a low accuracy rate of data commodity recommendation. Summary of the Invention

[0003] The embodiments of this application provide a data commodity recommendation method, apparatus, and computer device to at least solve the technical problem in the related technologies that the recommendation accuracy rate is low due to recommending only based on the fixed features of the data factors themselves.

[0004] According to one aspect of the embodiments of this application, a data commodity recommendation method is provided, including: obtaining original data, where the original data includes: historical transaction records of data commodities, data commodities to be traded, and information of users to be traded; performing feature extraction on the original data to obtain features of the data commodities to be traded and features of the users to be traded. Among them, the features of the data commodities to be traded and the features of the users to be traded both include: category features, numerical features, and sequence features; respectively inputting the features of the data commodities to be traded and the features of the users to be traded into the user tower and the commodity tower in the data commodity recommendation model for analysis to obtain a user tower embedding vector and a commodity tower embedding vector; determining the recommended data commodities corresponding to the users to be traded based on the user tower embedding vector and the commodity tower embedding vector.

[0005] Optionally, feature extraction is performed on the original data to obtain the features of the data item to be traded and the features of the user to be traded, including: encoding the categorical variables in the original data to obtain the categorical features, where the categorical features are used to represent the category information of the data item to be traded or the category information of the user to be traded; performing non-linear transformation on the numerical variables in the original data to obtain the transformed numerical variables, and encoding the transformed numerical variables to obtain the numerical features, where the numerical features are used to represent the numerical information of the original data; extracting the historical transaction records of the data item from the original data; and determining the sequential features according to the historical transaction records of the data item.

[0006] Optionally, determining the sequential features according to the historical transaction records of the data item includes: when the features of the user to be traded are obtained, converting the user identifier to be traded and the data item identifier in the historical transaction records of the data item into label encodings, and sorting the label encodings in chronological order to form the first sequential feature, where the first sequential feature is used to represent the change trend of the items of interest to the user to be traded; when the features of the data item to be traded are obtained, converting the user identifier to be traded and the data item identifier in the historical transaction records of the data item into label encodings, and sorting the label encodings in chronological order to form the second sequential feature, where the second sequential feature is used to represent the popularity of the data item to be traded, and the length of the sequential feature is the total number of identifiers included in the sequential feature plus one.

[0007] Optionally, the method further includes: obtaining the length of the sequential feature; when the length of the sequential feature is greater than a preset length, truncating the sequential feature to retain only the sequential feature of the preset length; when the length of the sequential feature is less than the preset length, padding zeros at the front of the sequential feature until the length of the sequential feature is equal to the preset length.

[0008] Optionally, determining the recommended data item corresponding to the user to be traded based on the user tower embedding vector and the item tower embedding vector includes: storing the item tower embedding vector in a preset vector library; determining the similarity between the user tower embedding vector and each item tower embedding vector in the preset vector library through the data item recommendation model; and determining the data item corresponding to the item tower vector with the highest similarity as the recommended data item corresponding to the user to be traded.

[0009] Optionally, the method further includes: obtaining a positive and negative sample set, where the positive samples include data products that users have interacted with in the service platform, and the negative samples include data products to be traded and data products that have been exposed but not interacted with in the service platform; training an initial model with the positive and negative sample set to obtain the data product recommendation model.

[0010] Optionally, training the initial model with the positive and negative sample set to obtain the data product recommendation model includes: pairing the product data features corresponding to the positive samples with user features to form positive sample pairs; pairing the product data features corresponding to the negative samples with user features to form negative sample pairs; training the initial model with multiple positive sample pairs and multiple negative sample pairs according to a preset loss function to obtain the data product recommendation model, where the preset loss function is used to represent the difference between the predicted value of the similarity between the user tower embedding vector and the product tower embedding vector and the actual value of the similarity.

[0011] According to another aspect of the embodiments of the present application, there is also provided a data product recommendation device, including: an acquisition module, configured to acquire original data, where the original data includes: historical transaction records of data products, data products to be traded, and information of users to be traded; an extraction module, configured to extract features from the original data to obtain features of data products to be traded and features of users to be traded, where the features of data products to be traded and the features of users to be traded both include: category features, numerical features, and sequence features; a vector module, configured to respectively input the features of data products to be traded and the features of users to be traded into the user tower and the product tower in the data product recommendation model to obtain a user tower embedding vector and a product tower embedding vector; a recommendation module, configured to determine recommended data products corresponding to the users to be traded based on the user tower embedding vector and the product tower embedding vector.

[0012] According to yet another aspect of the embodiments of the present application, there is also provided a computer device, including: a memory and a processor, where the memory is used to store program instructions; the processor is connected to the memory and is configured to execute the above data product recommendation method.

[0013] According to still another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium, which includes a stored computer program, where the device where the non-volatile storage medium is located executes the above data product recommendation method by running the computer program.

[0014] According to still another aspect of the embodiments of the present application, there is also provided a computer program product, including computer instructions, where when the computer instructions are executed by a processor, the above data product recommendation method is implemented.

[0015] In an embodiment of the present application, original data is obtained. The original data includes: historical transaction records of data commodities, data commodities to be traded, and information of users to be traded; feature extraction is performed on the original data to obtain features of data commodities to be traded and features of users to be traded. Among them, the features of data commodities to be traded and the features of users to be traded both include: category features, numerical features, and sequence features; the features of data commodities to be traded and the features of users to be traded are respectively input into the user tower and the commodity tower in a data commodity recommendation model for analysis to obtain a user tower embedding vector and a commodity tower embedding vector; based on the user tower embedding vector and the commodity tower embedding vector, recommended data commodities corresponding to the users to be traded are determined, thereby achieving the purpose of not only using the fixed attribute features of data commodities and users but also combining the dynamic change features used to reflect data commodities and users for data commodity recommendation, thereby realizing the technical effect of improving the accuracy of data commodities, and further solving the technical problem in the related art that the recommendation accuracy is relatively low due to only recommending based on the fixed features of data elements themselves. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0017] Figure 1 is a hardware structure block diagram of a computer terminal for implementing a data commodity recommendation method according to an embodiment of the present application;

[0018] Figure 2 is a flowchart of a data commodity recommendation method according to an embodiment of the present application;

[0019] Figure 3 is a schematic diagram of a two-tower model structure according to an embodiment of the present application;

[0020] Figure 4 is a flowchart of another data commodity recommendation method according to an embodiment of the present application;

[0021] Figure 5 is a structure diagram of a data commodity recommendation device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0023] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0024] The information collected in the embodiments of this application is information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure and application, complies with the relevant laws, regulations and standards of the relevant regions, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered.

[0025] To solve the problems existing in the related art, the embodiments of this application provide a method for recommending data commodities. This method can run on Figure 1 the computer terminal shown below. The following is an explanation of this computer terminal.

[0026] The method embodiments for recommending data commodities provided by the embodiments of this application can be executed on a mobile terminal, a computer terminal or a similar computing device. Figure 1 The following shows a hardware structure block diagram of a computer terminal for implementing the method for recommending data commodities. As Figure 1As shown, the computer terminal 10 may include one or more processors (shown as 102a, 102b, ……, 102n in the figure) (the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions connected through wired and / or wireless networks. In addition, it may further include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown in, or have a different configuration from Figure 1 that shown.

[0027] It should be noted that the above one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" herein. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).

[0028] The memory 104 may be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data product recommendation method in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned data product recommendation method. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor, and these remote memories may be connected to the computer terminal 10 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0029] The transmission module 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0030] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10.

[0031] It should be noted here that in some alternative embodiments, the above Figure 1 illustrated computer terminal may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance and is intended to illustrate the types of components that may exist in the above computer terminal.

[0032] Under the above operating environment, an embodiment of a data commodity recommendation method is provided in the embodiments of the present application. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0033] Figure 2 is a flowchart of a data commodity recommendation method according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:

[0034] Step S202, obtaining original data, where the original data includes: historical transaction records of data commodities, data commodities to be traded, and information of users to be traded;

[0035] Step S204, performing feature extraction on the original data to obtain features of data commodities to be traded and features of users to be traded. Among them, the features of data commodities to be traded and the features of users to be traded both include: category features, numerical features, and sequence features;

[0036] Step S206, respectively inputting the features of data commodities to be traded and the features of users to be traded into the user tower and the commodity tower in the data commodity recommendation model for analysis to obtain a user tower embedding vector and a commodity tower embedding vector;

[0037] Step S208: Determine the recommended data product corresponding to the user to be traded based on the user tower embedding vector and the product tower embedding vector.

[0038] Through the above steps S202 to S208, the original data is obtained, and the original data includes: historical transaction records of data products, data products to be traded, and information of users to be traded; feature extraction is performed on the original data to obtain features of data products to be traded and features of users to be traded. Among them, the features of data products to be traded and the features of users to be traded both include: category features, numerical features, and sequence features; the features of data products to be traded and the features of users to be traded are respectively input into the user tower and the product tower in the data product recommendation model for analysis to obtain a user tower embedding vector and a product tower embedding vector; based on the user tower embedding vector and the product tower embedding vector, the recommended data product corresponding to the user to be traded is determined, thereby achieving the purpose of not only using the fixed attribute features of data products and users but also combining the dynamic change features used to reflect data products and users for data product recommendation, thus realizing the technical effect of improving the accuracy of data products, and further solving the technical problem in the related art that the recommendation accuracy is relatively low due to only recommending based on the fixed features of data elements themselves. The following is a detailed description.

[0039] It should be noted that the data product recommendation model is a two-tower model, including: a user tower and a product tower. Among them, the user tower is used to analyze the features of the user to be traded, and the product tower is used to analyze the features of the data product to be traded.

[0040] It can be understood that the original data can be obtained in various ways. For example: obtaining the buried point data of multiple users and data products through the platform, including the exposure of data products to users, the clicks of users on data products, purchase and other behavior buried points.

[0041] In some embodiments of the present application, feature extraction is performed on the original data to obtain features of data products to be traded and features of users to be traded. The categorical variables in the original data are encoded to obtain the categorical features, where the categorical features are used to represent the category information of the data product to be traded or the category information of the user to be traded; the numerical variables in the original data are non-linearly transformed to obtain the transformed numerical variables, and the transformed numerical variables are encoded to obtain the numerical features, where the numerical features are used to represent the numerical information of the original data; the historical transaction records of data products are extracted from the original data; the sequence features are determined according to the historical transaction records of the data products.

[0042] It should also be noted that Spare Feature (category feature): This type of feature is used to represent classification information, such as the category of data products, the industry to which the user belongs, etc. Such information can be found in the historical transaction records of data products. Through encoding conversion, category variables can be processed by the neural network in numerical form.

[0043] Dense Feature (numerical feature): This type of feature is used to represent specific numerical information, such as the price of data products, the activity of users, etc. Numerical features can be found in the historical transaction records of data products or calculated through other means. Through numerical conversion, such as non-linear transformation (logarithmic processing), the numerical distribution can be adjusted so that features can be effectively learned at different magnitudes.

[0044] Sequence Feature (sequence feature): This feature is determined based on the historical transaction records of data products and reflects the dynamic changes in user behavior. For data products, it records which users have had what kind of interactions and the chronological order of the interactions. The sequence feature is obtained by sorting and encoding the interaction buried-point information related to users or data products in the historical transaction data, and can capture the behavioral patterns in the time series. It is a very crucial part in personalized recommendation. For the user tower, sequence features are constructed using the interaction history between users and products. These sequence values are arranged in chronological order to reflect the changes in user interests. For the product tower, the sequence of user IDs that have interacted with the product is used as the behavioral sequence feature of the product to describe the dynamic changes of the product.

[0045] Among them, the specific steps for determining the sequence feature according to the historical transaction records of the data product are as follows: In the case of obtaining the features of the user to be traded, the user identifier to be traded and the data product identifier in the historical transaction records of the data product are converted into label encodings, and the label encodings are sorted in chronological order to form a first sequence feature, where the first sequence feature is used to represent the change trend of the products of interest to the user to be traded; In the case of obtaining the features of the data product to be traded, the user identifier to be traded and the data product identifier in the historical transaction records of the data product are converted into label encodings, and the label encodings are sorted in chronological order to form a second sequence feature, where the second sequence feature is used to represent the popularity of the data product to be traded. The length of the sequence feature is the total number of identifiers included in the sequence feature plus one.

[0046] In the process of constructing sequence features, to ensure the uniformity of the length of sequence features, the sequence features can be processed in the following ways. For example, obtain the length of the sequence feature; when the length of the sequence feature is greater than the preset length, truncate the sequence feature and only retain the sequence feature with the preset length; when the length of the sequence feature is less than the preset length, pad zeros at the front of the sequence feature until the length of the sequence feature is equal to the preset length.

[0047] The specific steps for constructing sequence features are as follows: Sequence construction: First, based on the historical transaction records of data commodities, construct a sequence that contains the ids of data commodities interacted by users or the ids of users interacted by commodities. These id values are arranged in the order of interaction time. For example, for a user, his historical interaction sequence may be [item_id1, item_id2, item_id3,...], where item_id1 represents the data commodity interacted by the user. Encode each id in the sequence and convert it into an integer form. Since the id may be categorical data, map each different id to a unique integer. For example, item_id1 may be encoded as 0, item_id2 as 1, and so on.

[0048] Table 1 shows the historical transaction records of a data commodity. As shown in Table 1, the sequence feature of item_1 is [user_2, user_1, user_3, user_5, user_4].

[0049] Table 1

[0050] Item (Data Commodity) User Behavior_time (Interaction Time) item_1 user_2 2024.12.01 14:13:25 item_1 user_1 2024.12.01 14:14:16 item_1 user_3 2024.12.01 14:14:18 item_1 user_5 2024.12.01 14:14:20 item_1 user_4 2024.12.01 14:14:21

[0051] The setting process for the vocabulary scale is to set the vocabulary scale (i.e., the size of the vocabulary) of each feature to be equal to the number of types of ids in the sequence plus one. The plus one is to reserve an encoding position for padding or representing unknown ids. If there are 100 different item_ids (data commodity IDs) in the sequence, then the vocabulary scale will be set to 101.

[0052] Convert the encoded sequence into a low-dimensional dense vector through the embedding layer. The embedding layer can capture the similarity relationship between ids and encode this information into the vector. For example, if the embedding dimension is set to 16, then each id will be converted into a vector with a length of 16. Each ID in the sequence corresponds to an embedding vector, and finally all the dense vectors are concatenated.

[0053] Sequence length adjustment: Based on experience, if the sequence length exceeds 50, only the last 50 interaction events are retained for truncation; if the sequence length is less than 50, it is padded with 0s at the front of the sequence to keep the sequence length consistent. This step ensures that all sequences have the same dimension when input into the neural network, facilitating model processing.

[0054] For each sequence feature, average pooling is performed on all embeddings. This means averaging each dimension of all vectors in the sequence to obtain a fixed-length vector, which reflects the average feature of the entire sequence. Average pooling can capture the overall features of the sequence while reducing the computational amount and the risk of overfitting.

[0055] In some embodiments of the present application, the specific steps for determining the recommended data product corresponding to the user to be traded based on the user tower embedding vector and the product tower embedding vector include: storing the product tower embedding vector in a preset vector library; determining the similarity between the user tower embedding vector and each product tower embedding vector in the preset vector library through the data product recommendation model; and determining the data product corresponding to the product tower vector with the highest similarity as the recommended data product corresponding to the user to be traded.

[0056] The process of constructing the two-tower model is as follows: Obtain a positive and negative sample set, where the positive samples include: data products that users have interacted with in the business platform, and the negative samples include: data products to be traded and data products that have been exposed but not interacted with in the business platform; use the positive and negative sample set to train the initial model to obtain the data product recommendation model.

[0057] Among them, the specific steps for training the initial model with the positive and negative sample set to obtain the data product recommendation model are as follows: Pair the product data features corresponding to the positive samples with the user features to form positive sample pairs; pair the product data features corresponding to the negative samples with the user features to form negative sample pairs; use multiple positive sample pairs and multiple negative sample pairs to train the initial model according to a preset loss function to obtain the data product recommendation model, where the preset loss function is used to represent the difference between the predicted value of the similarity between the user tower embedding vector and the product tower embedding vector and the actual similarity value.

[0058] Specifically, positive sample selection: Positive samples are products with which users have direct interaction records. This includes data products for user click, purchase, download, collection, etc. These positive samples are sorted according to the historical interaction time, and the interaction product at the next moment is used as a positive sample, which helps the model learn the evolution of user interests.

[0059] Negative sample selection: Directly associated with the data products provided by the user himself: This part of the samples directly come from the user's data product list on the platform. The user is both the provider and potential consumer of the products. Therefore, these products as negative samples can increase the complexity of the model and the ability of personalized recommendation. Random sampling of products that were exposed but not clicked: Randomly select a part of the products that were exposed to the user but not clicked as negative samples. This reflects the user's potential lack of interest in the products and can help the model learn the key features of whether a product can attract the user.

[0060] Among them, the "data products that were exposed but not clicked" refer to the data products that were shown to the user (i.e., exposed) in the data element trading platform, but the user did not further perform subsequent interaction behaviors such as clicking, viewing details, purchasing, etc.

[0061] In recommendation systems and advertising systems, when a product or advertisement is shown to a user, it is called "exposure". This usually occurs when the user browses a page, and the system, according to certain rules and algorithms, shows a series of data products in front of the user. If the user does not perform any further operations after seeing these data products, such as clicking to enter the details page, downloading, purchasing, etc., then these products are marked as "exposed but not clicked".

[0062] Data products that were exposed but not clicked are also important in the training of recommendation systems because they can provide negative feedback information about the user's interests and behavior patterns. When building a two-tower model, using some data products that were exposed but not clicked as negative samples can help the model learn to distinguish which products are truly attractive to a specific user and which are not. This approach helps the model avoid overfitting to the user's known preferences and enables it to more widely explore data products that the user may be interested in but have not discovered yet, thus improving the diversity of recommendations and the ability to discover new interests.

[0063] In addition, data products that were exposed but not clicked can also help the model understand what factors make a product fail to attract the user, so as to avoid similar mistakes in future recommendations or optimize these products to improve their attractiveness. In this way, the recommendation system can continuously learn and improve, enhancing the overall recommendation effect.

[0064] Such as Figure 3As shown in the figure, the dual - tower model includes two independent neural networks, namely the User tower (user tower) and the Item tower (product tower). Each tower is composed of three layers of neural networks. Each layer of the neural network includes a fully - connected layer, a batch normalization layer, an activation function (such as ReLU), and a Dropout regularization layer. Finally, there is a fully - connected layer that outputs a one - dimensional vector value, which is used to represent the user's embedding (user embedding) and the product's embedding (item embedding), and finally processed by an activation function (softmax). The inputs are user features (user feature) and product features (item feature), and the output dimensions of each layer of the neural network are 128, 64, and 16 respectively.

[0065] Feature input: The User tower and the Item tower respectively receive the feature representations of positive and negative samples, specifically including categorical features, numerical features, and sequential features. Generate low - dimensional dense vector representations of users and products, that is, embeddings.

[0066] Use the cosine similarity function to calculate the similarity of the embeddings output by the User tower and the Item tower. The obtained similarity score represents the user's interest or matching degree for the product.

[0067] Output and sorting: Softmax normalize the similarity for each product to obtain a probability distribution, which represents the user's interest degree for all products. Then, according to this probability distribution, the products can be sorted for subsequent recommendations.

[0068] Finally, use the nearest - neighbor search to return the data products that the user is interested in. Specifically, maintain a data product library to store product embeddings. For new products, the product embedding can be directly added to the data product library according to the output of the item tower. For users, use a retrieval tool according to the user embedding output by the user tower to quickly find the items that match their interests.

[0069] For the acquisition of new data products or user embeddings at a new time, after data acquisition and feature processing steps, directly input them into the corresponding towers respectively, and the corresponding embeddings can be obtained. For the item embedding, directly add it to the vector database; for the user embedding, calculate the similarity with the item embeddings saved in the vector database and return the items that match it.

[0070] In constructing the dual-tower model, negative samples are determined to be data products to be traded and a small number of data products that have been exposed but not interacted with. Starting from the interactive behavior between users and products, combined with the fact that users of data element trading platforms are different from traditional e-commerce users, and that the companies settled in the platform are often both providers and developers of data products, as well as buyers of data products, it is determined that the negative samples for training the model are data products provided by settled companies and a small number of items that have been exposed but not clicked. The advantages are: directly avoiding the problems of inconsistent sample distribution and overcooling of negative samples caused by random negative sampling; the samples provided by the provider directly include all items in the real recommendation scenario; increasing the difficulty of model training and enhancing personalized recommendation capabilities;

[0071] The data product recommendation method provided in this application, based on the fact that the product tower input features have only fixed attributes, sorts the label codes corresponding to the user IDs of historical interactions into a row of sequence values ​​according to the interaction time as the behavioral sequence features of the product, reflects the dynamic changes of data products, and more deeply characterizes the characteristics of items in different periods, while solving the problems of item cold start and item embedding retraining.

[0072] Figure 4 Another data commodity recommendation method is shown, such as Figure 3 As shown, including:

[0073] Step S01: Obtain the historical data of data element transactions on the platform, basic information of the settled enterprises, and basic information of data elements provided by the settled enterprises;

[0074] Step S02: construct user and item tower features;

[0075] Step S03: Establish a dual-tower model to output the embeddings of the user and item towers;

[0076] Step S04: The nearest neighbor search returns the data products that the user is interested in.

[0077] Step S05: Update user and item embedding vectors.

[0078] Figure 5 A data commodity recommendation device according to an embodiment of the present application includes:

[0079] The acquisition module 50 is used to acquire original data, wherein the original data includes: historical transaction records of data commodities, data commodities to be traded, and user information to be traded;

[0080] The extraction module 52 is used to extract features from the original data to obtain features of the commodity to be traded and features of the user to be traded, wherein the features of the commodity to be traded and the features of the user to be traded both include: category features, numerical features and sequence features;

[0081] A vector module 54 for respectively inputting the to-be-traded data commodity features and the to-be-traded user features into a user tower and a commodity tower in a data commodity recommendation model to obtain a user tower embedding vector and a commodity tower embedding vector;

[0082] A recommendation module 56 for determining recommended data commodities corresponding to the to-be-traded user based on the user tower embedding vector and the commodity tower embedding vector.

[0083] Through the above data commodity recommendation device, by acquiring original data, the original data includes: historical transaction records of data commodities, to-be-traded data commodities, and to-be-traded user information; performing feature extraction on the original data to obtain to-be-traded data commodity features and to-be-traded user features, wherein the to-be-traded data commodity features and the to-be-traded user features both include: categorical features, numerical features, and sequential features; respectively inputting the to-be-traded data commodity features and the to-be-traded user features into the user tower and the commodity tower in the data commodity recommendation model for analysis to obtain a user tower embedding vector and a commodity tower embedding vector; determining recommended data commodities corresponding to the to-be-traded user based on the user tower embedding vector and the commodity tower embedding vector, thereby achieving the purpose of not only using the fixed attribute features of data commodities and users but also combining them with the dynamic change features reflecting data commodities and users for data commodity recommendation, thus realizing the technical effect of improving the accuracy of data commodities, and further solving the technical problem in the related art that the recommendation accuracy is relatively low due to only recommending based on the fixed features of data elements themselves.

[0084] The extraction module 52 includes: an extraction sub-module for performing feature extraction on the original data to obtain to-be-traded data commodity features and to-be-traded user features, including: encoding categorical variables in the original data to obtain the categorical features, wherein the categorical features are used to represent the category information of the to-be-traded data commodity or the category information of the to-be-traded user; performing non-linear transformation on numerical variables in the original data to obtain transformed numerical variables, and encoding the transformed numerical variables to obtain the numerical features, wherein the numerical features are used to represent the numerical information of the original data; extracting historical transaction records of data commodities from the original data; and determining the sequential features according to the historical transaction records of the data commodities.

[0085] The extraction sub-module includes: a determination unit, configured to determine the sequence feature according to the historical transaction record of the data commodity, including: when obtaining the feature of the user to be traded, converting the user identifier to be traded and the data commodity identifier in the historical transaction record of the data commodity into label encodings, and sorting the label encodings in chronological order to form a first sequence feature, where the first sequence feature is used to represent the change trend of the commodities that the user to be traded is interested in; when obtaining the feature of the data commodity to be traded, converting the user identifier to be traded and the data commodity identifier in the historical transaction record of the data commodity into label encodings, and sorting the label encodings in chronological order to form a second sequence feature, where the second sequence feature is used to represent the popularity of the data commodity to be traded, and the length of the sequence feature is the total number of identifiers included in the sequence feature plus one.

[0086] The determination unit includes: an acquisition subunit, configured to acquire the length of the sequence feature; when the length of the sequence feature is greater than a preset length, truncate the sequence feature, and only retain the sequence feature with the preset length; when the length of the sequence feature is less than the preset length, pad zeros at the front of the sequence feature until the length of the sequence feature is equal to the preset length.

[0087] The recommendation module 56 includes: a recommendation sub-module, configured to determine the recommended data commodity corresponding to the user to be traded based on the user tower embedding vector and the commodity tower embedding vector, including: storing the commodity tower embedding vector in a preset vector library; determining the similarity between the user tower embedding vector and each commodity tower embedding vector in the preset vector library through the data commodity recommendation model; determining the data commodity corresponding to the commodity tower vector with the highest similarity as the recommended data commodity corresponding to the user to be traded.

[0088] The recommendation module 56 further includes: a training sub-module, configured to obtain a positive and negative sample set, where the positive samples include: the data commodities that the users in the business platform have interacted with, and the negative samples include: the data commodities to be traded and the data commodities that have been exposed but not interacted with in the business platform; training the initial model with the positive and negative sample set to obtain the data commodity recommendation model.

[0089] The training sub-module includes: a training unit for training an initial model using the positive and negative sample sets to obtain the data commodity recommendation model, including: pairing the commodity data features corresponding to the positive samples with user features to form positive sample pairs; pairing the commodity data features corresponding to the negative samples with user features to form negative sample pairs; training the initial model using multiple positive sample pairs and multiple negative sample pairs according to a preset loss function to obtain the data commodity recommendation model, where the preset loss function is used to represent the difference between the predicted value of the similarity between the user tower embedding vector and the commodity tower embedding vector and the actual value of the similarity.

[0090] It should be noted that Figure 5 the data commodity recommendation device shown is used to execute Figure 2 the data commodity recommendation method shown. Therefore, the relevant explanations in the above data commodity recommendation method also apply to this data commodity recommendation device, and will not be elaborated here.

[0091] The embodiment of the present application also provides a computer device, including: a memory and a processor. Among them, the memory is used to store program instructions; the processor is connected to the memory and is used to execute the above data commodity recommendation method.

[0092] The method executed by the above computer device adopts obtaining original data, where the original data includes: historical transaction records of data commodities, data commodities to be traded, and information of users to be traded; performing feature extraction on the original data to obtain features of data commodities to be traded and features of users to be traded. Among them, the features of data commodities to be traded and the features of users to be traded both include: category features, numerical features, and sequence features; respectively inputting the features of data commodities to be traded and the features of users to be traded into the user tower and the commodity tower in the data commodity recommendation model for analysis to obtain a user tower embedding vector and a commodity tower embedding vector; determining the recommended data commodities corresponding to the users to be traded based on the user tower embedding vector and the commodity tower embedding vector, thereby achieving the purpose of not only using the fixed attribute features of data commodities and users but also combining the dynamic change features used to reflect data commodities and users for data commodity recommendation, thus realizing the technical effect of improving the accuracy of data commodities, and further solving the technical problem in the related art that the recommendation accuracy is low due to only recommending based on the fixed features of data elements themselves.

[0093] The embodiment of the present application also provides a non-volatile storage medium, which includes a stored computer program. Among them, the device where the non-volatile storage medium is located executes the above data commodity recommendation method by running the computer program.

[0094] The embodiments of the present application also provide a computer program product, including computer instructions, which implement the steps of the data product recommendation method in the present application when executed by a processor.

[0095] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0096] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0097] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.

[0098] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0099] In addition, the functional units in the respective embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0100] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0101] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A data commodity recommendation method, characterized in that: include: Acquire original data, the original data including: historical transaction records of data products, data products to be traded, and user information to be traded; Extracting features from the original data to obtain features of the commodity to be traded and features of the user to be traded, wherein the features of the commodity to be traded and the features of the user to be traded both include: category features, numerical features and sequence features; Input the characteristics of the data commodity to be traded and the characteristics of the user to be traded into the user tower and the commodity tower in the data commodity recommendation model for analysis, respectively, to obtain a user tower embedding vector and a commodity tower embedding vector; The recommended data commodity corresponding to the to-be-traded user is determined based on the user tower embedding vector and the commodity tower embedding vector.

2. The method according to claim 1, characterized in that Extracting features from the original data to obtain features of the commodity to be traded and features of the user to be traded, including: Encoding the category variables in the original data to obtain the category features, wherein the category features are used to represent the category information of the commodity to be traded or the category information of the user to be traded; Performing nonlinear transformation on the numerical variables in the original data to obtain transformed numerical variables, and encoding the transformed numerical variables to obtain the numerical features, wherein the numerical features are used to represent the numerical information of the original data; Extracting historical transaction records of data commodities from the raw data; The sequence feature is determined according to the historical transaction record of the data commodity.

3. The method according to claim 2, characterized in that Determining the sequence feature according to the historical transaction record of the data commodity includes: In the case of obtaining the characteristics of the user to be traded, converting the identifier of the user to be traded and the identifier of the data commodity in the historical transaction record of the data commodity into a label code, and sorting the label codes in chronological order to form a first sequence feature, wherein the first sequence feature is used to represent the change trend of the commodity that the user to be traded is interested in; When the characteristics of the data commodity to be traded are obtained, the user identifier to be traded and the data commodity identifier in the historical transaction record of the data commodity are converted into label codes, and the label codes are sorted in chronological order to form a second sequence feature, wherein the second sequence feature is used to indicate the popularity of the data commodity to be traded, and the length of the sequence feature is the total number of identifiers contained in the sequence feature plus one.

4. The method according to claim 3, characterized in that The method further comprises: Obtaining the length of the sequence feature; When the length of the sequence feature is greater than a preset length, the sequence feature is truncated to retain only the sequence feature of the preset length; When the length of the sequence feature is less than the preset length, zeros are padded in the front of the sequence feature until the length of the sequence feature is equal to the preset length.

5. The method according to claim 1, characterized in that Determining the recommended data commodity corresponding to the to-be-traded user based on the user tower embedding vector and the commodity tower embedding vector includes: Storing the commodity tower embedding vector in a preset vector library; Determine the similarity between the user tower embedding vector and each product tower embedding vector in the preset vector library through the data product recommendation model; The data commodity corresponding to the commodity tower vector with the highest similarity is determined as the recommended data commodity corresponding to the to-be-traded user.

6. The method according to claim 1, characterized in that The method further comprises: Obtaining positive and negative sample sets, where positive samples include: data commodities that users have interacted with in the business platform, and negative samples include: data commodities to be traded in the business platform and data commodities that have been exposed but not interacted with; The positive and negative sample sets are used to train the initial model to obtain the data product recommendation model.

7. The method according to claim 6, characterized in that The initial model is trained using the positive and negative sample sets to obtain the data commodity recommendation model, including: Pairing the commodity data features and user features corresponding to the positive sample to form a positive sample pair; Pair the product data features corresponding to the negative samples with the user features to form negative sample pairs; The initial model is trained using a plurality of the positive sample pairs and a plurality of the negative sample pairs according to a preset loss function to obtain the data product recommendation model, wherein the preset loss function is used to represent the difference between the predicted value of the similarity between the user tower embedding vector and the product tower embedding vector and the actual value of the similarity.

8. A data commodity recommendation device, characterized in that: include: An acquisition module is used to acquire original data, wherein the original data includes: historical transaction records of data commodities, data commodities to be traded, and user information to be traded; An extraction module is used to extract features from the original data to obtain features of the commodity to be traded and features of the user to be traded, wherein the features of the commodity to be traded and the features of the user to be traded both include: category features, numerical features and sequence features; A vector module, used to input the characteristics of the data commodity to be traded and the characteristics of the user to be traded into the user tower and the commodity tower in the data commodity recommendation model respectively, to obtain the user tower embedding vector and the commodity tower embedding vector; The recommendation module is used to determine the recommended data commodity corresponding to the to-be-traded user based on the user tower embedding vector and the commodity tower embedding vector.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory is used to store program instructions; The processor is connected to the memory, and is used to execute the data commodity recommendation method described in any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the data commodity recommendation method described in any one of claims 1 to 7 is implemented.