Data element recommendation method and device and nonvolatile storage medium

By combining the common characteristics of users and the characteristics of data elements themselves, the set of data elements to be recommended is solved, and the problem of personalization inadequate caused by the data element recommendation method in the prior art relying on the common characteristics of users is achieved, and more efficient data element recommendation is achieved.

CN120011644APending Publication Date: 2025-05-16CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510176611.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, the data element recommendation method only relies on the common characteristics of users, resulting in the recommendation results being not personalized enough and cannot fully meet user needs.

Method used

By determining the historical behavior data set of the user to be recommended, two methods are used to determine the data element set to be recommended: the first method determines the data element set by a set of target users with the same characteristics as the user to be recommended, and the second method determines the data element set by calculating the similarity between the feature vectors of the historical behavior data set and the feature vectors in the vector database, and combines the two to improve the personalization of the recommendation.

Benefits of technology

By determining the data elements to be recommended from both the common characteristics of the user and the data elements themselves, the degree of personalization of the recommendation results can be improved and the user needs can be more fully met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011644A_ABST
    Figure CN120011644A_ABST
Patent Text Reader

Abstract

The invention discloses a data element recommendation method and device and a nonvolatile storage medium. The method comprises the following steps: determining a historical behavior data set corresponding to data elements accessed by a to-be-recommended user in a preset time period; determining a first to-be-recommended data element set based on the historical behavior data set by using a first mode; determining a second to-be-recommended data element set based on the historical behavior data set by using a second mode; and based on the first to-be-recommended data element set and the second to-be-recommended data element set, determining a target to-be-recommended data element set, and pushing data elements ranked in a front preset digit in the target to-be-recommended data element set to a user terminal corresponding to the to-be-recommended user. The data element recommendation method and device solve the technical problems that a recommendation result is not personalized enough and cannot fully meet user requirements due to the fact that a data element recommendation method in the related technology only depends on common characteristics of users for recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing, and more specifically, to a method and device for recommending data elements and a non-volatile storage medium. Background Art

[0002] As the digital economy develops rapidly, the value of data elements as an important means of production is becoming increasingly prominent. Data elements not only carry rich information, but are also the core driving force for innovation and optimization in all walks of life. With the rise of data element trading platforms, how to accurately recommend data elements to meet the personalized needs of users has become a key issue in improving transaction efficiency and promoting data circulation. As a new type of production factor, data elements are new quality productivity in the digital economy era. They have the characteristics of being replicable, shareable, and without physical loss, and have the advantage of increasing returns to scale under the multiplier effect. In order to give full play to the market value of data elements, it is necessary to strengthen the transaction and circulation of data elements under the guidance of a series of policies and regulations.

[0003] When recommending data elements in related technologies, the data elements that the recommended user has visited are identified, and then other users with similar access records are queried to count the access of these users to other data elements, and a list of recommended data elements is generated based on this. Although the recommendation method based on user behavior has improved the personalization of recommendations to a certain extent, this method has obvious limitations. It relies too much on the user's explicit behavior and ignores the content and semantic information of the data elements themselves. Since data elements often contain complex descriptions and classifications, it is difficult to accurately judge the correlation and similarity between data elements based on user behavior alone. Therefore, the data element recommendation method in related technologies only relies on the common characteristics of users for recommendation, resulting in insufficient personalization of the recommendation results and failure to fully meet user needs.

[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0005] The embodiments of the present application provide a data element recommendation method, device and non-volatile storage medium to at least solve the technical problem that the data element recommendation method in the related art only relies on the common characteristics of users for recommendation, resulting in insufficient personalization of the recommendation results and inability to fully meet user needs.

[0006] According to one aspect of an embodiment of the present application, a method for recommending data elements is provided, including: determining a historical behavior data set corresponding to data elements accessed by a to-be-recommended user within a preset time period; determining a first to-be-recommended data element set based on the historical behavior data set using a first method, wherein the first method determines the first to-be-recommended data element set by a target user set having the same characteristics as the to-be-recommended user, the same characteristics referring to the data elements in the historical behavior data set included in the access records of users in the target user set within the preset time period; determining a second to-be-recommended data element set based on the historical behavior data set using a second method, wherein the second method determines the second to-be-recommended data element set by calculating the similarity between a first feature vector set corresponding to the historical behavior data set and all feature vectors in a vector database; determining a target to-be-recommended data element set based on the first to-be-recommended data element set and the second to-be-recommended data element set, and pushing the data elements ranked in the top preset number of the target to-be-recommended data element set to a user terminal corresponding to the to-be-recommended user.

[0007] In some embodiments of the present application, determining a historical behavior data set corresponding to data elements accessed by the to-be-recommended user within a preset time period includes: obtaining historical access records corresponding to data elements accessed by the to-be-recommended user within a preset time period; determining valid access records based on the historical access records, wherein the valid access records are access records in which the stay time of the access behavior in the historical access records exceeds a preset time threshold, and the access behavior includes at least one of the following: click, consultation, and add to purchase; determining the valid access records as the historical behavior data set.

[0008] In some embodiments of the present application, a first method is used to determine a first set of data elements to be recommended based on a historical behavior data set, including: determining a first set of historical data elements corresponding to the historical behavior data set; determining a target user set based on the first historical data element set, wherein user access records of users in the target user set within a preset time period have an intersection with the first historical data element set, and the number of data elements in the intersection is greater than a first preset threshold; determining the first set of data elements to be recommended based on the target user set and the user access records.

[0009] In some embodiments of the present application, a first set of data elements to be recommended is determined based on a target user set and user access records, including: obtaining user access records of all users in the target user set; determining the data elements contained in the user access records, and determining a data element list based on the data elements; counting the number of times each data element in the data element list appears in the user access records of all users to obtain the number of recommendations for each data element, and determining the average number of recommendations for each data element based on the number of recommendations and the total number of users in the target user set, wherein a case where the number of times the same data element appears in the user access record of a user is greater than 1 is recorded as only one occurrence; deleting data elements identified as not of interest from the data element list to obtain a target data element list; sorting the data elements in the target data element list in descending order based on the average number of recommendations for each data element to obtain a sorted target data element list; determining the data elements in the sorted target data element list that are ranked in front of the second preset threshold as the first set of data elements to be recommended.

[0010] In some embodiments of the present application, a second method is used to determine a second set of data elements to be recommended based on a historical behavior data set, including: determining a second set of historical data elements corresponding to the historical behavior data set; determining a first set of feature vectors corresponding to the second historical data element set in a vector database; and determining the second set of data elements to be recommended based on the vector database and the first set of feature vectors.

[0011] In some embodiments of the present application, a second set of data elements to be recommended is determined based on a vector database and a first set of feature vectors, including: for a first feature vector in the first set of feature vectors, calculating the cosine similarity between the first feature vector and each feature vector in the vector database, to obtain a cosine similarity set of the first feature vector, wherein the first feature vector is any feature vector in the first set of feature vectors; determining the data elements corresponding to the cosine similarities ranked before a third preset threshold in the cosine similarity set as the similar data element set of the first feature vector, wherein the similar data element set includes the data element ID of the data element and the cosine similarity corresponding to the data element; after traversing all the feature vectors in the first set of feature vectors, obtaining a similar data element set of each feature vector in the first set of feature vectors; and determining the second set of data elements to be recommended based on the similar data element set of each feature vector in the first set of feature vectors.

[0012] In some embodiments of the present application, a second set of data elements to be recommended is determined based on a similar data element set of each feature vector in a first set of feature vectors, including: for each similar data element set, removing data elements in the similar data element set that overlap with the second historical data element set and data elements identified as not of interest by the user to be recommended, to obtain a set of similar data elements to be recommended for each feature vector in the first set of feature vectors; sorting the similar data element sets to be recommended in descending order according to cosine similarity to obtain a sorted set of similar data elements to be recommended; and determining the data elements in the sorted set of similar data elements to be recommended that are ranked in front of a fourth preset threshold as the second set of data elements to be recommended.

[0013] In some embodiments of the present application, a target set of data elements to be recommended is determined based on a first set of data elements to be recommended and a second set of data elements to be recommended, including: performing a first normalization process on the average number of recommendations of each data element in the first set of data elements to be recommended to obtain the normalized first set of data elements to be recommended, wherein the first normalization process is used to convert the average number of recommendations of each data element in the first set of data elements to be recommended into a numerical value between 0 and 1; performing a second normalization process on the cosine similarity corresponding to each data element in the second set of data elements to be recommended to obtain the normalized second set of data elements to be recommended, wherein the second normalization process is used to convert the cosine similarity corresponding to each data element in the second set of data elements to be recommended into a numerical value between 0 and 1; based on the normalized first set of data elements to be recommended and the normalized second set of data elements to be recommended, determining the target set of data elements to be recommended.

[0014] In some embodiments of the present application, the method also includes: a sub-feature vector in a vector database is obtained in the following manner, wherein the sub-feature vector is any feature vector in the vector database: responding to a user's data element addition or modification instruction, and saving the data element carried in the data element addition or modification instruction to a metadata database, and obtaining the data element ID in the metadata database; determining multiple feature texts of data element attribute features corresponding to the data element ID; determining a first sub-feature vector corresponding to each feature text in a plurality of feature texts based on a preset model and a word segmenter, wherein the preset model is used to convert the feature text into a first sub-feature vector; performing weighted summation processing on the first sub-feature vector corresponding to each feature text based on a preset weight to obtain a sub-feature vector.

[0015] According to another aspect of an embodiment of the present application, a data element recommendation device is also provided, including: a first determination module, used to determine a historical behavior data set corresponding to data elements accessed by a to-be-recommended user within a preset time period; a second determination module, used to determine a first to-be-recommended data element set based on the historical behavior data set using a first method, wherein the first method determines the first to-be-recommended data element set by a target user set having the same characteristics as the to-be-recommended user, and the same characteristics refer to the data elements in the historical behavior data set included in the access records of users in the target user set within a preset time period; a third determination module, used to determine a second to-be-recommended data element set based on the historical behavior data set using a second method, wherein the second method determines the second to-be-recommended data element set by calculating the similarity between a first feature vector set corresponding to the historical behavior data set and all feature vectors in a vector database; a fourth determination module, used to determine a target to-be-recommended data element set based on the first to-be-recommended data element set and the second to-be-recommended data element set, and push the data elements ranked in the front preset number of the target to-be-recommended data element set to the user terminal corresponding to the to-be-recommended user.

[0016] According to another aspect of an embodiment of the present application, a non-volatile storage medium is provided, in which a program is stored, wherein when the program is running, a device where the non-volatile storage medium is located is controlled to execute the recommended method of the above-mentioned data elements.

[0017] According to another aspect of an embodiment of the present application, there is also provided an electronic device, including: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the above-mentioned method for recommending data elements is executed when the program is run.

[0018] According to another aspect of an embodiment of the present application, a computer program product is also provided, including computer instructions, which implement the above-mentioned method for recommending data elements when executed by a processor.

[0019] In an embodiment of the present application, a historical behavior data set corresponding to the data elements accessed by the to-be-recommended user within a preset time period is used to determine; a first method is used to determine a first set of data elements to be recommended based on the historical behavior data set, wherein the first method determines the first set of data elements to be recommended by a target user set having the same characteristics as the to-be-recommended user, and the same characteristics refer to the fact that the access records of the users in the target user set within the preset time period contain the data elements in the historical behavior data set; a second method is used to determine a second set of data elements to be recommended based on the historical behavior data set, wherein the second method determines the second set of data elements to be recommended by calculating the similarity between the first feature vector set corresponding to the historical behavior data set and all the feature vectors in the vector database set; determine the target data element set to be recommended based on the first data element set to be recommended and the second data element set to be recommended, and push the data elements ranked in the front preset number in the target data element set to the user terminal corresponding to the recommended user, obtain different data element sets to be recommended by two methods respectively, and then determine the final target data element set to be recommended based on the data element sets to be recommended determined by the two methods, thereby achieving the purpose of determining the data elements to be recommended from two aspects of the common characteristics of users and the characteristics of the data elements corresponding to the recommended users, thereby solving the technical problem that the data element recommendation method in the related art only relies on the common characteristics of users for recommendation, resulting in the recommendation results not being personalized enough and unable to fully meet user needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0021] Figure 1 It is a hardware structure block diagram of a computer terminal for implementing a method for recommending data elements according to an embodiment of the present application;

[0022] Figure 2 is a flow chart of a method for recommending data elements according to an embodiment of the present application;

[0023] Figure 3 is a flow chart for determining a first set of data elements to be recommended according to an embodiment of the present application;

[0024] Figure 4 is a flow chart of generating and storing feature vectors of a data element according to an embodiment of the present application;

[0025] Figure 5 is a flow chart for determining a second set of data elements to be recommended according to an embodiment of the present application;

[0026] Figure 6This is a data element comprehensive recommendation scoring flow chart according to an embodiment of the present application;

[0027] Figure 7 It is a structural diagram of a data element recommendation device according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0029] The information collected in the embodiments of the present application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or reject automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered.

[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0031] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0032] Data elements: refers to computer data and its derivative forms that are gathered, organized, and processed according to specific production needs in the information age. They have the characteristics of intangibility, reproducibility, economies of scale, multidimensionality, timeliness, security and privacy. They are data assets that can participate in social production and operation activities and bring economic benefits to owners or users.

[0033] Bidirectional Encoder Representations from Transformers (BERT): is a powerful language representation model that learns rich language features through a bidirectional Transformer encoder and can simultaneously consider the text information before and after a word, thereby more accurately understanding the contextual meaning of the text.

[0034] Milvus Vector Database (Milvus for short): is an open source vector database that can store, retrieve and analyze large amounts of vector data. It is designed for large-scale vector data management and can efficiently store and retrieve high-dimensional vectors. It is particularly suitable for image recognition, recommendation systems, natural language processing and other application scenarios that require vector similarity search. It has efficient vector search capabilities and supports millisecond-level nearest neighbor search. It can maintain high performance even at the scale of billions of vectors. It also supports a variety of distance metrics, such as Euclidean distance and cosine similarity, which can adapt to different application scenarios.

[0035] When recommending data elements in the related art, it will identify which data elements the recommended user has visited, and then query other users with similar access records to count the access of these users to other data elements, and generate a data element recommendation list based on this. Therefore, there is a problem that the data element recommendation method in the related art only relies on the common characteristics of users for recommendation, resulting in insufficient personalization of the recommendation results and failure to fully meet user needs. In order to solve this problem, the embodiments of the present application provide relevant solutions, which are described in detail below.

[0036] According to an embodiment of the present application, an embodiment of a method for recommending data elements is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0037] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal for implementing a recommended method for data elements. Figure 1As shown, the computer terminal 10 may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0038] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10. As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0039] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the method for recommending data elements in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the method for recommending data elements mentioned above. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0040] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0041] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0042] In the above-mentioned operating environment, an embodiment of the present application provides an embodiment of a method for recommending data elements. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0043] like Figure 2 FIG. 1 is a flowchart of a method for recommending data elements according to an embodiment of the present application. The method includes:

[0044] Step S202: determine the historical behavior data set corresponding to the data elements accessed by the user to be recommended within a preset time period.

[0045] In the technical solution provided in step S202, there are many ways to implement the determination of the historical behavior data set corresponding to the data elements accessed by the to-be-recommended user within a preset time period, for example: obtaining the historical access records corresponding to the data elements accessed by the to-be-recommended user within a preset time period; determining the valid access records based on the historical access records, wherein the valid access records are access records in which the stay time of the access behavior in the historical access records exceeds a preset time threshold, and the access behavior includes at least one of the following: click, consultation, and add to purchase; determining the valid access records as the historical behavior data set.

[0046] The following is a specific embodiment: The data elements mentioned in this application are actually described and identified in the form of metadata in systems such as data markets or data element trading platforms. Metadata refers to data that describes the attribute characteristics of data elements, including the type of data element, data element ID, data description, source, usage scenario, format, quality, permission information, etc. These metadata transform intangible data elements into descriptive, searchable, and comparable information. For example, a historical behavior data in a historical behavior data set can be described as "retail industry sales data, including monthly sales and product categories from 2018 to 2022, provided in xxx format", and all metadata are stored in the metadata database.

[0047] The display of data elements in the data element system or data element platform is in the form of visual preview (such as charts, maps, time series, etc.), showing part of the information in the metadata (such as data description, source, format, etc.). The preview is generated based on the metadata form of the data element and the data element itself. The purpose of the preview is to help users intuitively understand the structure and content of the data. When users perform operations such as clicks, inquiries, add to cart, transactions, etc., these operations will be recorded by the platform's log system. The recorded content usually includes user ID, data element ID, operation type (click, inquiries, add to cart, transactions, etc.), operation start time, operation end time, etc. When the data element recommendation system is running, it will query the user operation log in the database, extract all operation records of a specific user (i.e., the user to be recommended mentioned above) within a preset time period (e.g., the past 30 days), and generate a historical access record of the data element corresponding to the data element ID in the operation record, and delete the access records in the historical access records whose residence time of the access behavior (access behavior is operation behavior, including but not limited to clicks, consultations, purchases and transactions, and the residence time of access behavior is the difference between the start time and the end time of the operation) is less than the preset time threshold, and obtain valid access records, and determine the valid access records as historical behavior data sets (mainly containing historical records of user interactions with data elements, usually including the following information: User ID: identifies the user who performs the operation. Data element ID: identifies the data element being accessed. Operation type: the type of interaction between the user and the data element, such as clicks, browsing, consultations, purchases, downloads, purchases, etc. Operation start time and end time: the start time and end time of the user's operation, used to calculate the residence time or determine the time sequence of the behavior. Residence time: For certain operations (such as clicks and browsing), the system will record the length of time the user stays on the data element page to evaluate the user's interest in the data element, etc.).

[0048] Step S204, using the first method to determine the first set of data elements to be recommended based on the historical behavior data set, wherein the first method determines the first set of data elements to be recommended by a target user set having the same characteristics as the user to be recommended, and the same characteristics refer to that the access records of users in the target user set within a preset time period contain data elements in the historical behavior data set.

[0049] In the technical solution provided in step S204, there are multiple ways to implement the first method of determining the first set of data elements to be recommended based on the historical behavior data set, for example: determining the first historical data element set corresponding to the historical behavior data set; determining the target user set based on the first historical data element set, wherein the user access records of users in the target user set within a preset time period have an intersection with the first historical data element set, and the number of data elements in the intersection is greater than a first preset threshold; determining the first set of data elements to be recommended based on the target user set and the user access records.

[0050] In the above steps, there are many ways to implement determining the first set of data elements to be recommended based on the target user set and the user access records, for example: obtaining the user access records of all users in the target user set; determining the data elements contained in the user access records, and determining a data element list based on the data elements; calculating the number of times each data element in the data element list appears in the user access records of all users to obtain the number of recommendations for each data element, and determining the average number of recommendations for each data element based on the number of recommendations and the total number of users in the target user set, wherein the case where the number of times the same data element appears in the user access record of a user is greater than 1 is recorded as only one occurrence; deleting the data elements identified as not of interest from the data element list to obtain a target data element list; sorting the data elements in the target data element list in descending order based on the average number of recommendations for each data element to obtain a sorted target data element list; determining the data elements in the sorted target data element list that are ranked in front of the second preset threshold as the first set of data elements to be recommended.

[0051] The following are specific embodiments:

[0052] Determine the first historical data element set corresponding to the historical behavior data set, that is, determine the data element metadata associated with the data element ID in the historical behavior data set as the first historical data element set, query all user operation logs of the log system, and take the users whose data element ID in the operation record in the operation log (that is, the above-mentioned user access record) and the data element of the first historical data element set have an intersection, and the number of data elements (that is, the same data element ID) in the intersection is greater than a first preset threshold (for example, 10) as the target user set (the target user set includes at least the user ID and the data element ID). For users in the target user set, the operation records in the operation log are obtained as user access records, and the data elements included in the user access records are determined. Different data elements (distinguished by data element IDs) form a data element list, and each data element in the data element list (specifically, the number of occurrences of the data element ID) is counted as the number of recommendations for each data element, and the average number of recommendations for each data element is determined based on the number of recommendations and the total number of users in the target user set. The average number of recommendations for each data element is equal to the quotient of the number of recommendations for the data element and the total number of users in the target user set; if the number of times the same data element appears in the user access record of a user is greater than 1, it is only recorded as appearing once, that is, the number of data elements in the access record of each user is 0 or 1, and the access records of all users are traversed to calculate the number of data element recommendations. The data elements with the identification field of "not interested" in the user access record are deleted from the data element list (that is, in response to the user instruction, the data element ID corresponding to the data element with the special identification "not interested" is deleted from the data element list), and the target data element list is obtained. Sort the data elements in the target data element list in descending order according to the average number of recommendations to obtain a sorted target data element list; determine the data elements in the sorted target data element list that are ranked before a second preset threshold (for example, the top 10) as the first set of data elements to be recommended, and the first set of data elements to be recommended includes {data element ID, average number of recommendations}.

[0053] like Figure 3As shown, it is a flowchart for determining a first set of data elements to be recommended according to an embodiment of the present application. First, the effective access records of the user to be recommended to the data elements in the recent period are obtained (that is, the historical access records corresponding to the data elements accessed by the user to be recommended within the preset time period are obtained above; the effective access records are determined based on the historical access records, and the effective access records are determined as the historical behavior data set), the set of data elements accessed by the user to be recommended (that is, the first historical data element set) is taken out, and other users who have accessed or traded the elements in the set of data elements accessed by the user to be recommended and the number of elements exceeds N are searched (that is, the target user set is determined based on the first historical data element set, and N is the first preset threshold value above), and according to the historical access and transaction records, the data elements accessed or traded by users with the same characteristics as the user to be recommended are respectively extracted, and the number of accesses or transactions is recorded as 1 (that is, the user access records of all users in the target user set are obtained above; the data elements contained in the user access records are determined, and the data elements are determined based on the data elements element list; count the number of times each data element in the data element list appears in the user access records of all users, and the historical access and transaction records are the above-mentioned user access records), filter out the data elements in the categories that are not of interest to the recommended users (that is, the data elements identified as not of interest are deleted from the data element list to obtain a target data element list), count the average number of visits and transactions to the remaining data elements by users with common characteristics, that is, the number of times these data elements are recommended by users with common characteristics, and take the top K (K = the final number to be recommended) with the highest average number of recommendations as candidate data elements to be recommended (that is, sort the data elements in the target data element list from large to small based on the average number of recommendations of each data element, to obtain a sorted target data element list; determine the data elements in the sorted target data element list that are ranked in front of the second preset threshold as the first set of data elements to be recommended, K is the second preset threshold, and the average number of visits and transactions to the remaining data elements by users with common characteristics is the above-mentioned average number of recommendations).

[0054] For example, the target user set includes user U1 and user U2. User U1 accesses data elements with data element IDs D1, D2, and D3, and D1 is marked as uninteresting by user U1. User U2 accesses data elements with data element IDs D3 and D4, and D1, D2, D3, and D4 are all data elements included in the first historical data element set. The data element list includes data element IDs: D1, D2, D3, and D4, and associated data elements (i.e., metadata). The number of recommendations for data elements corresponding to D1, D2, D3, and D4 are 1, 1, 2, and 1, respectively. Since D1 is marked as uninteresting by user U1, D1 is deleted from the data element list, and the target data element list is D2, D3, and D4. The data elements in the target data element list are sorted in descending order based on the average number of recommendations for each data element, and the sorted target data element list is obtained: {D3, recommended 1 time on average}; {D2, recommended 0.5 times on average}; {D4, recommended 0.5 times on average}.

[0055] Step S206, using a second method to determine a second set of data elements to be recommended based on the historical behavior data set, wherein the second method determines the second set of data elements to be recommended by calculating similarity between a first feature vector set corresponding to the historical behavior data set and all feature vectors in a vector database.

[0056] In the technical solution provided in step S206, there are multiple ways to implement the second method of determining the second set of data elements to be recommended based on the historical behavior data set, for example: determining the second set of historical data elements corresponding to the historical behavior data set; determining the first set of feature vectors corresponding to the second set of historical data elements in the vector database; determining the second set of data elements to be recommended based on the vector database and the first set of feature vectors.

[0057] In the above steps, there are multiple ways to implement determining the second set of data elements to be recommended based on the vector database and the first feature vector set, for example: for the first feature vector in the first feature vector set, calculate the cosine similarity between the first feature vector and each feature vector in the vector database to obtain the cosine similarity set of the first feature vector, wherein the first feature vector is any feature vector in the first feature vector set; determine the data element corresponding to the cosine similarity ranked before the third preset threshold in the cosine similarity set as the similar data element set of the first feature vector, wherein the similar data element set includes the data element ID of the data element and the cosine similarity corresponding to the data element; after traversing all the feature vectors in the first feature vector set, obtain the similar data element set of each feature vector in the first feature vector set; determine the second set of data elements to be recommended based on the similar data element set of each feature vector in the first feature vector set.

[0058] Determining the second set of data elements to be recommended based on the similar data element set of each feature vector in the first set of feature vectors can be achieved in the following manner: for each set of similar data elements, removing the data elements in the similar data element set that overlap with the second set of historical data elements and the data elements identified as not of interest to the user to be recommended, to obtain the set of similar data elements to be recommended for each feature vector in the first set of feature vectors; sorting the set of similar data elements to be recommended in descending order of cosine similarity to obtain the sorted set of similar data elements to be recommended; and determining the data elements in the sorted set of similar data elements to be recommended that are ranked in front of the fourth preset threshold as the second set of data elements to be recommended.

[0059] The following are specific embodiments:

[0060] Determine the second historical data element set corresponding to the historical behavior data set (the second historical data element set has the same content as the first historical data element set), and determine the first feature vector set corresponding to the second historical data element set in the vector database. Specifically: for each data element in the second historical data element set, read the feature vector corresponding to the data element ID from the vector database (for example, a Milvus vector database), and these feature vectors constitute the first feature vector set. For the first feature vector in the first feature vector set, calculate the cosine similarity between the first feature vector and each feature vector in the vector database to obtain the cosine similarity set of the first feature vector, and determine the data element corresponding to the cosine similarity of the first third preset threshold (for example, the third preset threshold is round(K / 2), K is a constant, round(K / 2): is to round the result of K / 2 to ensure that the selected number is an integer, for example, K is 20, then round(K / 2) is 10) in the cosine similarity set as the similarity data element of the first feature vector. Prime set, after traversing all feature vectors in the first feature vector set, obtain the similar data element set of each feature vector in the first feature vector set, that is, for each feature vector in the first feature vector set, calculate the cosine similarity between it and all other feature vectors in the vector database, obtain the cosine similarity set of each feature vector in the first feature vector set, take the data element corresponding to the cosine similarity ranked before the third preset threshold in the cosine similarity set as the similar data element set of the feature vector, and finally obtain the similar data element set of each feature vector in the first feature vector set.

[0061] For each vector feature's similar data element set, remove the data elements in the similar data element set that overlap with the second historical data element set and the data elements identified as not interesting by the user to be recommended, that is, remove the data elements that the recommended user has visited, so as to reduce the impact of completely consistent similarity on the recommendation effect, and at the same time remove the data elements in the similar data element set that are identified as not interesting by the user to be recommended, to obtain a similar data element set to be recommended for each feature vector in the first feature vector set, sort the data elements in the similar data element set to be recommended in descending order of cosine similarity, to obtain a sorted similar data element set to be recommended, the format of the sorted similar data element set to be recommended is a record list {data element ID, cosine similarity}, and the data elements in the sorted similar data element set to be recommended that are ranked in the top fourth preset threshold (for example, the top 10) are determined as the second data element set to be recommended (format is {data element ID, cosine similarity}).

[0062] The sub-feature vectors in the vector database (for example, Milvus vector database) mentioned in each of the above steps are obtained in the following manner, wherein the sub-feature vector is any feature vector in the vector database: responding to a user's data element addition or modification instruction, and saving the data element (in the form of metadata) carried in the data element addition or modification instruction to the metadata database, and obtaining the data element ID in the metadata database; determining multiple feature texts of data element attribute features corresponding to the data element ID; determining the first sub-feature vector corresponding to each feature text in the multiple feature texts based on a preset model (for example, BERT Chinese model (bert-base-chinese)) and a word segmentor, wherein the preset model is used to convert the feature text into a first sub-feature vector; performing weighted summation processing on the first sub-feature vector corresponding to each feature text based on a preset weight to obtain a sub-feature vector.

[0063] The following are specific embodiments:

[0064] When a user (any user using the data element system) needs to add or modify a data element, he / she logs in to the data element system, first saves the metadata of the data element to the metadata database, and obtains the data element ID. Metadata includes key attributes such as the type of data element, data element ID, data description, source, usage scenario, format, quality, and permission information. The purpose of saving metadata is to convert intangible data elements into specific information for easy description, search, and comparison. The data element attribute characteristics are obtained from the saved metadata, such as industry classification, application scenario, element introduction, element classification, calculation type, service provider, and other information. This information is stored in the metadata in the form of text. First, the data element attribute characteristics are spliced ​​into feature text 1 in the order of industry classification, application scenario, element introduction, element classification, calculation type, and service provider, such as "Industry classification is XXX|Application scenario is XXX|..."; the service introduction of the data element attribute characteristics is used as feature text 2; the service introduction video or audio is converted into text as feature text 3; in addition, other feature texts can be added according to actual conditions. The acquisition of feature text ensures that analysis can be performed based on the multidimensional attributes of data elements.

[0065] Then enter the preprocessing stage, remove unnecessary elements such as spaces, line breaks, punctuation marks, etc. in feature text 1, feature text 2, and feature text 3, shorten the text length, and facilitate processing. Then use the word segmenter to segment the preprocessed multiple feature texts and remove stop words. When segmenting, the word segmentation rules are as follows: Set the parameter padding to max_length to fill the sequence to the maximum length. Padding is max_length, which means that after segmentation, if the sequence length is less than the preset maximum length (i.e., max_length), the filler will be used to fill the sequence to the maximum length. The purpose of this is to make all input sequence lengths consistent, which is convenient for model processing. Set the parameter truncation to True, which means that the sequence exceeding the maximum length after word segmentation will be truncated and only the maximum length part will be retained; set the parameter return_tensors to pt (tensor of pytorch type, pytorch is an open source machine learning library for building and executing deep neural networks, tensor is the basic data structure used in pytorch and many other deep learning frameworks), which means that the sequence returned by the word segmenter will be converted to the pytorch tensor type, which is the input format accepted by the BERT model.

[0066] The multiple feature texts after word segmentation are encoded through the BERT Chinese model, and the first word or token vector of the last hidden state (last_hidden_state) in the output of the BERT Chinese model is obtained as the feature text vector of this data element (i.e., the first sub-feature vector mentioned above). In the BERT model, after the text input is processed by multiple layers of Transformer encoders, each layer will generate a hidden state (hidden_state), which contains the model's understanding and feature representation of the input text. In the output of the last layer, "last_hidden_state" refers to the last hidden state generated by the model after processing the input text. It is usually a multi-dimensional tensor, and each row corresponds to the vector representation of a word or token in the input sequence. After obtaining the first sub-feature vector of each feature text, all first sub-feature vectors are weighted summed based on the preset weights and normalized to obtain the sub-feature vector of the data element. The vector database stores the feature vectors of all data elements and the corresponding data element IDs. Specifically, in the data vector library, they are stored in a data element set (for example, named data_element set, stored in the form of a database table) pre-set in the data vector library (for example, Milvus vector database). At the same time, in order to speed up the calculation of cosine similarity, an index with a metric type (metric_type) of cosine similarity (COSINE) and an index type (index_type) of inverted file-flat index (IVF_FLAT) is created in advance for the feature vector (vector) column of the data_element set, and the data set is divided into multiple clusters (also called "vector clusters"), and then an index is created in the cluster center. IVF_FLAT uses a flat index (FLAT) in each cluster. For cosine similarity queries, establishing an IVF_FLAT index can significantly speed up the query process, especially when the amount of data is large.

[0067] like Figure 4As shown, it is a feature vector generation and storage flow chart of a data element according to an embodiment of the present application. First, the user logs in to the data element system, the user adds or modifies a data element, clicks to save the data element, connects to the metadata database, saves the data element to the metadata database and returns the data element ID (that is, the above-mentioned response to the user's data element addition or modification instruction, and saves the data element carried in the data element addition or modification instruction to the metadata database, and obtains the data element ID in the metadata database); the data element attributes are combined into multiple feature texts according to type and importance, and preprocessing such as removing stop words is performed (that is, the above-mentioned multiple feature texts of the data element attribute characteristics corresponding to the data element ID are determined); the transformers library is used to load the BERT model and the custom word segmenter for the multiple feature texts respectively. Perform word segmentation and encoding, take the first token vector of the last hidden state in the output as the feature text vector (i.e., the first sub-feature vector corresponding to each feature text in multiple feature texts is determined based on the preset model (e.g., BERT Chinese model (bert-base-chinese)) and the word segmentor, wherein the preset model is used to convert the feature text into the first sub-feature vector); perform weighted summation and normalization on multiple feature text vectors of the data element to obtain the data element feature vector (i.e., the weighted summation of the first sub-feature vector corresponding to each feature text based on the preset weight is performed to obtain the sub-feature vector); finally, connect to Milvus (i.e., the above-mentioned vector database), and save the data element ID and its feature vector to the pre-created data_element collection.

[0068] like Figure 5As shown, it is a flowchart for determining a second set of data elements to be recommended according to an embodiment of the present application. First, the valid access records of the user to be recommended to the data elements in the recent period are obtained (i.e., the historical access records corresponding to the data elements accessed by the user to be recommended within the preset time period are obtained as mentioned above; the valid access records are determined based on the historical access records), the set of data element IDs accessed by the user to be recommended is taken out (i.e., the second set of historical data elements corresponding to the historical behavior data set determined as mentioned above), the Milvus vector database is connected, and the feature vectors corresponding to the data element IDs accessed by the user to be recommended are searched from the data_element set (i.e., the first set of feature vectors corresponding to the second set of historical data elements determined as mentioned above in the vector database, and the Milvus vector database is the above-mentioned vector database); the feature vectors of the previous step are traversed, and for each feature vector, the TOP is taken out from the data_element set of Milvus according to the cosine similarity. round(K / 2) data element IDs and record the cosine similarity (i.e., the data elements corresponding to the cosine similarity ranked in the top third preset threshold in the cosine similarity set are determined as the similar data element set of the first feature vector, and after traversing all feature vectors in the first feature vector set, the similar data element set of each feature vector in the first feature vector set is obtained; the second data element set to be recommended is determined based on the similar data element set of each feature vector in the first feature vector set); the data elements visited by the user to be recommended and the data elements in the marked categories of no interest are removed (i.e., for each similar data element set, removing the data elements in the similar data element set that overlap with the second historical data element set and the data elements identified as not interested by the user to be recommended, to obtain the similar data element set to be recommended for each feature vector in the first feature vector set); sorting the cosine similarity values ​​from high to low and taking the first K as candidate data elements to be recommended (that is, the above-mentioned similar data element set to be recommended is sorted from large to small according to the cosine similarity to obtain the sorted similar data element set to be recommended; the data elements ranked in the front of the fourth preset threshold in the sorted similar data element set to be recommended are determined as the second data element set to be recommended).

[0069] Step S208, determining a target data element set to be recommended based on the first data element set to be recommended and the second data element set to be recommended, and pushing the data elements ranked in front of a preset number of places in the target data element set to the user terminal corresponding to the user to be recommended.

[0070] In the technical solution provided in step S208, there are multiple ways to implement determining the target set of data elements to be recommended based on the first set of data elements to be recommended and the second set of data elements to be recommended, for example: performing a first normalization process on the average number of recommendations of each data element in the first set of data elements to be recommended to obtain the normalized first set of data elements to be recommended, wherein the first normalization process is used to convert the average number of recommendations of each data element in the first set of data elements to be recommended into a numerical value between 0 and 1; performing a second normalization process on the cosine similarity corresponding to each data element in the second set of data elements to be recommended to obtain the normalized second set of data elements to be recommended, wherein the second normalization process is used to convert the cosine similarity corresponding to each data element in the second set of data elements to be recommended into a numerical value between 0 and 1; determining the target set of data elements to be recommended based on the normalized first set of data elements to be recommended and the normalized second set of data elements to be recommended.

[0071] The following are specific embodiments:

[0072] The average number of recommendations for each data element in the first set of recommended data elements is first normalized to make the features of the data elements comparable in value and avoid over-weighting features with larger values. The minimum-maximum (Min-Max) normalization method is used to convert the number of visits to the [0, 1] interval, and the formula is as follows:

[0073] Normalized average number of recommendations X ‘ =(XX min )÷(X max -X min ) Among them, X ‘ represents the normalized average number of recommendations for any data element in the first set of recommended data elements, X represents the average number of recommendations for the data element, and X min represents the minimum value of the average number of recommendations of all data elements in the first set of recommended data elements, X max Indicates the maximum value of the average recommendation times of all data elements in the first set of data elements to be recommended.

[0074] The normalized first set of data elements to be recommended includes the data element ID of each data element and the corresponding normalized average number of recommendations.

[0075] The cosine similarity corresponding to each data element in the second set of data elements to be recommended is subjected to a second normalization process to obtain the normalized second set of data elements to be recommended. Since the value range of cosine similarity is between -1 and 1, the second normalization is performed by the following formula: normalized cosine similarity = (cosine similarity + 1) / 2, and the value range of the cosine similarity of the normalized data element is converted from -1 to 1 to the interval [0, 1]. The normalized second set of data elements to be recommended includes the data element ID of the data element and the corresponding normalized cosine similarity.

[0076] Based on the normalized first set of data elements to be recommended and the normalized second set of data elements to be recommended, determine the target set of data elements to be recommended: multiply each normalized average number of recommendations and normalized cosine similarity by a preset weight, respectively, to obtain the weighted average number of recommendations of each data element in the normalized first set of data elements to be recommended and the weighted cosine similarity of each data element in the normalized second set of data elements to be recommended, and take the sum of the weighted average number of recommendations and the weighted cosine similarity of the same data element as the comprehensive recommendation score of the data element. Sort all data elements in the normalized first set of data elements to be recommended and the normalized second set of data elements to be recommended in descending order of the comprehensive recommendation score to obtain the target set of data elements to be recommended, and push the data elements in the target set of data elements to be recommended that are ranked in the first preset number of places (e.g., the first 10 places) to the user terminal corresponding to the user to be recommended.

[0077] like Figure 6 As shown, it is a flow chart of comprehensive recommendation scoring of data elements according to an embodiment of the present application, firstly, the average number of recommendations of the candidate data elements to be recommended recalled according to users with common characteristics is normalized (i.e., the first normalization is performed on the average number of recommendations of each data element in the first set of data elements to be recommended), and the cosine similarity of the candidate data elements to be recommended recalled according to the similarity of the data elements is normalized (i.e., the second normalization is performed on the cosine similarity corresponding to each data element in the second set of data elements to be recommended), and the comprehensive recommendation score of the data elements is calculated according to the preset weights for the candidate items recalled according to the similarity of users and data elements with common characteristics, and the top K with the highest scores are selected and recommended to the user (i.e., the target set of data elements to be recommended is determined based on the first set of data elements to be recommended and the second set of data elements to be recommended, and the data elements ranked in the top preset number in the target set of data elements to be recommended are pushed to the user terminal corresponding to the user to be recommended).

[0078] The present application embodiment provides a structural diagram of a device for recommending data elements, such as Figure 7 As shown, including:

[0079] The first determination module 702 is used to determine the historical behavior data set corresponding to the data elements accessed by the user to be recommended within a preset time period.

[0080] The second determination module 704 is used to determine the first set of data elements to be recommended based on the historical behavior data set using the first method, wherein the first method determines the first set of data elements to be recommended by a target user set having the same characteristics as the user to be recommended, and the same characteristics refer to that the access records of users in the target user set within a preset time period contain data elements in the historical behavior data set.

[0081] The third determination module 706 is used to determine the second set of data elements to be recommended based on the historical behavior data set using a second method, wherein the second method determines the second set of data elements to be recommended by calculating the similarity between the first feature vector set corresponding to the historical behavior data set and all feature vectors in the vector database.

[0082] The fourth determination module 708 is used to determine a target data element set to be recommended based on the first data element set to be recommended and the second data element set to be recommended, and push the data elements ranked in the top preset position in the target data element set to the user terminal corresponding to the user to be recommended.

[0083] It should be noted that Figure 7 The data elements shown are recommended for performing Figure 2 The recommended approach for the data elements shown is therefore Figure 7 The relevant explanations in the recommended method for the data elements in also apply to the recommended device for the data elements and will not be repeated here.

[0084] It should be noted that the various modules in the above-mentioned data element recommendation device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0085] The embodiment of the present application also provides a non-volatile storage medium, the non-volatile storage medium includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the above data element recommendation method. For example, determine the historical behavior data set corresponding to the data elements accessed by the user to be recommended within a preset time period; use the first method to determine the first data element set to be recommended based on the historical behavior data set, wherein the first method determines the first data element set to be recommended by a target user set having the same characteristics as the user to be recommended, and the same characteristics refer to the access records of users in the target user set within the preset time period containing the data elements in the historical behavior data set; use the second method to determine the second data element set to be recommended based on the historical behavior data set, wherein the second method determines the second data element set to be recommended by calculating the similarity between the first feature vector set corresponding to the historical behavior data set and all feature vectors in the vector database; determine the target data element set to be recommended based on the first data element set to be recommended and the second data element set to be recommended, and push the data elements ranked in the front preset number of the target data element set to the user terminal corresponding to the user to be recommended.

[0086] The embodiment of the present application also provides an electronic device, the electronic device includes a processor, the processor is used to run a program, wherein the above data element recommendation method is executed when the program is running. For example, determine the historical behavior data set corresponding to the data elements accessed by the user to be recommended within a preset time period; use a first method to determine the first data element set to be recommended based on the historical behavior data set, wherein the first method determines the first data element set to be recommended by a target user set having the same characteristics as the user to be recommended, and the same characteristics refer to the data elements in the historical behavior data set included in the access records of the users in the target user set within the preset time period; use a second method to determine the second data element set to be recommended based on the historical behavior data set, wherein the second method determines the second data element set to be recommended by calculating the similarity between the first feature vector set corresponding to the historical behavior data set and all feature vectors in the vector database; determine the target data element set to be recommended based on the first data element set to be recommended and the second data element set to be recommended, and push the data elements ranked in the front preset number of the target data element set to the user terminal corresponding to the user to be recommended.

[0087] According to another aspect of the embodiment of the present application, a computer program product is also provided, including a computer program, which implements the above data element recommendation method when executed by a processor. For example, determine the historical behavior data set corresponding to the data elements accessed by the user to be recommended within a preset time period; use a first method to determine the first data element set to be recommended based on the historical behavior data set, wherein the first method determines the first data element set to be recommended by a target user set having the same characteristics as the user to be recommended, and the same characteristics refer to the data elements in the historical behavior data set included in the access records of the users in the target user set within the preset time period; use a second method to determine the second data element set to be recommended based on the historical behavior data set, wherein the second method determines the second data element set to be recommended by calculating the similarity between the first feature vector set corresponding to the historical behavior data set and all feature vectors in the vector database; determine the target data element set to be recommended based on the first data element set to be recommended and the second data element set to be recommended, and push the data elements ranked in the front preset number of the target data element set to the user terminal corresponding to the user to be recommended.

[0088] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0089] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0090] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0091] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0092] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc. Various media that can store program codes.

[0093] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for recommending data elements, characterized in that: include: Determine a historical behavior data set corresponding to the data elements accessed by the user to be recommended within a preset time period; Determine a first set of data elements to be recommended based on the historical behavior data set using a first method, wherein the first method determines the first set of data elements to be recommended by a target user set having the same characteristics as the user to be recommended, wherein the same characteristics refer to that the access records of users in the target user set within the preset time period contain data elements in the historical behavior data set; Determine a second set of data elements to be recommended based on the historical behavior data set using a second method, wherein the second method determines the second set of data elements to be recommended by calculating similarity between a first feature vector set corresponding to the historical behavior data set and all feature vectors in a vector database; A target set of data elements to be recommended is determined based on the first set of data elements to be recommended and the second set of data elements to be recommended, and data elements ranked in front of a preset number of places in the target set of data elements to be recommended are pushed to a user terminal corresponding to the user to be recommended.

2. The method according to claim 1, characterized in that: The determination of the historical behavior data set corresponding to the data elements accessed by the to-be-recommended user within a preset time period includes: Obtaining historical access records corresponding to the data elements accessed by the to-be-recommended user within the preset time period; Determine a valid access record based on the historical access record, wherein the valid access record is an access record in which the stay time of the access behavior in the historical access record exceeds a preset time threshold, and the access behavior includes at least one of the following: click, consult, and add to purchase; The valid access record is determined as the historical behavior data set.

3. The method according to claim 1, characterized in that The determining the first set of data elements to be recommended based on the historical behavior data set using the first method includes: Determine a first historical data element set corresponding to the historical behavior data set; Determining the target user set according to the first historical data element set, wherein user access records of users in the target user set within the preset time period have an intersection with the first historical data element set, and the number of data elements in the intersection is greater than a first preset threshold; The first set of data elements to be recommended is determined based on the target user set and the user access records.

4. The method according to claim 3, characterized in that The determining the first set of data elements to be recommended based on the target user set and the user access records includes: Obtaining user access records of all users in the target user set; Determine the data elements included in the user access record, and determine a data element list based on the data elements; Counting the number of times each data element in the data element list appears in the user access records of all users to obtain the number of recommendations for each data element, and determining the average number of recommendations for each data element based on the number of recommendations and the total number of users in the target user set, wherein if the number of times the same data element appears in the user access record of a user is greater than 1, it is recorded as only appearing once; Deleting data elements identified as uninteresting from the data element list to obtain a target data element list; sorting the data elements in the target data element list in descending order based on the average recommendation times of each of the data elements to obtain a sorted target data element list; The data elements ranked in front of the second preset threshold in the sorted target data element list are determined as the first set of data elements to be recommended.

5. The method according to claim 1, characterized in that The method of using the second method to determine the second set of data elements to be recommended based on the historical behavior data set includes: Determine a second historical data element set corresponding to the historical behavior data set; Determine a first feature vector set corresponding to the second historical data element set in the vector database; The second set of data elements to be recommended is determined based on the vector database and the first set of feature vectors.

6. The method according to claim 5, characterized in that The determining the second set of data elements to be recommended based on the vector database and the first set of feature vectors includes: For a first feature vector in the first feature vector set, calculating the cosine similarity between the first feature vector and each feature vector in the vector database to obtain a cosine similarity set of the first feature vectors, wherein the first feature vector is any feature vector in the first feature vector set; Determine the data elements corresponding to the cosine similarities ranked before the third preset threshold in the cosine similarity set as the similar data element set of the first feature vector, wherein the similar data element set includes the data element ID of the data element and the cosine similarity corresponding to the data element; After traversing all feature vectors in the first feature vector set, obtaining a similar data element set of each feature vector in the first feature vector set; The second set of data elements to be recommended is determined based on a similar set of data elements of each feature vector in the first set of feature vectors.

7. The method according to claim 6, characterized in that The determining the second set of data elements to be recommended based on the similar data element set of each feature vector in the first set of feature vectors includes: For each of the similar data element sets, remove the data elements in the similar data element set that overlap with the second historical data element set and the data elements identified as not of interest by the to-be-recommended user, to obtain a to-be-recommended similar data element set for each feature vector in the first feature vector set; The similar data element set to be recommended is sorted in descending order according to the cosine similarity to obtain a sorted similar data element set to be recommended; The data elements ranked in front of the sorted similar data elements to be recommended by a fourth preset threshold are determined as the second data elements to be recommended.

8. The method according to claim 1, characterized in that The determining a target set of data elements to be recommended based on the first set of data elements to be recommended and the second set of data elements to be recommended includes: Performing a first normalization process on the average number of recommendations of each data element in the first set of data elements to be recommended, to obtain a normalized first set of data elements to be recommended, wherein the first normalization process is used to convert the average number of recommendations of each data element in the first set of data elements to be recommended into a value between 0 and 1; performing a second normalization process on the cosine similarity corresponding to each data element in the second set of data elements to be recommended, to obtain a normalized second set of data elements to be recommended, wherein the second normalization process is used to convert the cosine similarity corresponding to each data element in the second set of data elements to be recommended into a value between 0 and 1; The target set of data elements to be recommended is determined based on the normalized first set of data elements to be recommended and the normalized second set of data elements to be recommended.

9. The method according to claim 1, characterized in that: The method further includes: obtaining a sub-feature vector in the vector database by the following method, wherein the sub-feature vector is any feature vector in the vector database: Responding to a data element adding or modifying instruction from a user, saving the data element carried in the data element adding or modifying instruction to a metadata database, and obtaining the data element ID in the metadata database; Determine multiple feature texts of data element attribute features corresponding to the data element ID; Determine a first sub-feature vector corresponding to each feature text in the plurality of feature texts based on a preset model and a word segmenter, wherein the preset model is used to convert the feature text into the first sub-feature vector; A weighted sum process is performed on the first sub-feature vector corresponding to each feature text based on a preset weight to obtain the sub-feature vector.

10. A data element recommendation device, characterized in that: include: A first determination module is used to determine a historical behavior data set corresponding to the data elements accessed by the user to be recommended within a preset time period; A second determination module is configured to determine a first set of data elements to be recommended based on the historical behavior data set using a first method, wherein the first method determines the first set of data elements to be recommended by a target user set having the same characteristics as the user to be recommended, wherein the same characteristics refer to that the access records of users in the target user set within the preset time period contain data elements in the historical behavior data set; A third determination module is used to determine a second set of data elements to be recommended based on the historical behavior data set using a second method, wherein the second method determines the second set of data elements to be recommended by calculating similarity between a first feature vector set corresponding to the historical behavior data set and all feature vectors in a vector database; The fourth determination module is used to determine a target set of data elements to be recommended based on the first set of data elements to be recommended and the second set of data elements to be recommended, and push the data elements ranked in the front preset position in the target set of data elements to the user terminal corresponding to the user to be recommended.

11. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a program, wherein when the program is executed, the device where the non-volatile storage medium is located is controlled to execute the method for recommending data elements described in any one of claims 1 to 9.

12. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program, when running, executes the method for recommending data elements described in any one of claims 1 to 9.

13. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the method for recommending data elements described in any one of claims 1 to 9 is implemented.