Data Update Method, Apparatus, Electronic Device, and Computer Readable Medium
The method enhances collaborative filtering by using time and behavior record files to incrementally update indexes, addressing low real-time updates and high computational costs, ensuring efficient and timely data processing.
Patent Information
- Application Number
- CN202111321431.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-09
AI Technical Summary
The existing collaborative filtering technology has problems such as poor real-time and high computing costs, especially when data is updated, which requires full refresh, resulting in inefficiency.
By using time record files and behavior record files collections, batch incremental updates of inverted index files are realized. Combined with Redis as the buffer layer, change records of user behavior data are processed in stages to reduce the frequency of full refresh.
It improves the real-time and update efficiency of data, reduces the calculation cost, and ensures the real-time and accuracy of the recommendation system.
Smart Images

Figure CN114036163B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and in particular, to data update methods, apparatuses, electronic devices, and computer-readable media. Background Art
[0002] Collaborative filtering, generally speaking, usually uses the preferences of a group of people with similar interests and common experiences to recommend information that users are interested in. An individual gives a certain degree of response (such as rating) to the information through a cooperation mechanism and records it to achieve the purpose of filtering, thereby helping others to screen information. The response is not necessarily limited to particularly interesting information. The record of particularly uninteresting information is also quite important.
[0003] This recommendation method usually has the following problems:
[0004] First, the real-time performance is poor and the data update efficiency is low. Since data is always changing, and the changed data is often only a very small part of all the data. Therefore, it is very inefficient to make a new copy (reload all the data to update the model) of a data set containing millions or even tens of millions of data just for the update of this small part of data. Generally, the model is updated every few days or longer.
[0005] Second, the computing cost is relatively high. Generally, it is necessary to load all user, item, and preference data into the memory for similarity calculation. When the data volume is very large, the cost of calculating the similarity values of all users is extremely high (that is, it occupies a relatively high amount of the central processing unit and memory). Summary of the Invention
[0006] The content part of the present disclosure is used to briefly introduce concepts, which will be described in detail in the following detailed implementation part. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution. Some embodiments of the present disclosure propose data update methods, apparatuses, electronic devices, and computer-readable media to solve one or more of the technical problems mentioned in the above background art part.
[0007] In a first aspect, some embodiments of the present disclosure provide a data update method, which includes: processing the obtained user behavior log to obtain target data; according to the target data, changing the data in the relevant behavior record files in the time record file and the behavior record file set, wherein the time record file stores the user data of each behavior occurring within a preset time period, and each behavior record file in the behavior record file set stores the behavior data of different users; in response to determining that the current time reaches the update time, determining whether there is data change within the preset time period according to the time record file; in response to determining that there is, updating the inverted index file according to the time record file and the behavior record file set.
[0008] In some embodiments, the target data includes a user identifier, a content identifier of the behavior content, and a behavior time; and the file name of the behavior record file in the behavior record file set is the user identifier, and the file content includes the content identifier of the behavior content and the behavior time; the file name of the time record file is the preset time period, and the file content includes the user identifier.
[0009] In some embodiments, in response to determining that there is, updating the inverted index file according to the time record file and the behavior record file set includes: obtaining the set of user identifiers for which data has changed from the time record file; for each user identifier in the set of user identifiers, selecting from the behavior record file set the behavior record file corresponding to the user identifier, and generating the updated data of the user indicated by the user identifier according to the selected behavior record file; updating the inverted index file according to the generated updated data.
[0010] In some embodiments, generating the updated data of the user indicated by the user identifier according to the selected behavior record file includes: generating the updated data of the user indicated by the user identifier according to the user identifier and at least one content identifier in the selected behavior record file.
[0011] In some embodiments, changing the data in the relevant behavior record files in the time record file and the behavior record file set according to the target data includes: determining whether there is a target behavior record file in the behavior record file set, wherein the file name of the target behavior record file is the user identifier in the target data; in response to determining that there is, determining whether the target content identifier is included in the target behavior record file, wherein the target content identifier is the content identifier in the target data; in response to determining that it is included, changing the behavior time corresponding to the target content identifier to the behavior time in the target data.
[0012] In some embodiments, according to the target data, changing the data in the relevant behavior record files in the time record file and the set of behavior record files further includes: in response to determining that it is not included, adding the content identifier and the behavior time in the target data to the target behavior record file; determining whether the number of records in the target behavior record file reaches a preset number; in response to determining that it reaches, deleting the record with the earliest behavior time in the target behavior record file.
[0013] In some embodiments, according to the target data, changing the data in the relevant behavior record files in the time record file and the set of behavior record files further includes: determining whether the behavior time in the target data is within a preset time period; in response to determining that it is within, determining whether the user identifier in the target data is included in the time record file; in response to determining that it is not included, adding the user identifier in the target data to the time record file.
[0014] In some embodiments, the method further includes: obtaining a list of content identifiers of a target user from an inverted index file; using the content identifier at a preset position in the list of content identifiers as query information, and obtaining a set of lists of content identifiers containing the query information from the inverted index file; generating a recommended list of content identifiers for the target user based on the set of lists of content identifiers.
[0015] In some embodiments, generating a recommended list of content identifiers for the target user based on the set of lists of content identifiers includes: aggregating the content identifiers in each list of content identifiers in the set of lists of content identifiers to obtain a set of content identifier groups; removing the content identifier groups in the set of content identifier groups that contain the content identifiers in the list of content identifiers to obtain a set of candidate content identifier groups; sorting each candidate content identifier group in the set of candidate content identifier groups according to the number of candidate content identifiers within the group, and generating a recommended list of content identifiers for the target user according to the sorting result.
[0016] In some embodiments, processing the obtained user behavior log to obtain target data includes: determining whether the obtained user behavior log includes data representing a target behavior; in response to determining that it includes, extracting key data from the user behavior log to generate target data.
[0017] Second aspect, some embodiments of the present disclosure provide a data update device, which includes: a processing unit configured to process the acquired user behavior logs to obtain target data; a change unit configured to change the data in the relevant behavior record files in the time record file and the behavior record file set according to the target data, wherein the time record file stores the user data of each behavior occurring within a preset time period, and each behavior record file in the behavior record file set stores the behavior data of different users; a change determination unit configured to, in response to determining that the current time reaches the update time, determine whether there is data change within the preset time period according to the time record file; an update unit configured to, in response to determining that there is, update the inverted index file according to the time record file and the behavior record file set.
[0018] In some embodiments, the target data includes a user identifier, a content identifier of the behavior content, and a behavior time; and the file name of the behavior record file in the behavior record file set is the user identifier, and the file content includes the content identifier of the behavior content and the behavior time; the file name of the time record file is the preset time period, and the file content includes the user identifier.
[0019] In some embodiments, the update unit is further configured to obtain a set of user identifiers with data change from the time record file; for each user identifier in the set of user identifiers, select the behavior record file corresponding to the user identifier from the behavior record file set, generate update data of the user indicated by the user identifier according to the selected behavior record file; and update the inverted index file according to the generated update data.
[0020] In some embodiments, the update unit is further configured to generate update data of the user indicated by the user identifier according to the user identifier and at least one content identifier in the selected behavior record file.
[0021] In some embodiments, the change unit includes a first change subunit configured to: determine whether there is a target behavior record file in the behavior record file set, wherein the file name of the target behavior record file is the user identifier in the target data; in response to determining that there is, determine whether the target content identifier is included in the target behavior record file, wherein the target content identifier is the content identifier in the target data; and in response to determining that it is included, change the behavior time corresponding to the target content identifier to the behavior time in the target data.
[0022] In some embodiments, the first change subunit is further configured to: in response to determining that it does not contain, add the content identifier and the behavior time in the target data to the target behavior record file; determine whether the number of records in the target behavior record file reaches a preset number; in response to determining that it reaches, delete the record with the earliest behavior time in the target behavior record file.
[0023] In some embodiments, the change unit further includes a second change subunit, which is configured to: determine whether the behavior time in the target data is within a preset time period; in response to determining that it is within, determine whether the user identifier in the target data is included in the time record file; in response to determining that it is not included, add the user identifier in the target data to the time record file.
[0024] In some embodiments, the apparatus further includes: an acquisition unit, which is configured to acquire a list of content identifiers of a target user from an inverted index file; a query unit, which is configured to use the content identifier at a preset position in the list of content identifiers as query information and acquire a set of lists of content identifiers including the query information from the inverted index file; a generation unit, which is configured to generate a recommended list of content identifiers of the target user based on the set of lists of content identifiers.
[0025] In some embodiments, the generation unit further includes: an aggregation subunit, which is configured to aggregate each content identifier in each list of content identifiers in the set of lists of content identifiers to obtain a set of content identifier groups; a removal subunit, which is configured to remove the content identifier groups that include the content identifiers in the list of content identifiers from the set of content identifier groups to obtain a set of candidate content identifier groups; a generation subunit, which is configured to sort each candidate content identifier group in the set of candidate content identifier groups according to the number of candidate content identifiers within the group, and generate a recommended list of content identifiers of the target user according to the sorting result.
[0026] In some embodiments, the processing unit is further configured to determine whether the acquired user behavior log includes data characterizing the target behavior; in response to determining that it includes, extract key data from the user behavior log to generate target data.
[0027] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device, on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect above.
[0028] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium, on which a computer program is stored, and wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect above is implemented.
[0029] The above - mentioned various embodiments of the present disclosure have the following beneficial effects: The data update method of some embodiments of the present disclosure can effectively improve the real - time performance of data in the inverted index file (equivalent to the query file). Specifically, the reason for the poor real - time performance of collaborative filtering technology is that each update requires a full - volume refresh of data. Based on this, the time record file and the behavior record files in the set of behavior record files in the data update method of some embodiments of the present disclosure can achieve real - time recording of key data in user behavior data. When the update time is reached, the inverted index file can be incrementally updated in batches according to these files. Also, because of the time record file and the set of behavior record files, the update frequency of the inverted index file can be greatly reduced. This can not only ensure the real - time performance of data but also improve the data update efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In combination with the accompanying drawings and with reference to the following specific embodiments, the above - mentioned and other features, advantages, and aspects of each embodiment of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that elements and elements are not necessarily drawn to scale.
[0031] Figure 1 is an architectural diagram of an exemplary system to which some embodiments of the present disclosure can be applied;
[0032] Figure 2 is a flowchart of some embodiments of the data update method according to the present disclosure;
[0033] Figure 3 is a flowchart of some other embodiments of the data update method according to the present disclosure;
[0034] Figure 4 is a schematic diagram of an application scenario of the data update method according to some embodiments of the present disclosure;
[0035] Figure 5 is a schematic structural diagram of some embodiments of the data update device according to the present disclosure;
[0036] Figure 6 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0038] In addition, it should be noted that for ease of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0039] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.
[0040] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".
[0041] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0042] The present disclosure will be described in detail below with reference to the drawings and in combination with embodiments.
[0043] Figure 1 An exemplary system architecture 100 of a data update method or a data update device to which some embodiments of the present disclosure can be applied is shown.
[0044] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, a server 103, and database servers 104 and 105. The network 102 can be a medium for providing a communication link between the terminal device 101, the server 103, and the database servers 104 and 105. The network 102 can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0045] Users can use the terminal device 101 to interact with the server 103 and the database servers 104 and 105 through the network 102 to receive or send messages, etc. Various client applications can be installed on the terminal device 101, such as a web browser, an article (or novel) application, a shopping application, and an instant messaging tool, etc.
[0046] The terminal device 101 here can be either hardware or software. When the terminal device 101 is hardware, it can be various electronic devices with a display screen, including but not limited to smartphones, tablet computers, e-book readers, laptop computers, desktop computers, and so on. When the terminal device 101 is software, it can be installed in the above-listed electronic devices. It can be implemented as, for example, multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made here.
[0047] The database servers 104 and 105 can be servers that can provide data support for the applications installed on the terminal device 101 and store the user behavior logs in the applications.
[0048] The server 103 can be a server that provides various services. For example, it can be a background server that supports the search engine of the application platform. The background server can obtain the user behavior logs from the database servers 104 and 105 in real time, thereby changing the time record file and the relevant behavior record file. Furthermore, based on these files, it can periodically update the inverted index file in the search engine. And the background server can also simulate collaborative filtering to generate recommendation information (such as a list of recommended content identifiers) based on the updated inverted index file, and can send the recommendation information to the terminal device 101.
[0049] The server 103 and the database servers 104 and 105 here can also be either hardware or software. When the server 103 and the database servers 104 and 105 are hardware, they can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server 103 and the database servers 104 and 105 are software, they can be implemented as, for example, multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made here.
[0050] It should be noted that the method provided by the embodiments of the present disclosure can be executed by the server 103 or by the terminal device 101. Correspondingly, the device can be set in the server 103 or in the terminal device 101. No specific limitation is made here.
[0051] It should be noted that in the case where the server 103 has the functions of the database servers 104 and 105, the database servers 104 and 105 may not be set in the system architecture 100 either.
[0052] It should be understood, Figure 1The numbers of the terminal devices, networks, servers, and database servers therein are merely illustrative. According to actual needs, there may be any number of terminal devices, networks, servers, and database servers.
[0053] Continuing to refer to Figure 2 , a flow 200 of some embodiments of the data update method according to the present disclosure is shown. The method includes the following steps:
[0054] Step 201, process the obtained user behavior logs to obtain target data.
[0055] In some embodiments, the execution subject of the data update method (such as Figure 1 the server 103 shown in Figure 1 ) can obtain user behavior logs in different applications in various ways. For example, the execution subject can obtain them from a database server (such as
[0056] the database servers 104, 105 shown in
[0057] ) or the cloud through a wired connection or a wireless connection. Alternatively, user behavior logs can be collected by embedding points at the front end. The execution subject can interface with existing data processing technologies to collect and consume user behavior logs in real time.
[0058] Optionally, when determining the data characterizing the target behavior, the executing entity may further determine whether the operation duration of the target behavior in the user behavior log reaches a preset duration. If it is determined that the duration is reached, the executing entity may extract the key data in the user behavior log to generate the target data. For example, if the duration for a user to click and view an article exceeds 10 seconds, it indicates that the user did not click by mistake or is interested in this article. Only in this way, the target data obtained from this user behavior log is the data we need, which can further improve the accuracy of subsequent content recommendations to users.
[0059] It should be noted that since the information formats of user behavior logs generated by different applications may vary. Therefore, when the executing entity obtains the user behavior log, it can first perform information format conversion, that is, convert it into a format readable by the executing entity (or a unified format). This also helps with subsequent data processing.
[0060] Step 202: According to the target data, change the data in the relevant behavior record files in the time record file and the set of behavior record files.
[0061] In some embodiments, based on the target data obtained in step 201, the executing entity may change the data in the time record file and the data in the relevant behavior record files in the set of behavior record files. Here, the time record file stores the data of each user whose behavior occurs within a preset time period. Each behavior record file in the set of behavior record files stores the behavior data of different users. That is, different behavior record files store the behavior data of different users. The duration of the preset time period can be set according to actual needs, such as several minutes or one hour, etc. And for the convenience of subsequent update of the inverted index file, the preset time period can correspond to the update cycle. That is to say, the preset time period can be the time period between two adjacent update times or any time period in between.
[0062] As an example, the file name of the behavior record file in the set of behavior record files may be related to the user identifier (such as the file name is the user identifier), and the file content may include the content identifier of the behavior content and the behavior time. This can not only effectively distinguish the behavior record files but also help reduce the amount of key data that the file needs to store.
[0063] In this case, first, the execution subject can determine whether the target behavior record file exists in the behavior record file set. The file name of the target behavior record file can be the user identifier in the target data. If it is determined to exist, the execution subject can then determine whether the target behavior record file contains the target content identifier. The target content identifier is the content identifier in the target data. If it is determined to be included, it means that the user has previously operated the behavior content. At this time, the execution subject can change the behavior time corresponding to the target content identifier to the behavior time in the target data.
[0064] Optionally, if it is determined that it is not included, it means that the user has not previously performed the operation of the behavior content. At this time, the execution entity can add the content identifier and behavior time in the target data to the target behavior record file. In addition, in order to ensure the real-time and effectiveness of the data, the execution entity can further determine whether the number of records in the target behavior record file reaches a preset number (such as 20). If it is determined that it is reached, the execution entity can delete the record with the earliest behavior time in the target behavior record file. This can further reduce the memory space occupied. At the same time, since the records of the user's most recent behavior content are retained, it is also beneficial to improve the accuracy of subsequent content recommendations (to the user or other users).
[0065] It is understandable that if the target behavior record file does not exist in the behavior record file set, the execution entity can create a target behavior record file according to the above description and add the content identifier and behavior time in the target data to the target behavior record file.
[0066] In some embodiments, the file name of the above-mentioned time record file can be related to a preset time period (such as the name is a preset time period), and the file content can include a user identifier. At this time, the execution subject can first determine whether the behavior time in the target data is within the preset time period. If it is determined to be within, then it can be determined whether the time record file contains the user identifier in the target data. If it is determined not to be included, the user identifier in the target data can be added to the time record file. If it is determined to be included, the execution subject can do nothing.
[0067] Optionally, the file content of the time record file may also include the behavior time. In this case, if it is determined to be included, the execution subject may change the behavior time corresponding to the user ID to the behavior time in the target data, or may not change the behavior time. It is sufficient to record the corresponding behavior time when the user ID is first added to the time record file, that is, the behavior time when the user indicated by the user ID performs the target behavior for the first time within a preset time period. In this way, while ensuring the real-time nature of the data, the processing process of the execution subject can be reduced, thereby improving processing efficiency.
[0068] In some application scenarios, the executing entity can use Redis (Remote Dictionary Server) to record data changes in real time. For example, for behavior record files, each user can correspond to an ordered set object (zset, sorted set) in Redis. The key of the set (the key in the key-value pair) can be user:{user identifier}. The members in the set can be content identifiers, and the scores in the set can be behavior times. In the file, the date format will be converted to the timestamp format, and at the same time, the uniqueness of the members in the set (such as article IDs) should be ensured. In addition, to improve the data reading efficiency, the records in the file can be arranged in descending order of timestamps.
[0069] As an example, user 100061 viewed article 66 at 14:03:28 on August 28, 2021. At this time, through the Redis command (zadd user:100061 1630130608000 66), the data can be stored in the behavior record file of this user, as shown in the following table:
[0070]
[0071] For another example, for time record files, each hour can correspond to an ordered set object. The key of the set can be changes:{year month day hour}. The members in the set can be user identifiers, and the scores in the set can be behavior times. At the same time, the uniqueness of the members in the set (such as user IDs) should be ensured.
[0072] As an example, user 100061 viewed article 66 at 14:03:28 on August 28, 2021. Through the Redis command (zadd change:2021082814 1630130608000 100061), the member 100061 can be added to the file for the first time. If user 100061 viewed article 78 at 14:05:26 on August 28, 2021. Since the member already exists in the file, the Redis command (zadd change:2021082814 1630130726000 100061) can be used to update its timestamp. The data format of the time record file is shown in the following table:
[0073]
[0074] For the processing procedures of the above steps 201 and 202, the executing entity can interface with a streaming computing platform (such as JD.com Streaming Computing Platform JRC) for data processing and changes.
[0075] Step 203, in response to determining that the current time reaches the update time, determine whether there is data change within a preset time period according to the time record file.
[0076] In some embodiments, when determining that the current time reaches the update time, the execution entity may determine whether there is data change within a preset time period according to the above time record file. That is, determine whether there is recorded data in the time record file. If there is a record, it can indicate that there is data change. At this time, the execution entity may continue to execute Step 204. If there is no record, it indicates that there is no data change.
[0077] Step 204, in response to determining that there is data change, update the inverted index file according to the time record file and the set of behavior record files.
[0078] In some embodiments, if it is determined in Step 203 that there is data change, the execution entity may update the inverted index file according to the time record file and the set of behavior record files.
[0079] Optionally, first, the execution entity may obtain the set of user identifiers for which the data has changed from the time record file. Then, for each user identifier in the set of user identifiers, the corresponding behavior record file can be selected from the set of behavior record files, such as the behavior record file with the file name being the user identifier. And according to the selected behavior record file, the updated data of the user indicated by the user identifier can be generated. For example, the execution entity may generate the updated data of the user according to the user identifier and at least one content identifier in the file. Here, the content identifier in the updated data may be all the content identifiers included in the file, or multiple content identifiers with the corresponding behavior time arranged in the front. After that, according to the generated updated data, the execution entity may update the inverted index file. The structure of the inverted index file may be as shown in the following table:
[0080]
[0081] In the inverted index file, the content identifiers associated with each user identifier are sorted in ascending order of the behavior time (timestamp). From the above description, it can be seen that the request merging can be completed through Steps 203 and 204. That is, multiple single requests can be merged into a batch request. This can reduce the update frequency of the inverted index file, reduce the pressure on the execution entity, and improve the processing efficiency.
[0082] In some embodiments, in order to further improve the update efficiency, the execution entity may obtain the set of user identifiers for which the data has changed through paged query. And take the user identifiers included in each page as a group to construct a batch insertion request for the inverted index file.
[0083] Some embodiments of the present disclosure provide a data update method that proposes a combination of Redis and inverted indexes to implement collaborative filtering technology in a recommendation system, thereby overcoming the problem of poor real-time performance of collaborative filtering technology. By dividing data updates into two stages: in the first stage, change records of user behavior data are implemented through a time record file and a behavior record file; in the second stage, the inverted index file is updated regularly. In the first stage, Redis can be used as a buffer layer, which can support high concurrency. And in the second stage, request merging can be performed based on the first stage. Also, due to the existence of the Redis buffer layer, the number of update requests in the second stage is often much smaller than that in the first stage. This not only ensures the real-time performance of the data but also reduces the data update frequency and improves the update efficiency.
[0084] Please refer to Figure 3 , which shows a flow 300 according to other embodiments of the data update method of the present disclosure. The method includes the following steps:
[0085] Step 301, obtain a list of content identifiers of a target user from an inverted index file.
[0086] In some embodiments, the execution entity of the data update method (such as Figure 1 the server 103 shown in
[0087] ) can obtain a list of content identifiers of a target user from the inverted index file (in the above embodiments). The target user here can be any user on the platform. For example, the execution entity can search for the list of content identifiers associated with the user identifier in the inverted index file according to the user identifier of the target user.
[0088] It should be noted that the above inverted index file can be an inverted index based on Elasticsearch. Elasticsearch is a search server. It provides a full-text search engine with distributed multi-user capabilities. In this way, content recommendations can be simulated through the aggregation API (Application Programming Interface) of Elasticsearch.
[0088] Step 302, use the content identifier at a preset position in the list of content identifiers as query information, and obtain a set of lists of content identifiers containing the query information from the inverted index file.
[0089] In some embodiments, based on the list of content identifiers obtained in step 301, the execution entity can use the content identifier at a preset position therein as query information. Among them, the preset position can be set according to actual needs. Based on the query information, the execution entity can obtain a set of lists of content identifiers containing the query information from the inverted index file.
[0090] For example Figure 4 As shown, the list of article IDs recently viewed by user A is [98, 68, 4, 7, 40, 9, 11, 78, 16, 51]. The article IDs in it are arranged in descending order of timestamp. At this time, the preset position can be the first four in the list (i.e., the first 40% of the number of article IDs). That is to say, the query information is the article ID list [98, 68, 4, 7]. According to the query information, other users who have viewed the same articles as user A can be found from the inverted index file. The lists of article IDs recently viewed by these users can form a set of article ID lists (i.e., a set of content identifier lists). For example, the list of article IDs recently viewed by user B is [54, 4, 22, 6, 68, 98, 11, 7, 13, 26], the list of article IDs recently viewed by user C is [68, 7, 54, 4, 26, 20, 98, 35, 9, 40], and the list of article IDs recently viewed by user D is [54, 98, 68, 4, 7].
[0091] It can be understood that for the above query process, the execution entity can construct a filter query condition. Compared with another query method (query query), this query method only filters data without performing score calculation, thereby further improving the query speed.
[0092] Step 303: Generate a recommended content identifier list for the target user based on the set of content identifier lists.
[0093] In some embodiments, based on the set of content identifier lists obtained in step 302, first, the execution entity can aggregate the content identifiers in each content identifier list to obtain a set of content identifier groups. Then, the content identifier groups that contain the content identifiers in the content identifier list of the target user can be removed from the set of content identifier groups, thereby obtaining a set of candidate content identifier groups.
[0094] Optionally, in order to reduce the amount of data to be processed and improve the processing efficiency of the execution entity, the first content identifiers included in each content identifier list in the set of content identifier lists can also be removed first. The first content identifier is the content identifier in the content identifier list of the target user. Then, the execution entity can aggregate the remaining content identifiers to obtain a set of candidate content identifier groups.
[0095] Further, according to the number of candidate content identifiers within a group, the execution entity can sort each candidate content identifier group in the candidate content identifier group set. A recommended content identifier list for the target user is generated based on the sorting result. For example, each candidate content identifier group is sorted in descending order of the number within the group. At this time, the execution entity can use the candidate content identifiers in each candidate content identifier group in sequence as the recommended content identifiers, thereby obtaining a recommended content identifier list. Or, the candidate content identifiers in several candidate content identifier groups with higher sorting positions can also be used in sequence as the recommended content identifiers in the recommended content identifier list.
[0096] It should be noted that for the case where the number within a group is the same, the execution entity can sort each candidate content identifier group in descending order (or ascending order, or randomly) of the candidate content identifiers. For the above aggregation process, the execution entity can construct a Terms Aggregation aggregation query. Such an aggregation query is generally based on a certain region, where each (unique token) within the region is a bucket, and the number of documents in each bucket is calculated. The default return order is sorted by the number of documents. Additionally, to prevent too much data from being returned, the number of buckets to be returned can also be defined, or buckets with the number of documents not less than a certain value can be returned.
[0097] For another example Figure 4 As shown, based on the list of article IDs recently viewed by users B, C, and D, the articles that these three users have viewed and user A has not viewed can be obtained. At this time, the set of candidate article ID groups obtained is [54, 54, 54], [26, 26],
[22] , [6],
[13] ,
[20] ,
[35] . Sorted in descending order of the viewing quantity, finally, the article ID list [54, 26, 35] that user A is most likely to be interested in (i.e., recommended) can be obtained.
[0098] The data update method provided by some embodiments of the present disclosure can also implement collaborative filtering technology, that is, it can recommend content for users. At the same time, since the inverted index file only records key data and does not need to load all data into memory for calculation. Therefore, the disadvantage of high computational cost of collaborative filtering technology can be overcome, and the computational cost can be effectively reduced.
[0099] Further referring to Figure 5 as an implementation of the methods shown above Figure 2 and 3 the present disclosure provides some embodiments of a data update device. These device embodiments correspond to those method embodiments shown in Figure 2 and 3 and this device can be specifically applied to various electronic devices.
[0100] Such as Figure 5As shown, the data update device 500 in some embodiments may include: a processing unit 501 configured to process the acquired user behavior logs to obtain target data; a change unit 502 configured to change the data in the relevant behavior record files in the time record file and the set of behavior record files according to the target data, where the time record file stores the user data of each behavior occurring within a preset time period, and each behavior record file in the set of behavior record files stores the behavior data of different users; a change determination unit 503 configured to determine whether there is data change within the preset time period according to the time record file in response to determining that the current time reaches the update time; and an update unit 504 configured to update the inverted index file according to the time record file and the set of behavior record files in response to determining that there is data change.
[0101] In some embodiments, the target data includes a user identifier, a content identifier of the behavior content, and a behavior time; and the file name of the behavior record file in the set of behavior record files is the user identifier, and the file content includes the content identifier of the behavior content and the behavior time; the file name of the time record file is the preset time period, and the file content includes the user identifier.
[0102] In some embodiments, the update unit 504 is further configured to obtain a set of user identifiers with data change from the time record file; for each user identifier in the set of user identifiers, select the behavior record file corresponding to the user identifier from the set of behavior record files, generate update data of the user indicated by the user identifier according to the selected behavior record file; and update the inverted index file according to the generated update data.
[0103] In some embodiments, the update unit 504 is further configured to generate update data of the user indicated by the user identifier according to the user identifier and at least one content identifier in the selected behavior record file.
[0104] In some embodiments, the change unit 502 includes a first change subunit ( Figure 5 not shown in the figure) configured to: determine whether there is a target behavior record file in the set of behavior record files, where the file name of the target behavior record file is the user identifier in the target data; in response to determining that there is, determine whether the target content identifier is included in the target behavior record file, where the target content identifier is the content identifier in the target data; and in response to determining that it is included, change the behavior time corresponding to the target content identifier to the behavior time in the target data.
[0105] In some embodiments, the first change subunit is further configured to: in response to determining that it is not included, add the content identifier and the behavior time in the target data to the target behavior record file; determine whether the number of records in the target behavior record file reaches a preset number; in response to determining that it reaches, delete the record with the earliest behavior time in the target behavior record file.
[0106] In some embodiments, the change unit 502 further includes a second change subunit ( Figure 5 not shown in the figure), configured to: determine whether the behavior time in the target data is within a preset time period; in response to determining that it is within, determine whether the user identifier in the target data is included in the time record file; in response to determining that it is not included, add the user identifier in the target data to the time record file.
[0107] In some embodiments, the apparatus 500 further includes: an acquisition unit ( Figure 5 not shown in the figure), configured to acquire a list of content identifiers of a target user from an inverted index file; a query unit ( Figure 5 not shown in the figure), configured to use the content identifier at a preset position in the content identifier list as query information, and acquire a set of content identifier lists including the query information from the inverted index file; a generation unit ( Figure 5 not shown in the figure), configured to generate a recommended content identifier list of the target user based on the set of content identifier lists.
[0108] In some embodiments, the generation unit may further include: an aggregation subunit, configured to aggregate each content identifier in each content identifier list in the set of content identifier lists to obtain a set of content identifier groups; a removal subunit, configured to remove the content identifier groups that include the content identifiers in the content identifier lists from the set of content identifier groups to obtain a set of candidate content identifier groups; a generation subunit, configured to sort each candidate content identifier group in the set of candidate content identifier groups according to the number of candidate content identifiers in the group, and generate a recommended content identifier list of the target user according to the sorting result.
[0109] In some embodiments, the processing unit 501 is further configured to determine whether the acquired user behavior log includes data characterizing a target behavior; in response to determining that it includes, extract key data from the user behavior log to generate target data.
[0110] It can be understood that the various units described in the apparatus 500 correspond to the respective steps in the method described with reference to Figure 2 and 3 Therefore, the operations, features, and beneficial effects described above for the method also apply to the apparatus 500 and the units included therein, and will not be repeated here.
[0111] Next, with reference toFigure 6 , which shows a schematic structural diagram of an electronic device (such as a server in Figure 1 ) 600 suitable for implementing some embodiments of the present disclosure. Figure 6 The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0112] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0113] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had. Figure 6 Each block shown in
[0114] may represent a device or may represent multiple devices as needed.
[0115] It should be noted that the computer-readable media described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0116] In some embodiments, the client and the server may communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0117] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: process the obtained user behavior log to obtain target data; according to the target data, change the data in the relevant behavior record files in the time record file and the set of behavior record files, where the time record file stores the data of each user whose behavior occurs within a preset time period, and each behavior record file in the set of behavior record files stores the behavior data of different users; in response to determining that the current time reaches the update time, determine whether there is data change within the preset time period according to the time record file; in response to determining that there is, update the inverted index file according to the time record file and the set of behavior record files.
[0118] In addition, computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0120] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a processing unit, a change unit, a change determination unit, and an update unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the processing unit can also be described as "the unit for processing the obtained user behavior logs".
[0121] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0122] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A data update method, wherein, The method includes: Processing the obtained user behavior log to obtain target data, where the target data includes a user identifier, a content identifier of the behavior content, and a behavior time; According to the target data, changing the user data in the time record file and the behavior data in the relevant behavior record files in the behavior record file set, where the time record file stores the user identifiers of each user whose behavior occurs within a preset time period, and each behavior record file in the behavior record file set stores the behavior data of different users, where the behavior data includes a content identifier of the behavior content and a behavior time; In response to determining that the current time reaches the update time, determining whether there is data change within the preset time period according to the time record file; In response to determining that there is, obtaining a set of user identifiers whose data has changed from the time record file, for each user identifier in the set of user identifiers, selecting the behavior record file corresponding to the user identifier from the behavior record file set, generating update data for the user indicated by the user identifier according to the selected behavior record file, and updating the inverted index file according to the generated update data.
2. The data update method according to claim 1, wherein, The file name of the behavior record file in the behavior record file set is the user identifier, and the file content includes a content identifier of the behavior content and a behavior time; The file name of the time record file is the preset time period, and the file content includes user identifiers.
3. The data update method according to claim 1, wherein, The generating the update data for the user indicated by the user identifier according to the selected behavior record file includes: Generating the update data for the user indicated by the user identifier according to the user identifier and at least one content identifier in the selected behavior record file.
4. The data update method according to claim 2, wherein, The changing the user data in the time record file and the behavior data in the relevant behavior record files in the behavior record file set according to the target data includes: Determining whether there is a target behavior record file in the behavior record file set, where the file name of the target behavior record file is the user identifier in the target data; In response to determining that there is, determining whether the target behavior record file contains a target content identifier, where the target content identifier is the content identifier in the target data; In response to determining that it contains, changing the behavior time corresponding to the target content identifier to the behavior time in the target data.
5. The data update method according to claim 4, wherein, The changing the user data in the time record file and the behavior data in the relevant behavior record files in the behavior record file set according to the target data further includes: In response to determining that it does not contain, adding the content identifier and behavior time in the target data to the target behavior record file; Determining whether the number of records in the target behavior record file reaches a preset number; In response to determining that it reaches, deleting the record with the earliest behavior time in the target behavior record file.
6. The data update method according to claim 3, wherein, The changing the user data in the time record file and the behavior data in the relevant behavior record files in the behavior record file set according to the target data further includes: Determine whether the behavior time in the target data is within a preset time period; In response to determining that it is within, determine whether the user identifier in the target data is included in the time record file; In response to determining that it is not included, add the user identifier in the target data to the time record file.
7. The data update method according to claim 1, wherein, The method further includes: Obtain a list of content identifiers of the target user from the inverted index file; Use the content identifier at a preset position in the list of content identifiers as query information, and obtain a set of lists of content identifiers containing the query information from the inverted index file; Generate a recommended list of content identifiers for the target user based on the set of lists of content identifiers.
8. The data update method according to claim 7, wherein The generating a recommended list of content identifiers for the target user based on the set of lists of content identifiers includes: Aggregate each content identifier in each list of content identifiers in the set of lists of content identifiers to obtain a set of content identifier groups; Remove the content identifier groups in the set of content identifier groups that contain the content identifiers in the list of content identifiers to obtain a set of candidate content identifier groups; Sort each candidate content identifier group in the set of candidate content identifier groups according to the number of candidate content identifiers within the group, and generate a recommended list of content identifiers for the target user according to the sorting result.
9. The data update method according to any one of claims 1-8, wherein, The processing the obtained user behavior log to obtain target data includes: Determine whether the obtained user behavior log includes data representing the target behavior; In response to determining that it includes, extract the key data in the user behavior log to generate target data.
10. A data update device, wherein, The apparatus includes: A processing unit configured to process the obtained user behavior log to obtain target data, where the target data includes a user identifier, a content identifier of the behavior content, and a behavior time; A change unit configured to change the user data in the time record file and the behavior data in the relevant behavior record files in the set of behavior record files according to the target data, where the time record file stores the user identifiers of each user whose behavior occurs within a preset time period, and each behavior record file in the set of behavior record files stores the behavior data of different users, where the behavior data includes a content identifier of the behavior content and a behavior time; A change determination unit configured to, in response to determining that the current time reaches the update time, determine whether there is data change within the preset time period according to the time record file; An update unit configured to, in response to determining that there is, obtain a set of user identifiers whose data has changed from the time record file, for each user identifier in the set of user identifiers, select the behavior record file corresponding to the user identifier from the set of behavior record files, generate update data for the user indicated by the user identifier according to the selected behavior record file, and update the inverted index file according to the generated update data.
11. An electronic device, including: One or more processors; A storage device on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-9.
12. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by a processor, the method according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Personalized service recommendation system and method
CN101334792A
User data processing method and device, medium and equipment
CN112657198A