A data acquisition method and apparatus
By combining dimensional indexing and time-series indexing, consumer data can be quickly identified and obtained, solving the problem of low data acquisition efficiency in existing technologies and achieving efficient consumer data acquisition.
Patent Information
- Application Number
- CN202110062707.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-18
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2041-01-18
AI Technical Summary
Existing methods for acquiring consumer data require querying and aggregating all real-time consumer data, resulting in low data acquisition efficiency.
By quickly identifying content identifiers and target time-series indexes using dimensional index information and time-series index information, consumption data can be directly obtained from the target time-series index, and multi-granularity pre-aggregation and storage can be performed in advance.
This significantly reduces the time required to acquire target consumer data and improves data acquisition efficiency.
Smart Images

Figure CN114817344B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, and more specifically to a data acquisition method and apparatus. Background Technology
[0002] In recent years, with the rapid development of content recommendation methods, more and more content is being accurately recommended. Users consume this recommended content through their devices, generating real-time consumption data. Analyzing the content quality or recommendation effectiveness requires obtaining consumption data for a specific consumption time from the real-time consumption data. Existing methods for obtaining consumption data involve calculating all consumption data within the specified time period from a database storing real-time consumption data to obtain the target consumption data.
[0003] In the process of researching and practicing existing technologies, the inventors of this invention discovered that for the method of acquiring consumer data, it is necessary not only to query all consumer data within the consumption time period from the full amount of real-time consumer data, but also to aggregate these consumer data in order to obtain the target consumer data, which greatly increases the acquisition time of the target consumer data, thus leading to a significant reduction in the efficiency of data acquisition. Summary of the Invention
[0004] This invention provides a data acquisition method and apparatus that can improve the efficiency of data acquisition.
[0005] A data acquisition method, comprising:
[0006] The receiving terminal sends a consumption data acquisition request for recommended content, the consumption data acquisition request carrying the attribute information and consumption time of the recommended content;
[0007] Based on the attribute information, at least one content identifier corresponding to the recommended content is identified in the dimension index information, wherein the dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content;
[0008] Based on the content identifier, the target time index corresponding to the recommended content is determined in the time index information. The time index information is used to indicate the association between the content identifier of the recommended content and the time index. The time index includes consumption data corresponding to at least one preset consumption time unit.
[0009] The target consumption data of the recommended content within the consumption time period is obtained from the target time series index, and the target consumption data is returned to the terminal.
[0010] Accordingly, embodiments of the present invention provide a data acquisition device, including:
[0011] A receiving unit is used to receive a consumption data acquisition request for recommended content sent by a terminal, wherein the consumption data acquisition request carries the attribute information and consumption time of the recommended content;
[0012] The identification unit is used to identify at least one content identifier corresponding to the recommended content in the dimension index information based on the attribute information, wherein the dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content;
[0013] The determining unit is used to determine the target time-series index corresponding to the recommended content in the time-series index information based on the content identifier. The time-series index information is used to indicate the association between the content identifier of the recommended content and the time-series index. The time-series index includes consumption data corresponding to at least one preset consumption time unit.
[0014] The acquisition unit is used to acquire the target consumption data of the recommended content within the consumption time from the target time-series index, and return the target consumption data to the terminal.
[0015] Optionally, in some embodiments, the data acquisition device may further include a construction unit, which may be used to acquire real-time content data of at least one recommended content, the real-time content data including basic consumption data and basic content data; construct time-series index information of the recommended content based on the basic consumption data; and construct dimensional index information of the recommended content based on the basic content data.
[0016] Optionally, in some embodiments, the construction unit may specifically be used to acquire real-time consumption data of at least one recommended content and classify the real-time consumption data; perform multi-row to column transformation on the classified real-time consumption data to obtain processed real-time consumption data; filter at least one real-time consumption data corresponding to a preset time window from the processed real-time consumption data to obtain a window real-time consumption data set; aggregate the real-time consumption data in the window real-time consumption data set to obtain the real-time basic consumption data of the recommended content; associate the content attribute data of the recommended content with the real-time basic consumption data, and encode the associated data according to a preset encoding strategy to obtain the real-time content data of the recommended content.
[0017] Optionally, in some embodiments, the construction unit may be specifically used to synchronize the content attribute data of the recommended content in the content database to the content cache database; obtain the content identifier of the recommended content, and filter out the content attribute data corresponding to the content identifier in the content cache database to obtain the content attribute data of the recommended content.
[0018] Optionally, in some embodiments, the construction unit may be specifically used to obtain preset time-series information for constructing time-series index information, the preset time-series information including at least one consumption time unit; aggregate the basic consumption data according to the consumption time unit to obtain consumption data corresponding to the consumption time unit; and construct the time-series index information of the recommended content based on the consumption data corresponding to the consumption time unit.
[0019] Optionally, in some embodiments, the construction unit may be specifically used to filter out at least one basic consumption data corresponding to the consumption time unit from the basic consumption data to obtain target basic consumption data; aggregate the target basic consumption data to obtain the data aggregation value within the consumption time unit; and compress the data aggregation value to obtain the consumption data corresponding to the time unit.
[0020] Optionally, in some embodiments, the construction unit may be specifically used to determine the storage area of the data aggregation value of the recommended content within the consumption time unit based on the content identifier of the recommended content; and store the data aggregation value in the storage area.
[0021] Optionally, in some embodiments, the construction unit may be specifically used to construct an initial time-series index corresponding to the content identifier based on the preset time-series information and the content identifier of the recommended content, the initial time-series index including index information corresponding to the consumption time unit; add the consumption data corresponding to the consumption time unit to the index information to obtain the time-series index corresponding to the content identifier; and fuse the time-series index corresponding to the content identifier to obtain the time-series index information of the recommended content.
[0022] Optionally, in some embodiments, the construction unit may be specifically used to: filter out the content dimension corresponding to the recommended content from a preset dimension set based on the basic content data; filter out target recommended content with the same content dimension from the recommended content; and associate the content identifier of the target recommended content with the content dimension to obtain the dimension index information of the recommended content.
[0023] Optionally, in some embodiments, the identification unit may be specifically used to identify at least one content identifier corresponding to the recommended content in the content identifier information when the attribute information contains the content identifier information of the recommended content; and when the attribute information does not contain the content identifier information of the recommended content, to filter out at least one content identifier corresponding to the recommended content in the dimension index information according to the attribute information.
[0024] Optionally, in some embodiments, the identification unit may be specifically used to determine the target content dimension of the recommended content based on the attribute information; filter out the content identifier information corresponding to the target content dimension from the dimension index information; and identify at least one content identifier of the recommended content from the content identifier information.
[0025] Optionally, in some embodiments, the acquisition unit may be specifically used to acquire the target consumption data of the recommended content within the consumption time from the index information corresponding to the candidate consumption time when the candidate consumption time is the same as the consumption time; and to aggregate the consumption data in the index information according to the consumption time to obtain the target consumption data of the recommended content within the consumption time when the candidate consumption time is different from the consumption time.
[0026] Optionally, in some embodiments, the acquisition unit may be specifically used to filter out at least one target candidate consumption time to form the consumption time from the candidate consumption time; filter out the index information corresponding to the target candidate consumption time from the target time series index to obtain target index information; and aggregate the consumption data in the target index information to obtain the target consumption data of the recommended content within the consumption time.
[0027] Optionally, in some embodiments, the acquisition unit may be specifically used to acquire historical consumption data of the recommended content within a preset consumption time; update the historical consumption data according to the real-time consumption data of the recommended content; calculate the data error between the updated historical consumption data and the target consumption data; and return the target consumption data to the terminal, including: returning the target consumption data to the terminal when the data error does not exceed a preset error threshold.
[0028] Furthermore, embodiments of the present invention also provide an electronic device, including a processor and a memory, wherein the memory stores an application program, and the processor is used to run the application program in the memory to implement the data acquisition method provided in embodiments of the present invention.
[0029] Furthermore, embodiments of the present invention also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the data acquisition methods provided in embodiments of the present invention.
[0030] In this embodiment of the invention, after receiving a request from a terminal to acquire consumption data of recommended content, the request carries attribute information and consumption time of the recommended content. Then, based on the attribute information, at least one content identifier corresponding to the recommended content is identified in the dimension index information. This dimension index information indicates the association between the content dimension and the content identifier of the recommended content. Next, based on the content identifier, a target time-series index corresponding to the recommended content is determined in the time-series index information. This time-series index indicates the association between the content identifier of the recommended content and the time-series index. The time-series index includes consumption data corresponding to at least one preset consumption time unit. The target consumption data of the recommended content within the consumption time is acquired from the target time-series index and returned to the terminal. Because this scheme can quickly acquire the target time-series index of the recommended content through dimension index information and time-series index information, and performs multi-granularity pre-aggregation of real-time consumption data and stores the pre-aggregated data in the target time-series index, the target consumption data can be directly filtered from the index information of the target time-series index, greatly reducing the acquisition time of the target consumption data. Therefore, the efficiency of data acquisition can be improved. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is a schematic diagram of a scenario for the data acquisition method provided in an embodiment of the present invention;
[0033] Figure 2 This is a flowchart illustrating the data acquisition method provided in an embodiment of the present invention;
[0034] Figure 3 This is a schematic diagram illustrating the splitting of real-time consumption data according to an embodiment of the present invention;
[0035] Figure 4 This is a schematic diagram of the process for real-time calculation of real-time consumption data provided in an embodiment of the present invention;
[0036] Figure 5 This is a schematic diagram of the storage architecture for consumer data provided in an embodiment of the present invention;
[0037] Figure 6 This is a schematic diagram of the time-series index provided in an embodiment of the present invention;
[0038] Figure 7 This is a schematic diagram of the dimensional index information provided in an embodiment of the present invention;
[0039] Figure 8 This is a system architecture diagram of the data acquisition device provided in an embodiment of the present invention;
[0040] Figure 9 This is another schematic diagram of the data acquisition process provided in an embodiment of the present invention;
[0041] Figure 10 This is a schematic diagram illustrating an application scenario of the data acquisition device provided in an embodiment of the present invention;
[0042] Figure 11 This is a schematic diagram of the structure of the data acquisition device provided in an embodiment of the present invention;
[0043] Figure 12 This is another structural schematic diagram of the data acquisition device provided in an embodiment of the present invention;
[0044] Figure 13 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0046] This invention provides a data acquisition method, apparatus, and computer-readable storage medium. The data acquisition apparatus can be integrated into an electronic device, which may be a server or a terminal, etc.
[0047] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.
[0048] For example, see Figure 1Taking the integration of a data acquisition device into an electronic device as an example, after receiving a consumption data acquisition request for recommended content sent by a terminal, the electronic device carries the attribute information and consumption time of the recommended content. Then, based on the attribute information, at least one content identifier corresponding to the recommended content is identified in the dimension index information. This dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content. Then, based on the content identifier, the target time-series index corresponding to the recommended content is determined in the time-series index information. This time-series index information is used to indicate the association between the content identifier of the recommended content and the time-series index. This time-series index includes consumption data corresponding to at least one preset consumption time unit. The target consumption data of the recommended content within the consumption time is obtained from the target time-series index, and the target consumption data is returned to the terminal.
[0049] Among them, consumption data can be data generated from consumption behavior of recommended content, such as data generated from consumption behavior of exposure, click, sharing, liking or commenting on recommended content. For example, it can be data such as the number of clicks, click-through rate, number of exposures or exposure rate within a consumption period.
[0050] Consumption time can be understood as the time it takes to complete a consumption behavior. When a consumption behavior is performed in real time, the consumption time at that time can be understood as the time of the current consumption behavior.
[0051] The time-series index information, dimensional index information, and target consumption data can be stored on a cloud platform. A cloud platform, also known as a cloud computing platform, refers to a service that provides computing, networking, and storage capabilities based on hardware and software resources. Cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, the resources in the "cloud" are infinitely scalable, readily available, on-demand, expandable, and pay-as-you-go.
[0052] As a provider of fundamental cloud computing capabilities, a cloud resource pool (referred to as a cloud platform, generally called an IaaS (Infrastructure as a Service) platform) is established. Various types of virtual resources are deployed in the resource pool for external customers to choose from. The cloud resource pool mainly includes: computing devices (virtualized machines containing operating systems), storage devices, and network devices.
[0053] Based on logical function, a PaaS (Platform as a Service) layer can be deployed on top of the IaaS (Infrastructure as a Service) layer, and a SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. Alternatively, SaaS can be deployed directly on top of IaaS. PaaS is a platform for running software, such as databases and web containers. SaaS refers to various types of business software, such as web portals and bulk SMS senders. Generally speaking, SaaS and PaaS are upper layers compared to IaaS.
[0054] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.
[0055] This embodiment will be described from the perspective of a data acquisition device, which can be integrated into an electronic device, such as a server or a terminal. The terminal can include tablet computers, laptops, personal computers (PCs), wearable devices, virtual reality devices, or other smart devices that can acquire data.
[0056] A data acquisition method, comprising:
[0057] The receiving terminal sends a consumption data acquisition request for recommended content. This request carries the attribute information and consumption time of the recommended content. Based on the attribute information, at least one content identifier corresponding to the recommended content is identified in the dimension index information. This dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content. Based on the content identifier, the target time-series index corresponding to the recommended content is determined in the time-series index information. This time-series index information is used to indicate the association between the content identifier of the recommended content and the time-series index. The time-series index includes consumption data corresponding to at least one preset consumption time unit. The target consumption data of the recommended content within the consumption time is obtained from the target time-series index, and the target consumption data is returned to the terminal.
[0058] like Figure 2 As shown, the specific process of this data acquisition method is as follows:
[0059] 101. Receive the consumption data acquisition request sent by the terminal.
[0060] The consumption data retrieval request carries the attribute information of the recommended content and the consumption time. Consumption time can be understood as the time it takes to complete one or more consumption behaviors, such as the time it takes to complete an exposure, click, share, or other consumption behavior. The consumption data retrieval request is used to obtain consumption data of the recommended content within the consumption time. Consumption data can be data on the consumption behaviors completed by users within the consumption time, such as data on exposure, likes, shares, or clicks completed within a specific time point or time interval.
[0061] For example, a device can directly receive consumption data retrieval requests sent by a terminal. For instance, a user triggers a consumption data retrieval request on the terminal to generate recommended content, adding attribute information and consumption time to the request. The terminal then sends this request to the data acquisition device, allowing the device to receive the request. When the memory required for the recommended content's attribute information and the consumption data is large, the storage address of the recommended content's attribute information and the consumption data can also be added to the consumption data retrieval request. The terminal then sends this request to the data acquisition device. Upon receiving the request, the data acquisition device extracts the storage address and retrieves the recommended content's attribute information and consumption time from the terminal's memory or cache based on the storage address.
[0062] Optionally, before receiving a consumption data acquisition request, the terminal can also construct time-series index information and dimensional index information of the recommended content based on the real-time consumption data of the recommended content. Therefore, the data acquisition method may also include:
[0063] Obtain real-time content data for at least one recommended content item. This real-time content data includes basic consumption data and basic content data. Based on the basic consumption data, construct the time-series index information of the recommended content. Based on the basic content data, construct the dimensional index information of the recommended content. Specifically, it can be as follows.
[0064] S1. Obtain real-time content data for at least one recommended item.
[0065] Real-time content data can include basic consumption data and basic content data. Basic consumption data can be understood as data aggregated from real-time consumption data within a fixed time window. For example, with a 1-minute time window, basic consumption data would be the consumption data of recommended content within each 1-minute period. Basic content data can include the title, content, or source of the recommended content. Real-time content data from multiple recommended content sets can then be used to form a message queue.
[0066] For example, real-time consumption data for at least one recommended item can be obtained, and the real-time consumption data can be categorized. Based on a preset time window, the categorized real-time consumption data can be aggregated to obtain the real-time basic consumption data of the recommended item. The content attribute data of the recommended item can then be associated with the real-time basic consumption data, and the associated data can be encoded according to a preset encoding strategy to obtain the real-time content data of the recommended item. Specifically, this can be done as follows:
[0067] (1) Obtain real-time consumption data of at least one recommended content and classify the real-time consumption data.
[0068] Real-time consumption data refers to the real-time data generated by users' actions of consuming recommended content on the terminal. For example, when a user likes a piece of recommended content on the terminal, consumption data for that recommended content is generated, and this consumption data can be considered real-time consumption data.
[0069] For example, one can obtain consumption data for at least one recommended content sent in real time by a content consumption platform, thus obtaining real-time consumption data for the recommended content. For instance, when a user consumes recommended content on a content consumption platform such as a content viewing platform, content browser, or content application, the platform sends the consumption data generated by the consumption behavior to a data acquisition device in real time, allowing the device to obtain real-time consumption data for at least one recommended content. This real-time consumption data can be categorized, for example, by business type. First, it can be categorized by the source of the real-time consumption data, and then further categorized by consumption behavior. Based on the categorization results, a real-time data warehouse can be built to store the categorized real-time consumption data. This allows the massive amount of real-time consumption data to be broken down into smaller message queues, such as... Figure 3 As shown, by classifying and splitting large queues of real-time consumption data, the burden on downstream consumers to obtain consumption data is reduced.
[0070] (2) Based on the preset time window, the real-time consumption data after classification is aggregated to obtain the real-time basic consumption data of the recommended content.
[0071] Real-time basic consumption data can be understood as consumption data aggregated from real-time consumption data within a preset time window. For example, if the preset time window is 1 minute, then the real-time basic consumption data can be the consumption data for 1 minute starting from the current time point.
[0072] For example, multi-row transpilation is performed on categorized real-time consumption data to obtain processed real-time consumption data. For instance, data decoding is performed on the message queue composed of categorized real-time consumption data to obtain the original data. Multi-row transpilation of the obtained original data yields the processed real-time consumption data. From the processed real-time consumption data, at least one real-time consumption data corresponding to a preset time window is selected to obtain windowed real-time consumption data. For example, with a preset time window of 1 minute, all real-time consumption data within 1 minute starting from the current time can be selected from the processed real-time consumption data, thus obtaining the windowed consumption data set. The real-time consumption data in the windowed real-time consumption data set is aggregated to obtain the basic real-time consumption data for the recommended content. For example, taking the consumption data of 5 clicks on a certain recommended content within 1 minute as an example, window aggregation is performed on the consumption data corresponding to these 5 clicks to obtain the basic real-time consumption data for the recommended content as 5 clicks on the recommended content within the current 1 minute.
[0073] The preset time window can be any time value, but the size of the preset time window is related to the aggregation rate of real-time consumption data within that time window. It needs to be set according to the order of magnitude of the real-time consumption data. When the preset time window is 1 minute, the aggregation speed of real-time consumption data can be optimized to 3 minutes.
[0074] (3) Associate the content attribute data of the recommended content with the real-time basic consumption data, and encode the associated data according to the preset encoding strategy to obtain the real-time content data of the recommended content.
[0075] For example, content attribute data of recommended content is retrieved from the content cache database and associated with real-time basic consumption data. This includes retrieving content attribute data such as the title, main content, source, and type of the recommended content from the content cache database and associating this data with the real-time basic consumption data. According to a preset encoding strategy, the associated data is encoded to obtain the real-time content data of the recommended content. This real-time content data can then be used to form a consumption queue.
[0076] Optionally, before associating real-time basic consumption data with content attribute data, it is also necessary to obtain the content attribute data of recommended content from the content cache database. The content attribute data of recommended content in the content cache database originates from data synchronization with the content database. Therefore, the data acquisition method may also include:
[0077] The content attribute data of recommended content in the content database is synchronized to the content cache database. The content identifier of the recommended content is obtained, and the content attribute data corresponding to the content identifier is filtered out in the content cache database to obtain the content attribute data of the recommended content.
[0078] For example, content attribute data for all recommended content in a distributed content database can be synchronized to a content cache database via logs. Taking HBase (an open-source distributed database) as an example, the HBase primary database synchronizes its content attribute data to the HBase secondary database via binary logs (binlog). The secondary database then synchronizes this data to the content cache database, which can be a Redis cache database. Data synchronization between the HBase primary, secondary, and Redis databases is primarily achieved through HBase-proxy (a database proxy service) writing logs and sending them to SyncSvr (a server synchronization service), ensuring cache consistency between the content cache database and the HBase primary and secondary databases. This ensures that when retrieving content attribute data for recommended content, it is retrieved directly from the Redis content cache database, minimizing access to HBase. Furthermore, Redis is a thousand times faster than HBase, significantly improving the speed of content attribute data retrieval.
[0079] Retrieve the content identifier of the recommended content, and then filter out the content attribute data corresponding to that identifier from the content cache database. For example, to retrieve the content identifier of the recommended content, the content identifier can be a unique identifier (Roekey) for that content. Each piece of content has a unique rowkey. Then, filter out the content attribute data corresponding to that rowkey from the Redis content cache database.
[0080] The process of assembling real-time consumed data into a message queue to obtain real-time content data can be viewed as a real-time computation of the consumed data. This real-time computation can be performed using Flink (a distributed streaming data engine), and the methods for real-time computation can include... Figure 4 As shown.
[0081] S2. Based on basic consumption data, construct a time-series index of recommended content.
[0082] The time-series index information indicates the relationship between the content identifier of the recommended content and the time-series index. For example, by querying the content identifier of the recommended content, the time-series index of the recommended content can be quickly obtained. A time-series index can be understood as a data structure that stores the consumption data of the recommended content across multiple consumption time units. A time-series index can contain multiple index information, with each time unit corresponding to one index information. The data structure corresponding to each index information records the consumption data of the recommended content within that time unit. Therefore, it can be seen that by using the content identifier of the recommended content, the corresponding time-series index can be found in the time-series index information. Based on the consumption time, the index information corresponding to that consumption time can also be found in the time-series index. By reading the consumption data recorded in the data structure corresponding to the index information, the target consumption data of the recommended content within the consumption time can be obtained.
[0083] For example, preset time-series information for constructing time-series index information can be obtained. This preset time-series information includes at least one consumption time unit. Based on the consumption time unit, the basic consumption data is aggregated to obtain the consumption data corresponding to the consumption time unit. Based on the consumption data corresponding to the consumption time unit, the time-series index information of the recommended content is constructed, as follows:
[0084] (1) Obtain the preset time sequence information for constructing the time sequence index.
[0085] The preset time series information includes at least one consumption time unit. The preset time series information can be the consumption time unit information contained in the constructed time series index information. For example, if the constructed time series index information contains days, hours, and 10 minutes, then the preset time series information can be days, hours, and 10 minutes, and the consumption time unit can be one of these three consumption time units: days, hours, and 10 minutes.
[0086] For example, configuration information of time-series index information can be directly obtained, preset time-series information used to build time-series index information can be filtered out from the configuration information, and at least one consumption time unit contained in the preset time-series information can be read.
[0087] (2) Aggregate the basic consumption data according to the consumption time unit to obtain the consumption data corresponding to the consumption time unit.
[0088] For example, by filtering at least one basic consumption data point corresponding to a consumption time unit from the basic consumption data, the target basic consumption data can be obtained. For instance, when the consumption time unit is a day, the basic consumption data within the 24 hours starting from the current time can be filtered out to obtain the target basic consumption data. Similarly, when the consumption time unit is an hour, the basic consumption data for each hour starting from the current time can be filtered out to obtain the target basic consumption data. The target basic consumption data is then aggregated to obtain the aggregated data value corresponding to the consumption time unit. For example, the aggregation method is the same as the window aggregation method described above, aggregating the target basic consumption data corresponding to each consumption time unit to obtain the aggregated data value for each consumption time unit. Finally, the aggregated data value is compressed to obtain the consumption data corresponding to the consumption time unit. For example, PB compression (a data compression method) can be used to compress the aggregated data value corresponding to the consumption time unit to obtain the consumption data for each consumption time unit.
[0089] Optionally, the aggregated data values of the recommended content within the consumption time unit need to be stored. The data acquisition methods may also include:
[0090] Based on the content identifier of the recommended content, determine the storage area for the data aggregation value of the recommended content within the consumption time unit, and store the data aggregation value in the storage area.
[0091] For example, the content identifier of the recommended content can be hashed, and the storage area for storing the aggregated data value can be determined based on the hash calculation result. For instance, a consistent hash calculation can be performed on the rowkey of the recommended content to obtain the data aggregated value to be stored in the corresponding shard (partition). The data aggregated value of the recommended content in each consumption time unit is stored in the corresponding shard.
[0092] The storage of consumption data can be handled by Zookeeper (a distributed application coordination service), and the specific storage architecture can be as follows: Figure 5 As shown, each shard can have both primary and backup storage, and the Zookeeper selection mechanism ensures high availability of the system. This storage architecture can interface with write and consumer services. The write service primarily writes aggregated data values to the corresponding shards using consistent hashing. The consumer service, which can be understood as a read service, is responsible for connecting to upstream interfaces and retrieving the necessary data from index information or the storage layer. For data storage, storage areas within the ClickHouse database can also be used.
[0093] (3) Based on the consumption data corresponding to the consumption time unit, construct the time-series index information of the recommended content.
[0094] For example, based on preset time-series information and the content identifiers of recommended content, an initial time-series index corresponding to the content identifier is constructed. This initial time-series index includes index information for consumption time units. For instance, an initial time-series index corresponding to a rowkey is created for each recommended content. This initial time-series index can include index information corresponding to consumption time units in the preset time-series information. For example, when the consumption time units are days, hours, and 10 minutes, the constructed initial time-series index can include index information corresponding to one day, one hour, and one 10-minute intervals from the current time. Consumption data corresponding to the consumption time units is added to the index information to obtain the time-series index corresponding to the content identifier. For example, adding consumption data corresponding to 10 minutes to the index information corresponding to 10 minutes ensures that the index information contains the consumption data corresponding to 10 minutes. The adding process can be understood as recording the consumption data in the data structure corresponding to the index. By adding the consumption data to the corresponding index information, the time-series index of the recommended content corresponding to the content identifier can be obtained. By fusing the time-series indices corresponding to content identifiers, the time-series index information of the recommended content can be obtained. For example, by fusing the time-series indices corresponding to each recommended content, the time-series index information of the recommended content can be obtained. The time-series index corresponding to the content identifier can be determined from the time-series index information. For instance, taking consumption time units as days, hours, and 10 minutes, and the content identifiers of the recommended content as key0 and key1, the time-series indices corresponding to key0 and key1 can be obtained as follows: Figure 6 As shown.
[0095] S3. Based on the basic content data, construct the dimensional index information of the recommended content.
[0096] The dimension index information indicates the relationship between content dimensions and the content identifiers of recommended content. Under a given content dimension, there can be a rowkey list consisting of one or more content identifiers. The function of this dimension index information is to query or retrieve the content identifiers of recommended content corresponding to that dimension using the recommended content's dimension information.
[0097] For example, based on basic content data, the content dimensions corresponding to the recommended content are filtered from a preset set of dimensions. This can be achieved by creating a dimension dictionary and its corresponding inverted list. The dimension dictionary can include a preset set of multiple dimensions, which can include the content category of the recommended content (e.g., sports or business) or the content source (e.g., WeChat, a browser, a content consumption platform). Each content dimension corresponds to an inverted list. Based on the basic content data, the dimension information of the recommended content is determined. By filtering the content dimensions corresponding to this information from the preset set of dimensions, the content dimensions of the recommended content can be determined. Finally, target recommended content with the same content dimensions can be filtered out from the recommended content. For example, recommended content belonging to the same type can be filtered out, and this content with the same content dimensions can be used as the target recommended content. By associating content identifiers within the target recommendation with content dimensions, we can obtain the dimensional index information of the recommended content. For example, we can add the content identifiers of the target recommended content to the inverted list corresponding to that content dimension, thus associating the content identifiers with the content dimensions and obtaining the dimensional index information of the recommended content. The dimensional index information can be an inverted index. For instance, taking the dimensions in the dimension dictionary as category and source, the dimensional index information of the recommended content could be as follows: Figure 7 As shown.
[0098] Steps S2 and S3 are relatively independent and have no fixed order. They can be completed simultaneously or sequentially.
[0099] 102. Based on the attribute information, identify at least one content identifier corresponding to the recommended content in the dimension index information.
[0100] Among them, attribute information can be information that indicates the attributes of the recommended content, such as content identifier, content information, source information, or content type.
[0101] For example, when querying attribute information to determine if recommended content identifiers exist, if such identifiers exist, the system identifies the recommended content's identifier within the attribute information. For instance, it identifies the recommended content's rowkey within the attribute information, thus obtaining the recommended content's identifier. The attribute information can include one or more recommended content identifiers. When the attribute information does not contain recommended content identifiers, the system filters at least one corresponding identifier from the dimension index information. For example, it determines the target content dimension of the recommended content based on the attribute information. This involves extracting the recommended content's dimension information from the attribute information, filtering the corresponding content dimension from a preset content dimension set, and using this content dimension as the target content dimension. The system then filters the dimension index information to identify the identifiers corresponding to the target content dimension. For instance, it filters the dimension dictionary of the dimension index information to identify a list of rowkeys associated with the target content dimension, and uses this list of rowkeys as the recommended content's identifiers. Finally, the system identifies at least one recommended content identifier, such as "person," within the attribute information, and uses the rowkey of the recommended content in the rowkey list as the recommended content's identifier.
[0102] 103. Based on the content identifier, determine the target time-series index corresponding to the recommended content in the time-series index information.
[0103] The time-series index information indicates the association between the content identifier of the recommended content and the time-series index; each content identifier of the recommended content corresponds to one time-series index. The time-series index stores consumption data corresponding to multiple consumption time units; therefore, it includes consumption data corresponding to at least one preset consumption time unit.
[0104] For example, the time-series index corresponding to the content identifier can be filtered out from the time-series index information, and this time-series index can be used as the target time-series index corresponding to the recommended content. For instance, the time-series index corresponding to the rowkey of each recommended content can be queried from the time-series index information, and this time-series index can be used as the target time-series index corresponding to the recommended content.
[0105] 104. Retrieve the target consumption data of the recommended content within the consumption time from the target time-series index, and return the target consumption data to the terminal.
[0106] The target time-series index includes index information corresponding to at least one candidate consumption time. Index information can be understood as information about the data structure that stores the aggregated values of consumption data within the candidate consumption time. The index information allows retrieval of the aggregated values of consumption data within a specific consumption time. For example, if the candidate consumption data is for one hour, the index information could store the aggregated values of the recommended content consumption data within that one hour. The candidate consumption time can be any consumption time unit used when constructing the time-series index information.
[0107] For example, when the candidate consumption time is the same as the consumption time, the target consumption data of the recommended content within the consumption time can be obtained from the index information corresponding to the candidate consumption time. For example, if the consumption time is 1 hour, when the candidate consumption time includes 1 hour, the data aggregation value can be directly extracted from the index information corresponding to the candidate consumption time that is the same as the consumption time, and the data aggregation value can be used as the target consumption data of the recommended content within the consumption time.
[0108] When candidate consumption times differ from actual consumption times, the consumption data in the index information is aggregated based on the consumption time to obtain the target consumption data for the recommended content within that consumption time. For example, at least one target candidate consumption time can be selected from the candidate consumption times to form the consumption time. For instance, if the consumption time is 20 minutes and candidate consumption times include days, hours, and 10 minutes, two 10-minute periods can be selected from the candidate consumption times to form the consumption time. Therefore, these two 10-minute periods constitute the target candidate consumption times. Then, the index information corresponding to the target candidate consumption times is selected from the target time-series index information to obtain the target index information. For example, if the target candidate consumption times are two 10-minute periods, the index information corresponding to these two 10-minute periods can be selected from the target time-series index information to obtain the target index information. The consumption data in the target index information is aggregated to obtain the target consumption data for the recommended content within the consumption time. For example, the aggregated data values in the target index information data structure can be aggregated again. The aggregation method is described above and will not be repeated here. The aggregated value after re-aggregation is used as the target consumption data for the recommended content within the consumption time. For example, taking the candidate consumption data corresponding to the target index information as two 10-minute intervals and a consumption time of 20 minutes as an example, it can be understood that the aggregated data values of the two 10-minute consumption data are aggregated again to obtain the target consumption data corresponding to 20 minutes. The target consumption data is then returned to the terminal. For example, the target consumption data can be sent directly to the terminal. When the number of target consumption data is large or the content is large, the storage address of the target consumption data can also be sent to the terminal, so that the terminal can retrieve the target consumption data within the consumption time from the memory or cache of the data acquisition device according to the storage address.
[0109] Optionally, before returning the target consumption data to the terminal, the target consumption data can be inspected to determine whether the target consumption data within the consumption time is accurate. Therefore, the data acquisition method may also include:
[0110] Obtain historical consumption data of recommended content within a preset consumption time, calculate the data error between the historical consumption data and the target consumption data, and return the target consumption data to the terminal when the data error does not exceed the preset error threshold.
[0111] Historical consumption data can be understood as consumption data of recommended content within a historical consumption period. This consumption data can be obtained through offline calculation, which is usually performed once a day. Target consumption data, on the other hand, is consumption data calculated in real time.
[0112] For example, historical consumption data for recommended content within a day or a minimum time unit can be obtained. Based on the obtained real-time consumption data of the recommended content, the historical consumption data is updated to obtain updated historical consumption data. For instance, real-time consumption data of recommended content in the content database is retrieved, added to the historical consumption data, and the consumption data within the consumption time period is recalculated, thus obtaining the updated historical consumption data. The data error between the updated historical consumption data and the target consumption data is calculated. For example, the updated historical consumption data and the target consumption data are reconciled to obtain the data error between them. This data error is compared with an error threshold. If the data error does not exceed the preset error threshold, it means the accuracy of the target consumption data meets the requirements, and the target consumption data can be sent to the terminal. If the data error exceeds the preset error threshold, it means the accuracy of the target consumption data does not meet the requirements. In this case, the system can either return to recalculate the target consumption data within the consumption time period in real time or stop obtaining the target consumption data.
[0113] The data acquisition device can be viewed as a content pop-up system. Compared to offline database calculations, which have fixed time ranges, granularities, and dimensions, making data acquisition inflexible and significantly lacking real-time latency (3-6 hours from content entry to POP viewing), this solution employs real-time calculation. Upon receiving a user's request for consumption data, it can display the target consumption data for the recommended content within the specified consumption timeframe within one minute. The entire data acquisition process can be divided into two stages: real-time calculation and real-time storage. Figure 8As shown, the real-time computing phase can include an access layer and a computing layer. The access layer is used to obtain real-time consumption data of recommended content, and the computing layer is used to perform window aggregation and associate content attribute data on the real-time consumption data to obtain a message queue composed of real-time content data. The real-time storage phase can include a writing layer, a storage layer, and an interface layer. The writing layer is used to write the data aggregation value corresponding to the consumption time unit calculated in the message queue to the storage area. The storage layer is used to construct dimensional index information and time-series index information. The interface layer is used to receive consumption data acquisition requests and return the target consumption data corresponding to the consumption data acquisition request to the terminal.
[0114] As can be seen from the above, in this embodiment of the invention, after receiving a request for consumption data acquisition of recommended content sent by the terminal, the request carries the attribute information and consumption time of the recommended content. Then, based on the attribute information, at least one content identifier corresponding to the recommended content is identified in the dimension index information. This dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content. Then, based on the content identifier, the target time-series index corresponding to the recommended content is determined in the time-series index information. This time-series index information is used to indicate the association between the content identifier of the recommended content and the time-series index. The time-series index includes consumption data corresponding to at least one preset consumption time unit. The target consumption data of the recommended content within the consumption time is obtained from the target time-series index, and the target consumption data is returned to the terminal. Since this scheme can quickly obtain the target time-series index of the recommended content through dimension index information and time-series index information, and performs multi-granularity pre-aggregation of real-time consumption data in advance and stores the pre-aggregated data in the target time-series index, the target consumption data can be directly filtered out from the index information of the target time-series index, which greatly reduces the acquisition time of the target consumption data. Therefore, the efficiency of data acquisition can be improved.
[0115] Based on the method described in the above embodiments, the following examples will provide further detailed explanations.
[0116] In this embodiment, the data acquisition device is specifically integrated into an electronic device, the electronic device is a server, the database of the content data platform is HBase, and the consumption time unit in the time-series index information is day, hour, and 10 minutes, as an example for explanation.
[0117] like Figure 9 As shown, a data acquisition method is described below, with the specific process as follows:
[0118] 201. The server obtains real-time content data for at least one recommended item.
[0119] For example, the server obtains real-time consumption data for at least one piece of recommended content, categorizes the real-time consumption data, performs window aggregation on the categorized real-time consumption data according to a preset time window, obtains the real-time basic consumption data of the recommended content, associates the content attribute data of the recommended content with the real-time basic consumption data, and encodes the associated data according to a preset encoding strategy to obtain the real-time content data of the recommended content. Specifically, it can be as follows:
[0120] (1) The server obtains real-time consumption data of at least one recommended content and classifies the real-time consumption data.
[0121] For example, when a user consumes recommended content on a content consumption platform such as a content viewing platform, content browser, or content application, the platform sends the consumption data generated by the consumption behavior to the server in real time. This allows the server to obtain real-time consumption data for at least one recommended piece of content. The real-time consumption data can be categorized according to business type; for example, it can be categorized first by source, and then further categorized by consumption behavior. Based on the categorization results, a real-time data warehouse is built to store the categorized real-time consumption data, thereby breaking down massive amounts of real-time consumption data into smaller message queues.
[0122] (2) The server performs window aggregation on the classified real-time consumption data according to the preset time window to obtain the real-time basic consumption data of the recommended content.
[0123] For example, the server retrieves raw data from a message queue composed of categorized real-time consumption data. By performing multi-row transformation on the retrieved raw data, processed real-time consumption data is obtained. Taking a preset time window of one minute as an example, all real-time consumption data within one minute, starting from the current time, can be filtered from the processed real-time consumption data, thus obtaining a window consumption data set. Aggregating the real-time consumption data in the window real-time consumption data set yields the basic real-time consumption data for the recommended content. For instance, taking the consumption data of a recommended content with 5 clicks within one minute as an example, window aggregation of the consumption data corresponding to these 5 clicks yields the basic real-time consumption data for that recommended content as 5 clicks on the recommended content within the current one minute.
[0124] (3) The server obtains the content attribute data of the recommended content.
[0125] For example, the server synchronizes content attribute data from the HBase primary content database to the HBase secondary database via binary logs. The secondary database then synchronizes the content attribute data to the Redis content cache database. The server retrieves the rowkey of the recommended content, filters the content attribute data corresponding to that rowkey in the Redis content cache database, and uses this content attribute data as the content attribute data for the recommended content.
[0126] (4) The server associates the content attribute data of the recommended content with the real-time basic consumption data, and encodes the associated data according to the preset encoding strategy to obtain the real-time content data of the recommended content.
[0127] For example, content attribute data such as the title, main content, source, and type of the recommended content can be retrieved from the content cache database and associated with real-time basic consumption data. According to a preset encoding strategy, the associated data is encoded to obtain the real-time content data of the recommended content, which can then be used to form a consumption queue.
[0128] 202. The server constructs a time-series index of recommended content based on basic consumption data.
[0129] For example, the server can directly obtain the configuration information of the time-series index, filter out the preset time-series information used to build the index, and read at least one consumption time unit contained in the preset time-series information. When the consumption time unit is a day, the server can filter out the basic consumption data within 24 hours starting from the current time from the basic consumption data, thus obtaining the target basic consumption data. Similarly, when the consumption time unit is an hour, the server can filter out the basic consumption data for each hour starting from the current time, thus obtaining the target basic consumption data. The target basic consumption data is then aggregated to obtain the aggregated data value corresponding to the consumption time unit. PB compression is then used to compress the aggregated data value corresponding to the consumption time unit, thus obtaining the consumption data for each consumption time unit.
[0130] For each recommended content, the server creates an initial time-series index corresponding to a rowkey. This initial time-series index can include index information corresponding to consumption time units in the preset time-series information. For example, when the consumption time units are days, hours, and 10 minutes, the constructed initial time-series index can include index information corresponding to one day, one hour, and one 10-minute intervals from the current time. Adding the consumption data corresponding to each 10-minute interval to the corresponding 10-minute index information ensures that the index information contains the consumption data corresponding to each 10-minute interval. The addition process can be understood as recording the consumption data in the data structure corresponding to the index. By adding the consumption data to the corresponding index information, the time-series index of the recommended content corresponding to the content identifier can be obtained. By merging the time-series indexes corresponding to each recommended content, the time-series index information of the recommended content can be obtained. The time-series index corresponding to the content identifier can be determined from the time-series index information.
[0131] Optionally, after obtaining the data aggregation value corresponding to the consumption time unit, the server also needs to store the data aggregation value to the corresponding shard. For example, the server can perform a consistent hash calculation on the rowkey of the recommended content to obtain the data aggregation value to be stored in the corresponding shard, and store the data aggregation value of the recommended content in each consumption time unit to the corresponding shard.
[0132] 203. The server constructs dimensional index information for recommended content based on basic content data.
[0133] For example, the server establishes a dimension dictionary and its corresponding inverted list. The dimension dictionary can include a preset set of multiple dimensions, which can include the content category of the recommended content (e.g., sports or business) or the content source (e.g., WeChat, a browser, a content consumption platform). Each content dimension corresponds to an inverted list. Based on the basic content data, the dimension information of the recommended content is determined. The content dimensions corresponding to the dimension information are then filtered from the preset dimension set to determine the content dimension of the recommended content. Recommended content belonging to the same type is selected, and these content dimensions are designated as target recommended content. The content identifier of the target recommended content is added to the inverted list corresponding to that content dimension, thus associating the content identifier with the content dimension and obtaining the dimension index information of the recommended content.
[0134] There is no fixed order for steps 202 and 203; they can be completed simultaneously or sequentially.
[0135] 204. The server receives a request from the terminal to retrieve consumption data.
[0136] For example, a user triggers a data retrieval request on their device to generate recommended content. This request includes the recommended content's attribute information and the consumption time. The device then sends this request to the server, allowing the server to receive the data retrieval request. When the recommended content's attribute information and the consumed data require a large amount of memory, the storage address of these two pieces of information can also be added to the data retrieval request. The device then sends this request to the server. Upon receiving the request, the server extracts the storage address and retrieves the recommended content's attribute information and consumption time from the device's memory or cache based on that address.
[0137] 205. Based on the attribute information, the server identifies at least one content identifier corresponding to the recommended content in the dimension index information.
[0138] For example, the server queries the attribute information to check if content identifier information for recommended content exists. If such information exists, the server identifies the rowkey of the recommended content within the identifier information, thus obtaining the content identifier. This identifier information can include one or more content identifiers for recommended content. If no such identifier exists, the server determines the target content dimension based on the attribute information. For instance, it extracts the dimension information from the attribute information, filters the corresponding content dimension from a preset set of content dimensions, and uses this dimension as the target content dimension. Finally, it filters the dimension dictionary of the dimension index information to find a list of rowkeys associated with the target content dimension, using this list as the content identifier for the recommended content, and then uses the rowkey of each recommended content in the rowkey list as its content identifier.
[0139] 206. Based on the content identifier, the server determines the target time-series index corresponding to the recommended content in the time-series index information.
[0140] For example, the server queries the time-series index corresponding to the rowkey of each recommended content in the time-series index information, and uses that time-series index as the target time-series index corresponding to the recommended content.
[0141] 207. The server retrieves the target consumption data of the recommended content within the consumption time from the target time-series index and returns the target consumption data to the terminal.
[0142] For example, when the candidate consumption time is the same as the consumption time, taking a consumption time of 1 hour as an example, when the candidate consumption time includes 1 hour, the data aggregation value can be directly extracted from the index information corresponding to the candidate consumption time that is the same as the consumption time, and the data aggregation value can be used as the target consumption data of the recommended content within the consumption time.
[0143] When candidate consumption times differ from actual consumption times, for example, if the consumption time is 20 minutes and the candidate consumption times only include days, hours, and 10 minutes, then two 10-minute periods can be selected from the candidate consumption times to form the consumption time. Therefore, these two 10-minute periods can be used as target candidate consumption times. The corresponding index information for these two 10-minute periods can then be selected from the target time-series index information to obtain the target index information. By further aggregating the aggregated values of the two 10-minute consumption data, the target consumption data corresponding to the 20-minute period can be obtained.
[0144] The server can directly send the target consumption data to the terminal. When the target consumption data is large in quantity or content, it can also send the storage address of the target consumption data to the terminal, so that the terminal can retrieve the target consumption data within the consumption time from the server's memory or cache according to the storage address.
[0145] Optionally, before returning the target consumption data to the terminal, the target consumption data can be checked to determine its accuracy within the consumption period. For example, the server can obtain historical consumption data for recommended content within a day or a minimum time unit. Based on the obtained real-time consumption data of the recommended content, the server retrieves the real-time consumption data of the recommended content from the content database, adds the real-time consumption data to the historical consumption data, and recalculates the consumption data within the consumption period to obtain updated historical consumption data. The updated historical consumption data is then reconciled with the target consumption data to determine the data error between the two. This data error is compared with an error threshold. If the data error does not exceed the preset error threshold, the accuracy of the target consumption data meets the requirements, and the target consumption data can be sent to the terminal. If the data error exceeds the preset error threshold, the accuracy of the target consumption data does not meet the requirements. In this case, the server can either return to recalculate the target consumption data within the consumption period in real time or stop acquiring the target consumption data.
[0146] The data acquisition device integrated into the server can be viewed as a backend server, belonging to the data platform. It works with the recommendation platform and content platform to complete tasks such as content recommendation, consumption, and querying. Specific application scenarios include... Figure 10As shown, the POP system, composed of data acquisition devices, pops up consumption data of content that meets the criteria to consumers for querying, or the content platform processes the recommended content based on the consumption data and sends the processed recommended content to the recommendation platform. The recommendation platform then recommends the processed content through various content consumption platforms, and returns the real-time consumption data generated during the recommendation process to the POP system in the data platform for further processing.
[0147] As can be seen from the above, after receiving a request for consumption data of recommended content sent by the terminal, the electronic device in this embodiment carries the attribute information and consumption time of the recommended content. Then, based on the attribute information, at least one content identifier corresponding to the recommended content is identified in the dimension index information. The dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content. Then, based on the content identifier, the target time-series index corresponding to the recommended content is determined in the time-series index information. The time-series index information is used to indicate the association between the content identifier of the recommended content and the time-series index. The time-series index includes consumption data corresponding to at least one preset consumption time unit. The target consumption data of the recommended content within the consumption time is obtained from the target time-series index, and the target consumption data is returned to the terminal. Since this scheme can quickly obtain the target time-series index of the recommended content through dimension index information and time-series index information, and performs multi-granularity pre-aggregation of real-time consumption data in advance and stores the pre-aggregated data in the target time-series index, the target consumption data can be directly filtered out from the index information of the target time-series index, which greatly reduces the acquisition time of the target consumption data. Therefore, the efficiency of data acquisition can be improved.
[0148] To better implement the above methods, embodiments of the present invention also provide a data acquisition device, which can be integrated into an electronic device, such as a server or terminal, and the terminal may include a tablet computer, a laptop computer, and / or a personal computer.
[0149] For example, such as Figure 11 As shown, the data acquisition device may include a receiving unit 301, an identification unit 302, a determining unit 303, and an acquisition unit 304, as follows:
[0150] (1) Receiving unit 301;
[0151] The receiving unit 301 is used to receive a consumption data acquisition request for recommended content sent by the terminal. The consumption data acquisition request carries the attribute information and consumption time of the recommended content.
[0152] For example, receiving unit 301 can be specifically used by a user to trigger a consumption data acquisition request for generated recommended content on the terminal. The user can add attribute information and consumption time of the recommended content to the consumption data acquisition request. The terminal then sends the consumption data acquisition request with the added attribute information and consumption time to the data acquisition device, allowing the data acquisition device to receive the request. When the memory for the attribute information and consumption data of the recommended content is large, the storage address of the attribute information and consumption data can also be added to the consumption data acquisition request. The terminal then sends the consumption data acquisition request with the added storage address to the data acquisition device, and the data acquisition device retrieves the attribute information and consumption time of the recommended content based on the storage address.
[0153] (2) Identification unit 302;
[0154] The identification unit 302 is used to identify at least one content identifier corresponding to the recommended content in the dimension index information based on the attribute information. The dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content.
[0155] For example, the identification unit 302 can be used to query whether there is content identifier information for recommended content in the attribute information. When there is content identifier information for recommended content in the attribute information, it identifies a content identifier corresponding to the recommended content in the content identifier information. When there is no content identifier information for recommended content in the attribute information, it filters at least one content identifier corresponding to the recommended content in the dimension index information based on the attribute information.
[0156] (3) Determine unit 303;
[0157] The determining unit 303 is used to determine the target time index corresponding to the recommended content in the time index information based on the content identifier. The time index information is used to indicate the association between the content identifier of the recommended content and the time index. The time index includes consumption data corresponding to at least one preset consumption time unit.
[0158] For example, the determination unit 303 can be used to query the time-series index corresponding to the rowkey of each recommended content in the time-series index information, and use the time-series index as the target time-series index corresponding to the recommended content.
[0159] (4) Obtain unit 304;
[0160] The acquisition unit 304 is used to retrieve the target consumption data of the recommended content within the consumption time from the target time-series index and return the target consumption data to the terminal.
[0161] For example, the acquisition unit 304 can be used to obtain the target consumption data of the recommended content within the consumption time from the index information corresponding to the candidate consumption time when the candidate consumption time is the same as the consumption time. When the candidate consumption time is different from the consumption time, the consumption data in the index information is aggregated according to the consumption time to obtain the target consumption data of the recommended content within the consumption time, and the target consumption data is returned to the terminal.
[0162] Optionally, the data acquisition device may also include a building unit 305, such as Figure 12 As shown, the specific details are as follows:
[0163] Construction unit 305 is used to construct the temporal index information and dimensional index information of the recommended content.
[0164] For example, the construction unit 305 can be used to obtain real-time content data of at least one recommended content, construct time-series index information of the recommended content based on basic consumption data, and construct dimensional index information of the recommended content based on basic content data.
[0165] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.
[0166] As can be seen from the above, in this embodiment, after the receiving unit 301 receives the consumption data acquisition request for recommended content sent by the terminal, the consumption data acquisition request carries the attribute information and consumption time of the recommended content. Then, the identification unit 302 identifies at least one content identifier corresponding to the recommended content in the dimension index information based on the attribute information. The dimension index information is used to indicate the association relationship between the content dimension and the content identifier of the recommended content. Then, the determining unit 303 determines the target time-series index corresponding to the recommended content in the time-series index information based on the content identifier. The time-series index information is used to indicate the association relationship between the content identifier of the recommended content and the time-series index. The index includes consumption data corresponding to at least one preset consumption time unit. The acquisition unit 304 retrieves the target consumption data of the recommended content within the consumption time from the target time-series index and returns the target consumption data to the terminal. Since this scheme can quickly obtain the target time-series index of the recommended content through dimensional index information and time-series index information, and performs multi-granularity pre-aggregation on real-time consumption data in advance and stores the pre-aggregated data in the target time-series index, the target consumption data can be directly filtered out from the index information of the target time-series index, which greatly reduces the acquisition time of the target consumption data. Therefore, the efficiency of data acquisition can be improved.
[0167] This invention also provides an electronic device, such as... Figure 13As shown, it illustrates a structural schematic diagram of the electronic device involved in an embodiment of the present invention, specifically:
[0168] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 13 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0169] The processor 401 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, it performs various functions and processes data, thereby performing overall detection of the electronic device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.
[0170] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0171] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0172] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0173] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:
[0174] The system receives a consumption data retrieval request for recommended content sent by a receiving terminal. This request carries attribute information and consumption time of the recommended content. Based on the attribute information, at least one content identifier corresponding to the recommended content is identified in the dimension index information. This dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content. Based on the content identifier, a target time-series index corresponding to the recommended content is determined in the time-series index information. This time-series index information is used to indicate the association between the content identifier of the recommended content and the time-series index. The time-series index includes consumption data corresponding to at least one preset consumption time unit. The system retrieves the target consumption data of the recommended content within the consumption time from the target time-series index and returns the target consumption data to the terminal.
[0175] For example, a user triggers a data retrieval request to generate recommended content on a terminal. The user adds attribute information and consumption time to the request. The terminal then sends this request to a data acquisition device, which receives the request. When the memory for the recommended content's attribute information and consumption data is large, the storage address of these two pieces of information can be added to the request. The terminal then sends this request to the data acquisition device, which retrieves the attribute information and consumption time based on the storage address. The system checks the attribute information for the existence of content identifiers for the recommended content. If a content identifier exists, it identifies the corresponding identifier. If not, it filters the dimension index information to find the identifier. Finally, it queries the time-series index for each recommended content's rowkey and uses this index as the target time-series index for that content. When the candidate consumption time is the same as the consumption time, the target consumption data of the recommended content within the consumption time is obtained from the index information corresponding to the candidate consumption time. When the candidate consumption time is different from the consumption time, the consumption data in the index information is aggregated according to the consumption time to obtain the target consumption data of the recommended content within the consumption time, and the target consumption data is returned to the terminal.
[0176] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0177] As can be seen from the above, in this embodiment of the invention, after receiving a request for consumption data acquisition of recommended content sent by the terminal, the request carries the attribute information and consumption time of the recommended content. Then, based on the attribute information, at least one content identifier corresponding to the recommended content is identified in the dimension index information. This dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content. Then, based on the content identifier, the target time-series index corresponding to the recommended content is determined in the time-series index information. This time-series index information is used to indicate the association between the content identifier of the recommended content and the time-series index. This time-series index includes consumption data corresponding to at least one preset consumption time unit. The target consumption data of the recommended content within the consumption time is obtained from the target time-series index, and the target consumption data is returned to the terminal. Since this scheme can quickly obtain the target time-series index of the recommended content through dimension index information and time-series index information, and performs multi-granularity pre-aggregation of real-time consumption data in advance and stores the pre-aggregated data in the target time-series index, the target consumption data can be directly filtered out from the index information of the target time-series index, greatly reducing the acquisition time of the target consumption data. Therefore, the efficiency of data acquisition can be improved.
[0178] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0179] Therefore, embodiments of the present invention provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the data acquisition methods provided in the embodiments of the present invention. For example, the instructions can execute the following steps:
[0180] The system receives a consumption data retrieval request for recommended content sent by a receiving terminal. This request carries attribute information and consumption time of the recommended content. Based on the attribute information, at least one content identifier corresponding to the recommended content is identified in the dimension index information. This dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content. Based on the content identifier, a target time-series index corresponding to the recommended content is determined in the time-series index information. This time-series index information is used to indicate the association between the content identifier of the recommended content and the time-series index. The time-series index includes consumption data corresponding to at least one preset consumption time unit. The system retrieves the target consumption data of the recommended content within the consumption time from the target time-series index and returns the target consumption data to the terminal.
[0181] For example, a user triggers a data retrieval request to generate recommended content on a terminal. The user adds attribute information and consumption time to the request. The terminal then sends this request to a data acquisition device, which receives the request. When the memory for the recommended content's attribute information and consumption data is large, the storage address of these two pieces of information can be added to the request. The terminal then sends this request to the data acquisition device, which retrieves the attribute information and consumption time based on the storage address. The system checks the attribute information for the existence of content identifiers for the recommended content. If a content identifier exists, it identifies the corresponding identifier. If not, it filters the dimension index information to find the identifier. Finally, it queries the time-series index for each recommended content's rowkey and uses this index as the target time-series index for that content. When the candidate consumption time is the same as the consumption time, the target consumption data of the recommended content within the consumption time is obtained from the index information corresponding to the candidate consumption time. When the candidate consumption time is different from the consumption time, the consumption data in the index information is aggregated according to the consumption time to obtain the target consumption data of the recommended content within the consumption time, and the target consumption data is returned to the terminal.
[0182] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0183] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0184] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the data acquisition methods provided in the embodiments of the present invention, the beneficial effects that any of the data acquisition methods provided in the embodiments of the present invention can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0185] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations of the data acquisition aspect described above.
[0186] The foregoing has provided a detailed description of a data acquisition method, apparatus, and computer-readable storage medium provided by embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A data acquisition method, characterized in that, include: Obtain real-time content data for at least one recommended content, wherein the real-time content data includes basic consumption data and basic content data; wherein the basic consumption data is data obtained by aggregating the real-time consumption data within a fixed time window, and the basic content data is the attribute information of the recommended content; Based on the aforementioned basic consumption data, construct the time-series index information of the recommended content; Based on the aforementioned basic content data, construct the dimensional index information for the recommended content; The receiving terminal sends a consumption data acquisition request for recommended content, wherein the consumption data acquisition request carries the attribute information and consumption time of the recommended content; Based on the attribute information, at least one content identifier corresponding to the recommended content is identified in the dimension index information, wherein the dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content; Based on the content identifier, the target time index corresponding to the recommended content is determined in the time index information. The time index information is used to indicate the association between the content identifier of the recommended content and the time index. The time index includes consumption data corresponding to at least one preset consumption time unit. The target consumption data of the recommended content within the consumption time period is obtained from the target time series index, and the target consumption data is returned to the terminal.
2. The data acquisition method according to claim 1, characterized in that, The acquisition of real-time content data for at least one recommended item includes: Obtain real-time consumption data for at least one recommended content item, and classify the real-time consumption data; The categorized real-time consumption data is subjected to multi-row to column transformation to obtain the processed real-time consumption data. Filter out at least one real-time consumption data corresponding to a preset time window from the processed real-time consumption data to obtain a window real-time consumption data set. The real-time consumption data in the window real-time consumption data set is aggregated to obtain the real-time basic consumption data of the recommended content; The content attribute data of the recommended content is associated with the real-time basic consumption data, and the associated data is encoded according to a preset encoding strategy to obtain the real-time content data of the recommended content.
3. The data acquisition method according to claim 2, characterized in that, Before associating the content attribute data of the recommended content with the real-time basic consumption data, the method further includes: Synchronize the content attribute data of the recommended content in the content database to the content cache database; Obtain the content identifier of the recommended content, and filter out the content attribute data corresponding to the content identifier in the content cache database to obtain the content attribute data of the recommended content.
4. The data acquisition method according to claim 3, characterized in that, The step of constructing the time-series index information of the recommended content based on the basic consumption data includes: Obtain preset time-series information for constructing time-series index information, wherein the preset time-series information includes at least one consumption time unit; Based on the consumption time unit, the basic consumption data is aggregated to obtain the consumption data corresponding to the consumption time unit; Based on the consumption data corresponding to the consumption time unit, the time-series index information of the recommended content is constructed.
5. The data acquisition method according to claim 4, characterized in that, The step of aggregating the basic consumption data according to the consumption time unit to obtain the consumption data corresponding to the consumption time unit includes: Filter out at least one basic consumption data point corresponding to the consumption time unit from the basic consumption data to obtain the target basic consumption data; The target basic consumption data is aggregated to obtain the aggregated data value within the consumption time unit; The aggregated data is compressed to obtain the consumption data corresponding to the time unit.
6. The data acquisition method according to claim 5, characterized in that, After aggregating the target basic consumption data to obtain the aggregated data value corresponding to the consumption time unit, the method further includes: Based on the content identifier of the recommended content, determine the storage area for the data aggregation value of the recommended content within the consumption time unit; The aggregated data value is stored in the storage area.
7. The data acquisition method according to claim 4, characterized in that, The step of constructing the time-series index information of the recommended content based on the consumption data corresponding to the consumption time unit includes: Based on the preset time sequence information and the content identifier of the recommended content, an initial time sequence index corresponding to the content identifier is constructed, and the initial time sequence index includes the index information corresponding to the consumption time unit; Add the consumption data corresponding to the consumption time unit to the index information to obtain the time-series index corresponding to the content identifier; The temporal indexes corresponding to the content identifiers are fused to obtain the temporal index information of the recommended content.
8. The data acquisition method according to claim 3, characterized in that, The step of constructing the dimensional index information of the recommended content based on the basic content data includes: Based on the basic content data, the content dimensions corresponding to the recommended content are selected from the preset dimension set; Filter out target recommended content with the same content dimension from the recommended content; The content identifier of the target recommended content is associated with the content dimension to obtain the dimension index information of the recommended content.
9. The data acquisition method according to claim 1, characterized in that, The step of identifying at least one content identifier corresponding to the recommended content in the dimension index information based on the attribute information includes: When the attribute information contains the content identifier information of the recommended content, at least one content identifier corresponding to the recommended content is identified in the content identifier information; When the content identifier information of the recommended content is not present in the attribute information, at least one content identifier corresponding to the recommended content is selected from the dimension index information based on the attribute information.
10. The data acquisition method according to claim 9, characterized in that, The step of filtering out at least one content identifier corresponding to the recommended content from the dimension index information based on the attribute information includes: Based on the attribute information, the target content dimension of the recommended content is determined; Filter out the content identifier information corresponding to the target content dimension from the dimension index information; At least one content identifier of the recommended content is identified in the content identifier information.
11. The data acquisition method according to claim 1, characterized in that, The target time-series index includes index information corresponding to at least one candidate consumption time. The step of obtaining the target consumption data of the recommended content within the consumption time from the target time-series index includes: When the candidate consumption time is the same as the consumption time, the target consumption data of the recommended content within the consumption time is obtained from the index information corresponding to the candidate consumption time. When the candidate consumption time is different from the consumption time, the consumption data in the index information is aggregated according to the consumption time to obtain the target consumption data of the recommended content within the consumption time.
12. The data acquisition method according to claim 11, characterized in that, The step of aggregating the consumption data in the index information according to the consumption time to obtain the target consumption data of the recommended content within the consumption time includes: At least one target candidate consumption time is selected from the candidate consumption times to form the consumption time. The target index information is obtained by filtering the index information corresponding to the target candidate consumption time from the target time series index; The consumption data in the target index information is aggregated to obtain the target consumption data of the recommended content within the consumption time.
13. The data acquisition method according to claim 1, characterized in that, Before returning the target consumption data to the terminal, the method further includes: Obtain historical consumption data of the recommended content within a preset consumption period; The historical consumption data is updated based on real-time consumption data of the recommended content; Calculate the data error between the updated historical consumption data and the target consumption data; Returning the target consumption data to the terminal includes: returning the target consumption data to the terminal when the data error does not exceed a preset error threshold.
14. A data acquisition device, characterized in that, include: A construction unit is used to acquire real-time content data of at least one recommended content, wherein the real-time content data includes basic consumption data and basic content data; wherein the basic consumption data is data obtained by aggregating real-time consumption data within a fixed time window, and the basic content data is attribute information of the recommended content; based on the basic consumption data, a time-series index information of the recommended content is constructed; based on the basic content data, a dimensional index information of the recommended content is constructed. A receiving unit is used to receive a consumption data acquisition request for recommended content sent by a terminal, wherein the consumption data acquisition request carries the attribute information and consumption time of the recommended content; The identification unit is used to identify at least one content identifier corresponding to the recommended content in the dimension index information based on the attribute information, wherein the dimension index information is used to indicate the association between the content dimension and the content identifier of the recommended content; The determining unit is used to determine the target time-series index corresponding to the recommended content in the time-series index information based on the content identifier. The time-series index information is used to indicate the association between the content identifier of the recommended content and the time-series index. The time-series index includes consumption data corresponding to at least one preset consumption time unit. The acquisition unit is used to acquire the target consumption data of the recommended content within the consumption time from the target time-series index, and return the target consumption data to the terminal.
15. An electronic device, characterized in that, It includes a processor and a memory, the memory storing an application program, and the processor running the application program in the memory to implement the data acquisition method as described in any one of claims 1-13.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the data acquisition method as described in any one of claims 1-13.
17. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium; the processor of the computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the data acquisition method as described in any one of claims 1-13.
Citation Information
Patent Citations
Information recommendation method, device, and server
CN105096151A
System and method for personalizing and recommending content
US20170064405A1