Data Processing Method, Device, Equipment and Storage Medium for Content-Based Push

The method improves real-time data processing and query efficiency in content push systems by aggregating interaction data in message queues, addressing the inefficiencies of delayed data processing and querying in existing systems.

CN113609374BActive Publication Date: 2025-07-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110160293.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-05
Publication Date
2025-07-15
Estimated Expiration
2041-02-05

AI Technical Summary

Technical Problem

In the prior art, the processing of content interaction data is low in real time, which causes the account to be unable to obtain feedback information from other accounts on the content in a timely manner, affecting the efficiency and accuracy of data query.

Method used

By pushing the basic content interactive data into the message queue in real time, and performing data association operations based on the associated information set, converting it into aggregated data, reducing the amount of data, improving processing timeliness and query efficiency.

Benefits of technology

Real-time processing of content interactive data is realized, timeliness and accuracy of data queries are improved, and query delay and computing resource consumption are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113609374B_ABST
    Figure CN113609374B_ABST
Patent Text Reader

Abstract

The present application provides a data processing method, apparatus, device and storage medium based on content push, which relates to the technical field of data processing and aims to improve the timeliness of processing content interaction data. The method includes: based on the operation type of the interaction operation, the basic content interaction data determined by the interaction operation triggered for the push content is pushed to the corresponding message queue in real time. For each message queue, the corresponding aggregated data is obtained respectively in the following manner: based on at least one association information set, an association information set association operation is performed on a message queue to obtain the corresponding aggregated data; one association information set association operation includes: based on an association information set, the basic content interaction data obtained by a message queue within a preset time window is converted into aggregated data. In this method, multi-dimensional association operation processing can be performed on the basic content interaction data in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and in particular, to a data processing method, apparatus, device, and storage medium based on content push. Background Art

[0002] In the era of self-media, a content push system can perform offline data processing on content interaction data generated by an account's interaction operations on push content, use the results of the offline data processing as intermediate data, and store it in a data warehouse. Subsequently, the account can query data related to the content it created based on the above intermediate data. However, in the above process, a large amount of content interaction data needs to be delayed for a certain period of time, such as one day, before data processing, resulting in low real-time performance in processing content interaction data. Therefore, how to improve the timeliness of processing content interaction data to enhance the timeliness of the account's querying of data related to the content based on the intermediate data obtained from the data processing has become an issue that needs to be considered. Summary of the Invention

[0003] Embodiments of the present application provide a data processing method, apparatus, device, and storage medium based on content push, which are used to improve the timeliness of processing interaction data for push content.

[0004] In the first aspect of the present application, a data processing method based on content push is provided, including:

[0005] In response to interaction operations triggered by each target account for the push content obtained by each, obtain basic content interaction data;

[0006] Based on the operation type of the interaction operation, push the basic content interaction data in real time to at least one message queue associated with the operation type;

[0007] For each message queue in the at least one message queue, perform data association operations respectively in the following manner to obtain corresponding aggregated data:

[0008] Based on at least one set of association information, perform an association information set association operation on a message queue to obtain corresponding aggregated data; where, one association information set association operation includes: based on a set of association information, convert the basic content interaction data obtained by the message queue within a preset time window into aggregated data, and the number of the converted aggregated data is not greater than the number of the basic content interaction data received within the preset time window.

[0009] In the second aspect of the present application, a data processing apparatus based on content push is provided, including:

[0010] A data collection unit, configured to obtain basic content interaction data in response to interaction operations triggered by each target account for the push content obtained by each;

[0011] A data splitting unit, configured to, based on the operation type of the interaction operation, push the basic content interaction data to at least one message queue associated with the operation type in real time;

[0012] A data aggregation unit, configured to perform data association operations on each of the at least one message queue respectively in the following manner to obtain corresponding aggregated data: performing an associated information set association operation on a message queue based on at least one set of associated information to obtain corresponding aggregated data; wherein, one associated information set association operation includes: based on a set of associated information, converting the basic content interaction data obtained by the message queue within a preset time window into aggregated data, and the number of the converted aggregated data is not greater than the number of the basic content interaction data received within the preset time window.

[0013] In a possible implementation manner, each basic content interaction data includes an account identifier associated with the target account that triggers the interaction operation and a content identifier associated with the push content that triggers the interaction operation, and the data aggregation unit is specifically configured to perform any one or a combination of the following operations:

[0014] If the set of associated information includes a set of account portraits, determining the basic content interaction data obtained by the message queue within a first preset time window as a first set of interaction data, and respectively performing the following operations on each account identifier included in the basic content interaction data in the first set of interaction data: based on the account portrait data of each target account recorded in the set of account portraits, obtaining the account portrait data of the target account associated with an account identifier, and aggregating the basic content interaction data including the account identifier in the first set of interaction data through the obtained account portrait data to obtain a first aggregated data;

[0015] If the set of associated information includes a set of content information, determining the basic content interaction data obtained by the message queue within a second preset time window as a second set of interaction data, and respectively performing the following operations on each content identifier included in the basic content interaction data in the second set of interaction data: based on the content information of each push content recorded in the set of content information, obtaining the content information of the push content associated with a content identifier, and aggregating the basic content interaction data including the content identifier in the second set of interaction data through the obtained content information to obtain a second aggregated data.

[0016] In a possible implementation, the data aggregation unit is specifically configured to:

[0017] Determine the basic content interaction data in the first interaction data set that contains the one account identifier;

[0018] Obtain the content identifiers included in the determined basic content interaction data, and generate a content identifier set;

[0019] Associate the obtained account portrait data, the one account identifier, and the content identifier set to obtain a corresponding first aggregated data.

[0020] In a possible implementation, the data aggregation unit is specifically configured to:

[0021] Determine the basic content interaction data in the second interaction data set that contains the one content identifier;

[0022] Obtain the account identifiers included in the determined basic content interaction data, and generate an account identifier set;

[0023] Associate the obtained content information, the one content identifier, and the account identifier set to obtain a corresponding second aggregated data.

[0024] In a possible implementation, the account portrait set is stored in a first key-value pair database, and the first key-value pair database is periodically updated based on a first period;

[0025] The content information set is stored in a second key-value pair database, and the second key-value pair database is obtained by periodically backing up a content database based on a second period, and the content database is used to record the content information of the pushed content in real time.

[0026] In a possible implementation, the data aggregation unit is further configured to:

[0027] For each message queue in the at least one message queue, perform a data association operation respectively in the following manner. After obtaining the corresponding aggregated data, for each obtained aggregated data, perform a data storage operation respectively in the following manner, and store each aggregated data into a corresponding disk slice: Store an aggregated data into the disk slice mapped by the pushed content associated with the one aggregated data, where one disk slice is used to store the aggregated data associated with one pushed content;

[0028] And

[0029] In response to a data query request for a to-be-query pushed content, obtain the aggregated data associated with the to-be-query pushed content from the disk slice mapped by the to-be-query pushed content;

[0030] Perform data processing on the obtained aggregated data according to the data processing rules associated with the target service requirements, obtain query data, and return the query data.

[0031] In a possible implementation manner, the content identifier of the push content to be queried and the query time period are carried in the data query request, and the data aggregation unit is specifically configured to:

[0032] Based on the content identifier, determine the disk shard mapped by the push content to be queried as the target disk shard; and

[0033] Divide the query time period into at least one sub-time period through a preset time granularity;

[0034] According to the storage addresses of the aggregated data mapped by each sub-time period in the target disk shard in the aggregated data associated with the push content to be queried, determine the data query index information corresponding to the data query request;

[0035] Based on the data index information, obtain the aggregated data mapped by each sub-time period from the storage addresses of the aggregated data mapped by each sub-time period in the target disk shard.

[0036] In the third aspect of the present application, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.

[0037] In the fourth aspect of the present application, a computer program product is provided. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in the first aspect above.

[0038] In the fifth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions. When the computer instructions run on a computer, the computer is enabled to execute the method described in the first aspect.

[0039] Since the embodiments of the present application adopt the above technical solutions, they have at least the following technical effects:

[0040] In the embodiments of the present application, on the one hand, the obtained basic content interaction data is pushed to the corresponding message queue for processing in real time to obtain the corresponding aggregated data, realizing the real-time processing of the basic content interaction data, improving the timeliness of data processing, and further improving the timeliness and accuracy of querying the feedback information of the target account on the pushed content based on the aggregated data; on the other hand, in the embodiments of the present application, based on the association information set, a large amount of basic content interaction data obtained in each message queue within each time window is converted into a relatively small amount of aggregated data, reducing the quantity of the aggregated data reflecting the feedback information of the target account on the pushed content. Thus, when the account queries the feedback information of the target account on the pushed content based on the aggregated data, the quantity of the aggregated data that needs to be queried and processed is significantly reduced, reducing the latency of querying the feedback information of the target account on the pushed content and improving the query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 FIG. is a schematic diagram of an application scenario of data processing based on content push provided by an embodiment of the present application;

[0042] Figure 2 FIG. is an example diagram of the process of a data processing method based on content push provided by an embodiment of the present application;

[0043] Figure 3 FIG. is a schematic diagram of the principle of updating an account portrait set and a content information set provided by an embodiment of the present application;

[0044] Figure 4 FIG. is a schematic diagram of the process of account dimension association provided by an embodiment of the present application;

[0045] Figure 5 FIG. is a schematic diagram of the process of information dimension association provided by an embodiment of the present application;

[0046] Figure 6 FIG. is a schematic diagram of the process of writing aggregated data into a disk shard provided by an embodiment of the present application;

[0047] Figure 7 FIG. is a schematic diagram of the process of obtaining aggregated data based on a data query request provided by an embodiment of the present application;

[0048] Figure 8 FIG. is a schematic diagram of the process of obtaining aggregated data from a target disk shard provided by an embodiment of the present application;

[0049] Figure 9 FIG. is a schematic diagram of the framework of a content push system provided by an embodiment of the present application;

[0050] Figure 10An example diagram of the process of a data processing method based on push content provided by an embodiment of the present application;

[0051] Figure 11 An example diagram of a process for aggregating basic content interactions provided by an embodiment of the present application;

[0052] Figure 12 An example structural diagram of a data processing device based on content push provided by an embodiment of the present application;

[0053] Figure 13 A structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0054] To better understand the technical solutions provided by the embodiments of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0055] To facilitate those skilled in the art to better understand the technical solutions of the present application, some concepts involved in the present application will be described below.

[0056] 1) Content and push content

[0057] In the era of self-media, content generally can refer to audio, video, graphics, text, etc.; in the embodiments of the present application, the push content is the content pushed by the content push system to an account; the push content in the embodiments of the present application can be the content created and published by a single user's account, or the content published by an account corresponding to a group composed of multiple users, etc. The push content can also be actively published by an account of Professionally-generated Content (PGC) or User-generated Content (UGC);

[0058] The push content in the embodiments of the present application can but is not limited to including at least one type of information such as text, audio, video, article, graphics, text, picture, etc. or a multimedia resource obtained by any combination; among them, the article can but is not limited to the graphics and text actively edited and published by a public account created by self-media, and the graphics and text can but is not limited to including vertical small graphics and text, horizontal short graphics and text, or graphics and text that can slide up and down or left and right, etc.; the video can but is not limited to being provided by a user of professionally-produced content or user-generated content, and finally provided in the form of a Feeds stream.

[0059] 2) Content publishing account, target account, and query account

[0060] Generally, an account represents a user's identity. In the embodiments of the present application, various accounts involved can also be regarded as users. In the embodiments of the present application, the account that sends push content to the content push server is called the content publishing account (which can also be referred to as the content producer or content production end), the account that receives the push content distributed by the content push server is called the target account (which can also be referred to as the content consumer or content consumption end), and the account that triggers a data query for the push content is called the query account. Among them, the query account can be the above-mentioned content producer or an account associated with each content space in the content push system. The above content can be, but is not limited to, content channels or content communities that publish push content in the content push system, etc.

[0061] 3) Basic content interaction data and association information set

[0062] In the embodiments of the present application, the basic content interaction data can be data generated by the interaction operations of the target account for the push content, and the basic content interaction data can characterize the feedback information of the target account on the push content. In the embodiments of the present application, the basic content interaction data can include, but is not limited to, the account identifier associated with the target account that triggers the interaction operation and the content identifier associated with the push content that triggers the above interaction operation.

[0063] The association information set in the embodiments of the present application contains information used to aggregate multiple basic content interaction data. For example, the association information set can include information about the target account that triggers the above interaction operation, and the association information set can also include information about the content that triggers the above interaction operation, etc. Among them, the association information set involved in the embodiments of the present application can have various forms, such as a dimension table containing information about the target account or content. In subsequent embodiments, the dimension table will be used as an example of the association information set for illustration.

[0064] 4) Professionally-generated Content (PGC) and User-generated Content (UGC)

[0065] The above PGC is an Internet term, representing an organization or entity that produces professional content, such as a video website, expert-generated content, such as Weibo, etc.

[0066] UGC refers to user-generated content, which emerged along with the Web 2.0 concept featuring personalized advocacy. It is not a specific business but a new way for users to use the Internet, that is, from the original Internet usage mode mainly focused on downloading content to an Internet usage mode that equally emphasizes downloading and uploading content, referring to user-generated content. In UGC platforms, anyone can upload relevant text, images, videos, etc. The articles and short videos shared by users in current instant messaging applications, as well as the short videos in various short video applications, all belong to the content generated in the UGC way.

[0067] 5) Feeds stream (i.e., information stream)

[0068] A Feeds stream is an information stream that continuously updates and presents content to users. It is the source of information and can also be called source material, feed, information provision, contribution, summary, source, news subscription, web feed, etc. A Feed is a data form that continuously provides content to users. It is a resource aggregator composed of multiple information sources that provide content. Users actively subscribe to the information sources that provide content and are provided with content. That is, a Feed combines several information sources subscribed to by users actively to form a content aggregator to help users continuously obtain the latest content of the subscribed information sources. The Feeds stream is a data format. Through the Feeds stream, websites transmit the latest information to users, usually arranged in a timeline. The timeline is the most primitive, intuitive, and basic display form in the Feeds stream.

[0069] The prerequisite for users to be able to subscribe to a website is that the website provides a source of information. Converging Feeds streams into one place is called aggregation, and the software used for aggregation is called an aggregator. For end-users, an aggregator is software specifically used to subscribe to websites and is generally also called an RSS reader, feed reader, news reader, etc.

[0070] 6) Page views per unique visitor (PV / UV)

[0071] Page views (PV) refer to that each access by a user to a website is recorded. Multiple accesses by a user to the same page refer to the page view or click volume of the website, and the access volume is accumulated. The number of unique visitors (UV): One computer client accessing a website is counted as one visitor. Within 24 hours, multiple accesses from the same address are only counted as one time.

[0072] PV / UV = Page Views per Unit UV (page views browsed by independent visitors), which reflects one of the factors of page view quality.

[0073] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.

[0074] The embodiments of this application relate to artificial intelligence (AI) and machine learning technologies, and are designed based on computer vision technology and machine learning (ML) in artificial intelligence, especially involving big data processing technology in artificial intelligence.

[0075] Artificial intelligence is to use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence.

[0076] Machine learning is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance.

[0077] The basic content interaction data involved in the embodiments of this application is massive big data. Big data generally refers to a data set that cannot be captured, managed, and processed by conventional software tools within a certain time range. It is a massive, high-growth-rate, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery ability, and process optimization ability. The main characteristics of big data are large volume, variety, complexity, difficult to process, and high value. The data processing method provided in the embodiments of this application involves big data processing technology. The process of big data processing mainly includes data collection, data preprocessing, data storage, data processing and analysis, data display or data visualization, data application, etc. Among them, data quality runs through the entire process of big data processing, and each data processing link will have an impact on the quality of big data. Big data processing technologies can include, but are not limited to, massively parallel processing (MPP) databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems, etc. Among them, the cloud computing platform, also known as the cloud platform, refers to a service based on hardware resources and software resources that provides computing, network, and storage capabilities.

[0078] Among them, big data processing technologies include offline data processing and real-time data processing. Among them: Offline data processing, also called offline computing, batch computing, or batch processing computing, refers to first extracting and storing data locally. Once the data is extracted, it is static and unchanged, and then subsequent processing and analysis are carried out. It is a calculation carried out on the premise that all input data is known before the calculation starts, the input data will not change, and the result should be obtained immediately after solving a problem.

[0079] Real-time data processing, also known as real-time computing, refers to regarding the production of data as a continuous and dynamic data stream, pre-defining processing rules, and performing processing and analysis in the direction according to the pre-defined rules when the data flows through; generally, it is carried out for massive data, and the efficiency of data processing is required to be at the second level. Real-time computing is mainly divided into two parts: real-time data warehousing and real-time data computing. Real-time computing is applied when the data source is real-time and uninterrupted, and the response time of the user is also required to be real-time (such as for the streaming data of large websites: website access PV / UV, what content the user accessed, what content was searched, etc. Real-time data computing and analysis can dynamically and real-time refresh the user access data, display the real-time traffic changes of the website, and analyze the traffic and user distribution of each hour of the day). The data volume is large and it is impossible or unnecessary to budget, but the response time to the user is required to be real-time.

[0080] In the process of performing data association operations in the embodiments of this application, based on the real-time data processing technology in big data processing technology, the basic content interaction data obtained in real time can be processed to improve the timeliness of processing the basic content interaction data.

[0081] The design concept of this application will be described below.

[0082] In the self-media era, an account can publish the content it creates to the content push system at any time and place. Then, the content push system can push the above content to other accounts. Other accounts can perform interaction operations on the obtained content to generate content interaction data and report the content interaction data to the content push system. The content push system can perform multi-layer offline data processing on the received content interaction data through data analysis tools such as Spark, and finally store the intermediate data obtained from the offline data processing in a data warehouse such as Mysql or ES. Then, the account can query the data related to the content it creates based on the above intermediate data. However, in the above process, the offline processing of a large amount of content interaction data is often delayed, resulting in low timeliness of content interaction data processing. It is impossible to timely obtain the feedback information of other accounts on the content based on the above intermediate data, so it is impossible to timely discover abnormal content based on the feedback information (such as content involving sensitive topics, etc., which can be but not limited to). In addition, the current processing dimension of the above intermediate data is low. When querying data related to the content based on the intermediate data, it is necessary to perform relatively complex statistical analysis processing on the intermediate data before the query data can be obtained, resulting in long data query time and low efficiency.

[0083] In view of this, the inventor designed a data processing method, device, equipment and storage medium based on content push, which is used to improve the real-time performance and processing dimension of processing the content interaction data of an account for content, so as to improve the efficiency and real-time performance of the account to obtain the feedback information of other accounts on the content based on the processed content interaction data. In this method, in order to reduce the delay of offline data processing of content interaction data, according to the operation type of the interaction operation, the obtained basic content interaction data is split in real time and pushed to the corresponding message queue (Message Queue, MQ). And based on one or more association information sets, a large amount of basic content interaction data obtained in each time window in each message queue is converted into a small amount of aggregated data to improve the data processing dimension, reduce the quantity of the obtained aggregated data, and then reduce the quantity of the aggregated data that needs to be statistically analyzed when querying related data, reducing the data query delay and improving the data query efficiency.

[0084] As an embodiment, in order to further improve the efficiency of content interaction data related to account query push content, in the embodiments of the present application, the aggregated data associated with the same push content can also be stored in the same disk partition. When querying the relevant data in the push, the aggregated data can be directly obtained from the corresponding disk partition for statistical analysis, reducing the time for querying the aggregated data, and thus improving the time for returning query data based on the aggregated data.

[0085] To more clearly understand the design concept of the present application, the following is an example introduction to the application scenarios in the embodiments of the present application.

[0086] Please refer to Figure 1 , which represents an application scenario of data processing based on content push. This scenario includes a terminal device 110, a content push server 120, and a data processing server 130; the terminal device 110, the content push server 120, and the data processing server 130 can communicate with each other through a network, where:

[0087] The terminal device 110 is used to receive the push content uploaded by the content publishing account and send the push content to the content push server 120; the terminal device 110 can also receive the push content distributed by the content push server 120, and the terminal device 110 can also receive the interaction operation triggered by the target account for the push content.

[0088] As an embodiment, a client of the content push system can be installed on the terminal device 110. The terminal device 110 can upload the push content to the content push server 120 through the client, and receive the push content distributed by the content push server 120 through the client, or receive the interaction operation triggered by the target account for the push content, etc.

[0089] The content push server 120 is used to receive the push content uploaded by the content publishing account through the terminal device 110, and distribute the push content to the terminal devices 110 used by one or more target accounts.

[0090] The data processing server 130 is used to obtain basic content interaction data in response to the interaction operations triggered by each target account for the push content obtained by each of them, and based on the operation type of the above interaction operation, push the above basic content interaction data to at least one message queue associated with the operation type in real time; and based on at least one associated information set, perform an associated information set association operation on the basic content interaction data received by each message queue to obtain the corresponding aggregated data.

[0091] The message queue involved in the embodiments of the present application can be, but is not limited to, a container for storing data (or messages), that is, a queue for storing the basic content interaction data to be transmitted; the message queue is an asynchronous inter-service communication method and an important component in a distributed system, which can be used for publishing and subscribing to messages, mainly solving problems such as application coupling, asynchronous messages, and traffic peak shaving, and realizing a high-performance, highly available, scalable, and eventually consistent architecture; the message queue involved in the embodiments of the present application can be, but is not limited to, at least one of RocketMQ, RabbitMQ, Kafka, etc.

[0092] The terminal device 100 in the embodiments of the present application can be a mobile terminal, a fixed terminal or a portable terminal, such as a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio or video player, a digital camera or a video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof.

[0093] The content push server 120 and the data processing server 130 in the embodiments of the present application can be the same server or different servers; and the content push server 120 and the data processing server 130 can be independent physical servers, or a server cluster or a distributed system composed of multiple physical servers, or multiple cloud servers providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms in cloud service technology (for example, the content push server 120 can be, but is not limited to, including the servers 120-1, 120-2, or 120-3 shown in the figure; for example, the data processing server 130 can be, but is not limited to, including the servers 130-1, 130-2, or 130-3 shown in the figure); the functions of the above content push server 120 can be implemented by one or more cloud servers, or can also be implemented by one or more cloud server clusters, etc.; the functions of the above data processing server 130 can be implemented by one or more cloud servers, or can also be implemented by one or more cloud server clusters, etc.

[0094] Among them, cloud service technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. applied based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud service technology is an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system background, which can only be achieved through cloud service technology.

[0095] As an embodiment, the above data processing server 130 can be, but is not limited to, the server in the data stream processing engine. The data stream processing engine is a data processing engine that processes streaming data (i.e., data streams), and can be, but is not limited to, including Apache, Flink, Storm, Samza, etc. Among them, Flink is a framework for stateful computing processing of unbounded and bounded data streams and is a relatively efficient data stream processing engine, which believes that all data exists in the form of streams.

[0096] As an embodiment, after the data processing server 130 in the embodiment of the present application obtains the corresponding aggregated data, it can also, but is not limited to, store the obtained aggregated data so that the query account can query the data for the push content, and obtain the query data associated with the target business requirement based on the aggregated data associated with the push content. Among them, the query data can be, but is not limited to, processing the obtained aggregated data according to the data processing rules associated with the target business requirement. The above target business requirement can be, but is not limited to, some business metrics, and the business metrics can be, but is not limited to, including at least one data metric such as the viewing rate of the target account viewing the push content, the probability of the target account sharing the push content, the popularity ranking of the push content, and the PV / UV of the push content.

[0097] As an embodiment, in order to improve the efficiency of data query for push content, the embodiment of the present application can also implement data query through a data analysis system. Specifically, the obtained aggregated data can be pushed to the storage space in the data analysis system for storage. The data analysis system responds to the data query request, processes the corresponding aggregated data according to the data processing rules associated with the target business requirement, obtains the query data, and returns the obtained query data to the query account, etc. The above data analysis system can be, but is not limited to, including Druid, the columnar storage database ClickHouse, or Zookeeper, where:

[0098] The above-mentioned Druid is a distributed data processing system that supports real-time multi-dimensional joint analysis processing (On-Line Analytical Processing, OLAP). It not only supports high-speed real-time data ingestion processing but also supports real-time and flexible multi-dimensional data analysis queries. Therefore, in the embodiments of this application, Druid is used to perform flexible and fast multi-dimensional OLAP analysis on large-volume aggregated data to improve the efficiency of data associated with the content pushed in account queries.

[0099] The above-mentioned ClickHouse is a column-oriented database management system (Database Management System, DBMS) with an MPP architecture, which is used for OLAP analysis and uses columnar storage. The data is organized by column, and the data belonging to the same column is saved together, and different files are used to save between columns respectively. In the embodiments of this application, ClickHouse is used to store the above-mentioned aggregated data, and it has the following advantages: it can dynamically create, modify, or delete databases, tables, and views of aggregated data without restarting the service, it can dynamically query, insert, modify, or delete the above-mentioned aggregated data, it can set the operation permissions of the database or table storing the aggregated data according to user granularity to ensure the security of the aggregated data, it can flexibly perform backup import and export of the aggregated data, and ClickHouse provides a cluster mode that can automatically manage multiple database nodes of the aggregated data; thus improving the flexibility of the aggregated data associated with the content pushed in account queries.

[0100] Next, the technical solutions of the embodiments of this application will be further introduced. It should be noted that the following introduced technical solutions are only exemplary.

[0101] Before introducing the detailed technical solutions, first, the basic content interaction data involved in the embodiments of this application will be described:

[0102] The basic content interaction data involved in the embodiments of this application is data containing feedback information of the target account on the pushed content, that is, the basic content interaction data can but is not limited to being data generated by the interaction operations triggered by the target account on the obtained pushed content. The basic content interaction data can but is not limited to include at least one of information such as the account identifier associated with the target account, the content identifier associated with the pushed content, and the operation information of the interaction operation performed by the target account on the pushed content.

[0103] In the embodiments of the present application, the basic content interaction data can be, but is not limited to, set in the form of an n-tuple, where n is a positive integer, and one of the n elements represents an information of the basic content interaction data; for example, in the embodiments of the present application, the basic content interaction data can be represented in the form of a triple {X1, X2, X3}, a multi-tuple {X1, X2, X3, X4...}, etc., where X1 can be, but is not limited to, an account identifier associated with the target account that triggers the interaction operation, X2 can be, but is not limited to, a content identifier associated with the push content targeted by the above interaction operation, X3 can be, but is not limited to, operation information of the interaction operation (such as, but not limited to, including the operation type, operation name, etc. of the interaction operation), and X4 can be, but is not limited to, information such as the time, place, or manner of the above interaction operation; it can also be set in the form of "Account A1 performed an interaction operation A4 on the push content A2", etc.; in the embodiments of the present application, the specific form of the basic content interaction data is not limited, and those skilled in the art can flexibly set it according to actual needs.

[0104] Secondly, the interaction operation on the push content in the embodiments of the present application will be described.

[0105] The interaction operation in the embodiments of the present application refers to the operation performed by the target account on the received push content. In the embodiments of the present application, the interaction operation can be, but is not limited to, classified according to the preference degree of the target account for the push content. In this case, the operation types of the interaction operation can include positive feedback operations, negative feedback operations, etc. The operation types of the interaction operation can also include types that do not show the preference degree of the target account for the push content, such as at least one operation type among exposure operations and interaction operations for the push content, where:

[0106] The above exposure operation can be, but is not limited to, an operation of exposing the push content.

[0107] The above positive feedback operation refers to the operation in which the target account gives positive feedback on the received push content. In the embodiments of the present application, the positive feedback operation may but is not limited to include at least one of the following operations: click operation on the push content, full text view operation of viewing the full text of the push content, follow operation of following the account that publishes the push content, like operation on the push content, electronic currency transfer operation of transferring electronic currency for the push content, multimedia resource transfer operation of transferring multimedia resources for the push content, favorite operation of the target account saving the push content to the favorite folder of the target account, download operation of downloading the push content, sharing operation of forwarding the network link of the push content, etc.; the above electronic currency refers to the currency stored in the electronic wallet held by the account in electronic form (such as the wallet in a payment application or the wallet in a banking application, etc.), and the electronic currency may but is not limited to include electronic bills, game resources (such as game coins, game equipment), etc.; the above multimedia resources may but is not limited to include dynamic emoticons, game resources, electronic heart emoticons, etc.

[0108] The above negative feedback operation refers to the operation in which the target account gives negative feedback on the received push content. In the embodiments of the present application, the negative feedback operation may but is not limited to include at least one of the following operations: blocking operation of blocking the push content, unfollow operation of canceling the follow of the account that publishes the push content, negative reading operation of the viewing duration of the push content being less than the preset duration, etc.

[0109] The above interaction operation refers to the operation in which the target account interacts with the push content, or the operation in which the target account interacts with the account that publishes the push content. In the embodiments of the present application, the interaction operation may but is not limited to the comment operation of commenting on the push content, etc.; if the push content is a live media stream, the interaction operation may also be the chat operation of the target account chatting with other accounts in the live broadcast room where the live media stream is published, etc.

[0110] As an embodiment, in the embodiments of the present application, the interaction operation may also be directly classified based on the specific operation content of the interaction operation. For example, the above exposure operation, click operation, electronic currency transfer operation, follow operation, multimedia resource transfer operation, favorite operation, download operation, sharing operation, blocking operation, negative reading operation, comment operation, etc. are respectively regarded as one operation type, etc.

[0111] Based on Figure 1 the application scenario, the following is an example description of a method for starting an application for instant messaging involved in the embodiments of the present application; please refer to Figure 2 , which shows a schematic diagram of a data processing method based on content push involved in the embodiments of the present application, and specifically includes the following steps:

[0112] Step S201: In response to the interaction operations triggered by each target account for the pushed content obtained respectively, obtain basic content interaction data.

[0113] As an embodiment, in this step, based on an original message queue, receive the basic content interaction data indicated by the interaction operations triggered by each target account. The above original message queue may include, but is not limited to, at least one of RocketMQ, RabbitMQ, Kafka, etc.

[0114] Step S202: Based on the operation type of the above interaction operation, push the above basic content interaction data to at least one message queue associated with the operation type in real time.

[0115] Specifically, in this step, one message queue may be associated with one operation type or multiple operation types. To improve the accuracy of processing basic content interaction data, in the embodiments of the present application, for each operation type, a message queue may be created as its associated message queue, and one message queue receives the basic content interaction data generated by the interaction operation of one operation type.

[0116] Step S203: For each of the above at least one message queue, perform data association operations respectively in the following manner to obtain corresponding aggregated data: Based on at least one set of association information, perform an association information set association operation on a message queue to obtain corresponding aggregated data. Among them, one association information set association operation includes: Based on a set of association information, convert the basic content interaction data obtained by the above one message queue within a preset time window into aggregated data, where the number of the converted aggregated data is not greater than the number of the basic content interaction data received within the preset time window.

[0117] Each preset time window involved in the embodiments of the present application may include, but is not limited to, the reference time period length for dividing time periods. That is, in step 203, time may be divided with the preset time window as the reference time period length to obtain each time period. Then, for the basic content interaction data obtained by a message queue in each time period, convert them into corresponding aggregated data respectively. Among them, the above preset time window is not limited, and those skilled in the art can set it according to actual needs. For example, it may include, but is not limited to, setting the preset time window to 1 minute, 5 minutes, 10 minutes, etc.

[0118] As an embodiment, in order to improve the accuracy and dimension of aggregating and processing the basic content interaction data, at least one of the account portrait set and the content information set may be included in the above-mentioned associated information set. In the embodiments of the present application, it may but is not limited to performing account dimension association on the basic content interaction data in each message queue based on the account portrait set to obtain first aggregation data representing the information of each target account's interaction operations on those push contents; or performing content dimension association on the basic content interaction data in each message queue based on the content information set to obtain second aggregation data representing the information of the specific accounts that perform interaction operations with each push content; wherein in this step, separate account dimension association or content dimension association may be performed on the basic content interaction data in each message queue, or both account dimension association and content dimension association may be performed on the basic content interaction data in each message queue simultaneously.

[0119] In the following content of this application, the above-mentioned account dimension association and content dimension association are further described.

[0120] First, the account portrait set and the content information set involved in the embodiments of this application are described.

[0121] As an embodiment, in order to improve the information richness of the data after account dimension association in the embodiments of the present application, the account portrait set in the embodiments of the present application includes the account portrait data of each target account registered in the content push system. In the embodiments of the present application, in order to improve the efficiency of obtaining the account portrait data of each target account, the account portrait data may be associated with the account identifier of the target account, and then based on the account identifier of the target account, the account portrait data of the target account can be quickly obtained.

[0122] The account portrait data involved in the embodiments of the present application is also called an account portrait or a user portrait (UserProfile), which refers to the tagging of the information of the user associated with the account. Currently, the account portrait of the target account is mainly extracted through the interaction operations between the target account and the push content (such as exposure operations, click operations, like operations, comment operations, etc.). The account portrait is precipitated on the tags of the push content and includes static information and dynamic information. Among them, the static information may be provided when the target account is first registered, such as demographic attributes and social attributes such as gender, age, permanent residence, native place, height, education level, marital status, education level, asset situation, income situation, occupation, etc.; the dynamic information can be mined from the behavior data of the target account, including interests such as photography, sports, food, beauty, clothing, tourism, education, etc. associated with the target account through content logs or third-party data, and including account psychology (i.e., the psychology of the user using the target account), motivation, values, life attitude, personality, etc. of awareness and cognition.

[0123] As an embodiment, in order to improve the information richness of the data after content dimension association, the content information set in the embodiments of the present application includes the content information (also known as content meta-information) of each push content published in the content push system; the content meta-information of the push content in the embodiments of the present application can be but is not limited to data describing some features and attributes of the push content, and the content meta-information can be but is not limited to including at least one of the file size of the push content, the cover image link, the content title, the content format, the release time, the account information of the account that released the push content, the picture information in the push content (such as but not limited to including the picture size, the picture format, the picture creator), the original mark indicating whether the push content is original, the first release mark indicating whether the push content is the first release, the classification information of the push content, etc.

[0124] As an embodiment, when the above push content is a video, the above content meta-information may further include at least one of the link of the cover image of the video, the file format of the video, the playing duration of the video, the bit rate of the video, etc.

[0125] Among them, the above classification information can be the classification and tag information of the push content when manually reviewing the push content, or the classification of the push content when automatically reviewing the push content by a machine; in the embodiments of the present application, the classification of the push content can be based on the format, release source, content field, etc. of the push content. For example, when the push content is an article, it can be but is not limited to classified based on the content field involved in the push content, and the push content can be classified at multiple levels. For example, an article explaining a certain brand and model of mobile phone, its first-level classification can be technology, the second-level classification can be smart phones, and the third-level classification can be domestic mobile phones, etc., and the tag information is "a certain brand, a certain model of mobile phone".

[0126] As an embodiment, in order to improve the efficiency of accessing the above account portrait set and content information set, in the embodiments of the present application, the account portrait set can be stored in a key-value database, and the content information set can be stored in another key-value database. In the following content, the key-value database storing the account portrait information set is referred to as the first key-value database, and the key-value data storing the content information set is referred to as the second key-value database.

[0127] The key-value pair database involved in the embodiments of the present application is a non-relational database that stores relevant data using a simple key-value method. The key-value pair database involved in the embodiments of the present application may but is not limited to including Redis cache. Redis is a distributed cache middleware. In the embodiments of the present application, the key-value pair database is used to store the above-mentioned account profile set and content information set. On the one hand, it can support the persistent caching of account profile data and content meta-information, and can save the account profile data and content meta-information in memory to disk, and can reload the account profile data and content meta-information for use when restarting. On the other hand, the key-value pair database not only supports storing the above-mentioned account profile data and content meta-information in a simple key-value data structure, but also supports storing the above-mentioned account profile data and content meta-information data in data structures such as List, Set, Zset, and Hash. In addition, the key-value pair database in the embodiments of the present application also supports the backup of account profile data and content meta-information, and supports fast access to a large amount of data. The key-value pair database Redis is in milliseconds, and the access speed of Redis is basically nearly 1000 times that of accessing other databases such as Hbase, which can significantly improve the efficiency of accessing the account profile set and content information set.

[0128] As an embodiment, please refer to Figure 3 , the account profile set can be in the form of a dimension table. In order to improve the accuracy of the account profile data obtained from the account profile set, in the embodiments of the present application, the account profile data of each target account in the above-mentioned account profile set can be updated periodically but not limited to based on the first period; no excessive limitation is imposed on the above-mentioned first period, and those skilled in the art can set it according to actual needs, such as it can be set to 1 day, 7 days, 10 days, etc. but not limited to.

[0129] As an embodiment, please continue to refer to Figure 3, the content information set can be in the form of a dimension table. After the content push system receives the content information of the push content, it can record the obtained content information in the content database in real time (such as but not limited to including the Hbase database). In the embodiments of the present application, in order to improve the accuracy of obtaining content information from the content information set, the content database can be periodically backed up based on the second period to obtain the above-mentioned second key-value pair database; in the embodiments of the present application, a layer of Redis cache can be set as the above-mentioned second key-value pair database before accessing the content in the HBase database that records the content information. When 1000 pieces of content information in the HBase database are pushed to the above-mentioned Redis cache, since accessing 1000 pieces of data in the HBase is in seconds, while accessing the Redis is in milliseconds, the speed of accessing the Redis is basically 1000 times that of accessing the HBase; at the same time, in order to prevent waste of cache for expired content information data, the second period of the cache can be set to 24 hours or 48 hours, etc., and the cache consistency can be ensured by listening to the write HBase Proxy. In this way, the access time is changed from more than ten minutes to seconds; the above-mentioned second period is not limited too much, and those skilled in the art can obtain it according to actual needs.

[0130] As an embodiment, in order to reduce the risk of cache penetration, in the embodiments of the present application, during the process of recording the content information of the push content associated with the content identifier in real time in the content database, abnormal content identifiers can be detected, and then the content information of the push content associated with the above-mentioned abnormal content identifiers is not recorded in the content database, so that during the process of performing content dimension association on the basic content interaction data, the content information of the push content associated with the above-mentioned abnormal content identifiers can be directly filtered out; among them, the abnormal content identifier can be but not limited to the content identifier of the push content that has been deleted from the content push system due to content security or copyright supervision reasons of the push content, such as the content identifiers of some non-existent push content due to security, policy, or consistency. The above-mentioned cache penetration refers to the phenomenon that all caches cannot be hit and all access pressures are transferred to the next-level storage.

[0131] As an embodiment, in order to reduce the phenomenon of cache avalanche during the process of periodically caching the first key-value pair database and the second key-value pair database, a peak shaving and valley filling operation can be performed during the process of periodically caching the first key-value pair database and the second key-value pair database, that is, for different key-value pair databases, different periodic update time periods are set to stagger the cache time, so as to reduce the possibility of avalanche during the process of periodically updating the first key-value pair database and the second key-value pair database; the above-mentioned cache avalanche refers to the phenomenon that all cache contents expire at the same time and the cache does not work at a certain moment.

[0132] The following content of the embodiments of this application describes the specific processes of the above-mentioned account dimension association and content dimension association.

[0133] (1) Account Dimension Association

[0134] In this association method, based on the account portrait set, the basic content interaction data obtained by each message queue within each preset time window can be aggregated to obtain the first aggregated data; specifically, refer to Figure 4 , it can, but is not limited to, perform the following operations for each message queue in at least one message queue:

[0135] Step S401: Determine the basic content interaction data obtained by a message queue within the first preset time window as the first interaction data set.

[0136] The above first preset time window is the reference time period length for dividing time periods in the process of account dimension association. The specific data of the first preset time window is not limited, and those skilled in the art can set it according to actual needs. For example, it can, but is not limited to, set the first preset time window to 1 minute, 5 minutes, 8 minutes, etc.

[0137] In this step, for a message queue, the basic content interaction data obtained within each first preset time window is respectively determined as a first interaction data set. If the first preset time window is 1 minute, the basic content interaction data obtained by the message queue within each minute is respectively determined as a first interaction data set.

[0138] As an embodiment, considering the association between different message queues and different operation types of interaction operations, in some business requirements, the requirements for the aggregation dimension of the basic content interaction data associated with different operation types may be different. Therefore, in order to improve the flexibility of account dimension association for the basic content interaction data obtained from different message queues, different first preset time windows can be set for different message queues in the embodiments of this application.

[0139] Step S402: For each account identifier included in the basic content interaction data in the above first interaction data set, obtain the first aggregated data associated with each account identifier.

[0140] Specifically, the first interaction data set may contain multiple basic content interaction data, and the account identifiers included in different basic content interaction data may be different. Therefore, a first interaction data set may contain one or more account identifiers. Furthermore, for a first interaction data set, the basic content interaction data containing the same account identifier is respectively aggregated to obtain a first aggregated data associated with the same account identifier.

[0141] As an embodiment, it is possible but not limited to perform the operations of the following steps S4021 and S4022 respectively for each of the above-mentioned account identifiers:

[0142] Step S4021: Based on the account portrait data of each target account recorded in the above-mentioned account portrait set, obtain the account portrait data of the target account associated with an account identifier.

[0143] Specifically, the account portrait data in the account portrait set may be associated with the account identifier. Therefore, in this step, the account portrait data of the target account associated with an account identifier can be directly obtained based on an account identifier.

[0144] Step S4022: Aggregate the basic content interaction data in the above-mentioned first interaction data set that contains the above-mentioned account identifier through the obtained account portrait data to obtain a first aggregated data.

[0145] As an embodiment, in order to increase the dimension of the information in the obtained first aggregated data, in the embodiments of the present application, the obtained account portrait data and the basic content interaction data in the above-mentioned first interaction data set that contains the above-mentioned account identifier may be aggregated. Specifically, the first aggregated data may be obtained but not limited to through the following methods:

[0146] Determine the basic content interaction data in the above-mentioned first interaction data set that contains the above-mentioned account identifier; obtain the content identifiers included in the determined basic content interaction data to generate a content identifier set; associate the obtained account portrait data, the above-mentioned account identifier, and the content identifier set to obtain a corresponding first aggregated data.

[0147] In the embodiments of the present application, the specific representation form of the first aggregated data is not limited, and those skilled in the art can set it according to actual needs. For example, the first aggregated data can be represented as a triple {account identifier, account portrait data, content identifier set}, or the obtained first aggregated data can be identified in the form of a data table;

[0148] Here, in order to facilitate understanding of the difference in the data volume of the relevant data before and after the association in the account dimension, please refer to Table 1, which gives an example of the basic content interaction data included in a first interaction data set. Each row of data in Table 1 is a basic content interaction data. Please refer to Table 2, which gives the schematic after aggregating the basic content interaction data in Table 1 into the first aggregated data. Each row of data in Table 2 is a first aggregated data.

[0149] Table 1: Basic content interaction data included in a first interaction data set

[0150] Account identifier Content identifier Account identifier 1 Content identifier 1 Account identifier 1 Content identifier 3 Account identifier 1 Content identifier 5 Account identifier 2 Content identifier 2 Account identifier 2 Content identifier 3 Account identifier 2 Content identifier 4 Account identifier 2 Content identifier 6 Account identifier 3 Content identifier 1 Account identifier 3 Content identifier 2 … …

[0151] Table 2: The first aggregated data determined by associating the basic content interaction data in Table 1 at the account dimension

[0152] Account identifier Account portrait data Set of content identifiers Account identifier 1 Account portrait data 1 Content identifier 1, Content identifier 3, Content identifier 5 Account identifier 2 Account portrait data 2 Content identifier 2, Content identifier 3, Content identifier 4, Content identifier 6 Account identifier 3 Account portrait data 3 Content identifier 1, Content identifier 2 … … …

[0153] Before the account dimension association in Table 1, there were a total of 9 basic content interaction data in the above first interaction data set. After the account dimension association, there are only 3 pieces of the first aggregated data. The quantity of the first aggregated data in Table 2 is significantly less than the quantity of the basic content interaction data in Table 1. Considering that in the actual application process, the basic content interaction data pushed to each message queue is massive, after performing the above account dimension association using the method provided in the embodiments of the present application, the data volume of the data including the feedback information of the target account on the pushed content will be significantly reduced, which can significantly reduce the occupation of data storage resources, and can reduce the quantity of accessing the relevant first aggregated data during later data query, so as to improve the efficiency of data query and reduce the latency of data query.

[0154] (2) Content dimension association

[0155] In this association method, based on the content information set, the basic content interaction data obtained by each message queue within each preset time window can be aggregated to obtain the second aggregated data; specifically, refer to Figure 5 , for each message queue in at least one message queue, the following operations can be performed:

[0156] Step S501: Determine the basic content interaction data obtained by a message queue within the second preset time window as the second interaction data set.

[0157] The above second preset time window is the reference time period length for dividing time periods during the account dimension association. The specific data of the second preset time window is not limited, and those skilled in the art can set it according to actual needs. For example, it can be but not limited to setting the second preset time window to 1 minute, 5 minutes, or 8 minutes, etc.

[0158] In this step, for a message queue, the basic content interaction data obtained within each second preset time window is respectively determined as a second interaction data set. If the second preset time window is 5 minutes, then the basic content interaction data obtained by this message queue every 5 minutes is respectively determined as a second interaction data set.

[0159] As an embodiment, considering that different message queues are associated with different operation types of interactive operations, in some business requirements, the requirements for the aggregation dimensions of the basic content interaction data associated with different operation types may be different. Therefore, in order to improve the flexibility of content dimension association for the basic content interaction data obtained from different message queues, different second preset time windows can be set for different message queues in the embodiments of the present application.

[0160] Step S502: For each content identifier included in the basic content interaction data in the second interaction data set, obtain the second aggregation data associated with each content identifier.

[0161] Specifically, the first interaction data set may include multiple pieces of basic content interaction data, and the content identifiers included in different pieces of basic content interaction data may be different. Therefore, the basic content interaction data in a second interaction data set may include one or more content identifiers. Furthermore, for a second interaction data set, the basic content interaction data including the same content identifier are respectively aggregated to obtain a second aggregation data associated with the same content identifier.

[0162] As an embodiment, it is possible but not limited to perform the operations of steps S5021 and S5022 respectively for each of the above account identifiers:

[0163] Step S5021: Based on the content information of each push content recorded in the content information set, obtain the content information of the push content associated with a content identifier.

[0164] Specifically, the content information in the content information set may be associated with the content identifier. Therefore, in this step, the content information of the push content associated with a content identifier can be directly obtained based on a content identifier.

[0165] Step S5022: Aggregate the basic content interaction data including the above one content identifier in the second interaction data set through the obtained content information to obtain a second aggregation data.

[0166] As an embodiment, in order to increase the dimension of the information in the obtained second aggregation data, in the embodiments of the present application, the obtained content information and the basic content interaction data including the above one content identifier in the first interaction data set can be aggregated. Specifically, the second aggregation data can be obtained in the following ways but is not limited thereto:

[0167] Determine the basic content interaction data in the second interaction data set that contains the above one content identifier; obtain the account identifiers included in the determined basic content interaction data, and generate an account identifier set; associate the obtained content information, the above one content identifier, and the account identifier set to obtain a corresponding second aggregation data.

[0168] In the embodiments of the present application, the specific representation form of the second aggregation data is not limited, and those skilled in the art can set it according to actual needs. For example, the second aggregation data can be represented as a triple {content identifier, content information, account identifier set}, or each obtained second aggregation data can be identified in the form of a data table;

[0169] Here, in order to facilitate understanding of the difference in the data volume of the relevant data before and after content dimension association, please refer to Table 3, which gives an example of the basic content interaction data included in a second interaction data set. Each row of data in Table 3 is a basic content interaction data. Please refer to Table 4, which gives a schematic representation of the first aggregation data after aggregating the basic content interaction data in Table 1. Each row of data in Table 4 is a first aggregation data.

[0170] Table 3: Basic content interaction data included in a second interaction data set

[0171] Account identifier Content identifier Account identifier 1 Content identifier 3 Account identifier 2 Content identifier 2 Account identifier 2 Content identifier 3 Account identifier 3 Content identifier 1 Account identifier 3 Content identifier 2 Account identifier 4 Content identifier 1 Account identifier 5 Content identifier 2 Account identifier 5 Content identifier 3 Account identifier 6 Content identifier 3 … …

[0172] Table 4: Second aggregation data determined by performing content dimension association on the basic content interaction data in Table 1

[0173] Content identifier Content information Set of account identifiers Content identifier 1 Content information 1 Account identifier 3, Account identifier 4 Content identifier 2 Content information 2 Account identifier 2, Account identifier 3, Account identifier 5 Content identifier 3 Content information 3 Account identifier 1, Account identifier 2, Account identifier 5, Account identifier 6 … … …

[0174] Before the content dimension association in Table 3, there are a total of 9 basic content interaction data in the second interaction data set. After the content dimension association, there are only 3 second aggregation data. The number of second aggregation data in Table 4 is significantly less than the number of basic content interaction data in Table 3. Considering that in the actual application process, the basic content interaction data pushed to each message queue is massive, therefore, after performing the above content dimension association using the method provided in the embodiments of the present application, the data volume of the data containing the feedback information of the target account on the pushed content will be significantly reduced, which can significantly reduce the occupation of data storage resources, and can reduce the number of related second aggregation data accessed during later data query, so as to improve the efficiency of data query and reduce the latency of data query.

[0175] As an embodiment, during the process of storing the aggregated data after obtaining it, as the data volume of the pushed content grows, the data volume of the basic content interaction data of the pushed content for massive content distribution in the information flow is very large. The data volume of each piece of aggregated data obtained based on the basic content interaction data may reach tens of billions. If the massive aggregated data is directly written into a data analysis system such as Clickhouse, it will cause the QPS of the Zookeeper cluster to be too high. Therefore, in the embodiments of the present application, the Batch batch method can be adopted to write the aggregated data into the storage space. At the same time, in order to relieve the pressure on the Zookeeper cluster, in the embodiments of the present application, a Batch with a size of hundreds of thousands (such as 300,000, which can be but is not limited to) can be selected to write the aggregated data into the storage space.

[0176] Generally, when storing data, the data is often written to a distributed table, which will cause a bottleneck in the disk of a single machine. For example, when writing data into Clickhouse, since the underlying of Clickhouse uses Mergetree, the principle is similar to the underlying LSM-Tree of HBase, and there will be a problem of write amplification during the merging process, which increases the disk pressure. The peak is tens of millions of data per minute, and it takes dozens of seconds to write. If a Merge is in progress, it will block the write request and the query will be very slow. In view of this, in the embodiments of the present application, before storing the aggregated data, disk Raid can be performed on the storage space in the above data processing engine or data analysis system to obtain multiple disk shards and improve the disk IO. At the same time, before writing the aggregated data, the aggregated data can be sharded by Hash (which can be but is not limited to this method), and a large amount of aggregated data can be written into different disk shards. In addition, considering that by the above method, when writing the aggregated data into the storage space of the distributed system, there may be a problem that the local maximum value Top is not the global Top, especially when statistically monitoring the global results. For example, the aggregated data associated with the same pushed content is stored in different storage shards respectively. When calculating the top 100 (Top100) pushed content with the highest reading volume globally, there is a pushed content that is Top100 on disk shard 1 but not Top100 on other disk shards, resulting in the loss of some data during data aggregation and affecting the final statistical result. In view of this, in order to improve the accuracy of data query for the pushed content, in the embodiments of the present application, the aggregated data associated with a pushed content can be stored in the same disk shard.

[0177] In view of the above, in the embodiments of the present application, but not limited to, for each piece of aggregated data obtained, the following operations may be respectively performed to store the above-mentioned aggregated data into corresponding disk slices: store an aggregated data into the disk slice mapped by the push content associated with the above-mentioned aggregated data, where one disk slice is used to store the aggregated data associated with one push content. The above-mentioned disk slice may be, but not limited to, obtained by performing disk Raid on the storage space in the above-mentioned data processing engine or data analysis system. The disk slice mapped by the push content associated with the aggregated data may be, but not limited to, obtained by performing Hash routing on the content identifier of the push content associated with the aggregated data.

[0178] Furthermore, when the query account queries the data related to the push content through a data query request, the query account may trigger a data query request for the push content to be queried. Then, the above-mentioned data processing server 130, the above-mentioned data processing engine or data analysis system may respond to the data query request for the push content to be queried, and obtain the aggregated data associated with the push content to be queried from the disk slice mapped by the push content to be queried; and perform data processing on the obtained aggregated data according to the data processing rules associated with the target business requirements to obtain the query data, and return the above-mentioned query data.

[0179] As an embodiment, please refer to Figure 6 , in the process of storing the aggregated data after obtaining the aggregated data, in order to ensure the consistency of the recorded aggregated data, in the embodiments of the present application, but not limited to, a highly available solution implemented by a Zookeeper cluster may be adopted. Write the aggregated data into a disk slice, only write one copy, and then write it into the Zookeeper cluster. The Zookeeper cluster tells other copies of the same disk slice, and other copies then come to pull the aggregated data to ensure the consistency of the aggregated data; since the Zookeeper cluster is lightweight, and when writing the aggregated data, any copy can be written, and other copies can obtain consistent data through the Zookeeper cluster. In addition, even if other nodes fail to obtain data for the first time, as long as it is found that the aggregated data it has is inconsistent with the aggregated data recorded on the Zookeeper cluster later, it will try to obtain the aggregated data again to ensure consistency.

[0180] As an embodiment, in order to improve the efficiency of data query and reduce the time of data query, the content identifier of the push content to be queried and the query time period are carried in the data query request in the embodiments of the present application. Furthermore, the data query index information corresponding to the data query request may be determined based on the information carried in the data query request, and then the corresponding aggregated data may be directly obtained from the respective storage addresses indicated in the data index information for processing. For details, please refer to Figure 7, specifically, it may but is not limited to include the following steps:

[0181] Step S701: Based on the content identifier of the push content to be queried carried in the data query request, determine the disk shard mapped by the above-mentioned push content to be queried as the target disk.

[0182] In the embodiments of the present application, when storing aggregated data, routing has been performed on the content identifier, and the aggregated data associated with a content identifier only exists on one disk shard. Therefore, in this step, the content identifier carried in the data query request is first subjected to Hash routing according to the same rule to determine the disk shard mapped by the push content to be queried as the target disk shard.

[0183] Step S702: Divide the query time period carried in the above-mentioned data query request into at least one sub-time period by a preset time granularity.

[0184] The above-mentioned preset time granularity may but is not limited to be a preset time length. In the embodiments of the present application, the specific value of the preset time granularity is not limited. Those skilled in the art can set it according to actual business requirements. For example, it may but is not limited to setting the above-mentioned preset time granularity to 1 minute, 5 minutes, 10 minutes, etc.

[0185] It should be noted that there is no fixed execution order for the above-mentioned step S701 and step S702.

[0186] Step S703: According to the storage addresses of the aggregated data mapped by each sub-time period in the aggregated data associated with the above-mentioned push content to be queried in the above-mentioned target disk shard, determine the data query index information corresponding to the above-mentioned data query request.

[0187] Specifically, it may but is not limited to determining the storage addresses of the aggregated data mapped by each sub-time period in the above-mentioned target disk shard as the data query index information, or determining the time period information of each sub-time period and the storage addresses of the aggregated data mapped by each sub-time period in the target disk shard as the above-mentioned data query index information. The above-mentioned time period information may include the time range of the sub-time period, or include information such as the time range of the sub-time period and the content identifier of the push content to be queried; for ease of understanding, an example of the data query index information corresponding to a data query request is given here. Please refer to Table 5. This data query request is triggered for the push content M to be queried.

[0188] Table 5: Example of data query index information

[0189] Hash(content identifier + hash(date of M, sub - time period 1 of M)) Save address 1 of aggregated data for sub - time period 1 of M Hash(content identifier + hash(date of M, sub - time period 2 of M)) Save address 2 of aggregated data for sub - time period 2 of M … …

[0190] Among them, the content identifier in Table 5 can be but is not limited to the content ID of the push content to be queried; the aggregated data in the table can include at least one of the first aggregated data and the second aggregated data associated with the push content to be queried.

[0191] Step S704, based on the above data index information, obtain the aggregated data mapped by each of the above sub-time periods from the storage addresses of the aggregated data mapped by each of the above sub-time periods in the above target disk slice.

[0192] Specifically, reference can be made to Figure 8 , and an example diagram of obtaining aggregated data from the target disk slice is given here. As can be seen from the figure, usually when querying the aggregated data related to a single content ID (content identifier), the distributed table will send the query to all disk slices, and then return the query results for summarization. In the embodiment of the present application, since the content ID has been routed, the aggregated data associated with a content ID only exists on one disk slice, and the remaining disk slices are running idle without receiving the above data query request; for such data query requests, first route the content ID carried in the data query request according to the same rule (a distribution addressing strategy), directly query the target disk slice (such as the disk slice - 2 shown in the figure), reduce the load of N - 1 / N, and greatly shorten the data query time; at the same time, since it is an OLAP query provided, the data only needs to meet the eventual consistency. By separating read and write of the master-slave replicas, the performance can be further improved. At the same time, through caching for the same results, the performance of the external service can be significantly improved.

[0193] As an embodiment, after determining the data query index information corresponding to the data query request in step S703, the determined data query index can be cached for a preset duration, so that within the preset duration after determining the data query index, when the same data query request is received, the corresponding aggregated data can be directly obtained based on the cached data query index; the above preset duration is not limited, and those skilled in the art can set it according to actual needs, such as setting it to 1 minute, 5 minutes, 10 minutes, or 1 hour, etc.

[0194] According to the scenario of the data processing method for content push involved in the embodiments of the present application, most data query requests are related to time and content ID (i.e., the above-mentioned content identifier). For example, for a certain pushed content, query the feedback information of the target accounts that received the pushed content within N (N is an integer as mentioned above) minutes after its release. In the embodiments of the present application, a data query index can be established according to the date, preset time granularity, and content ID. After establishing the data query index for the query of a certain pushed content, nearly 99% of the file scans can be reduced; in addition, in some business scenarios, if the amount of data to be queried is too large and the dimensions are too many, taking videos as the pushed content, there are tens of billions of videos pushed in a day, and some dimensions have hundreds of categories. If all dimensions are pre-aggregated at one time, the amount of data will exponentially expand, the query will become slower instead, and a large amount of memory space will be occupied; the solution is to establish corresponding pre-aggregated views for different dimensions, trading space for time, so as to shorten the query time.

[0195] Please refer to Figure 9 In addition, the embodiments of the present application further provide a framework of a content push system, which includes: a content production end and a content consumption end, an up and down content interface server, a content database, a scheduling center service, an artificial review system, a duplicate elimination system, a content distribution outlet service, a data operation system, a real-time distribution statistics access layer, a real-time storage engine layer, a real-time computing layer, a real-time data aggregation model, and a real-time data monitoring and display service.

[0196] Such as Figure 9 The flowchart of the content distribution monitoring method and system based on real-time multi-dimensional aggregation calculation as shown. In the information flow scenario of massive content distribution in the information flow, when the amount of data is huge, reaching a scale of trillions in a day, achieving extremely low-latency real-time computing and second-level multi-dimensional real-time query and monitoring is the core problem to be addressed in the present invention. Efficient real-time data monitoring and analysis and efficient processing of aggregated data, the real-time distribution statistics data include: exposure, PV / VV, comments, negative comments, reports, negative feedback, etc.; the content that needs to be reviewed is pushed to the artificial review system through the data operation system, and the review result of the pushed content by the artificial review system is obtained to determine whether to continue to distribute the pushed content in the recommendation pool or disable or take down the above-mentioned pushed content; among them:

[0197] The real-time distribution statistics access layer can be used to monitor and analyze the statistics of the basic content interaction data uploaded by the target account and the content push system for distributing and pushing content. For example, statistical data on abnormal performance of push content at the content consumption end (also known as the C side), such as a rapid increase in comment data for push content, an overly rapid growth rate of PV / VV, an overly rapid growth rate of the number of times of forwarding push content, a rapid increase in the operation of liking push content, etc. Furthermore, after the statistical information of a certain push content is monitored and shown by the real-time data monitoring and display service and meets the abnormal conditions, the submission interface can be called to push the push content to the manual review system for manual review. For the push content that is confirmed to be abnormal at the content consumption end after review, the push content will be directly taken off the shelf;

[0198] In the embodiments of the present application, the statistical data for monitoring push content may but is not limited to including the data in Table 6, etc. The statistical data in Table 6 is only for illustrative purposes, and those skilled in the art can set the statistical data for monitoring push content according to actual needs.

[0199] Table 6: Examples of statistical data for monitoring push content

[0200] Rapid growth of PV / VV Rapid growth of comments Content report or negative feedback Negative comments High bounce Low duration Low reading completion rate Overall market top PV / VV Top PV / VV of sensitive categories (such as society) Overall market top conversion rate Top conversion rate of sensitive categories Overall market top comment volume Overall market top like volume Overall market top Biu volume

[0201] As an embodiment, the data processing method based on content push provided by the embodiments of the present application can be implemented through Figure 9 the real-time distribution statistics access layer, the real-time storage engine, the real-time computing layer, and the real-time data aggregation model in, where:

[0202] The real-time distribution statistics access layer mainly pushes a large amount of basic content interaction data to the corresponding message queues associated with the operation types. For the pushed content of a certain multimedia channel (such as video, etc.), after splitting, the data is only in the order of millions per second; the real-time computing layer and the real-time data aggregation model are mainly responsible for performing real-time account dimension association based on the account portrait set and content dimension association based on the content information set on the basic content interaction data obtained from each message queue, and converting multiple rows of basic content interaction data into one row of multi-column aggregated data, etc.; the real-time storage engine mainly designs a real-time message queue that meets the target business requirements and is easy to use downstream. In the embodiments of the present application, two message queues can be provided as two layers of the real-time storage engine. One is the message queue of the Data WareHouse Middle (DWM) layer. The message queue of the DWM layer stores a series of intermediate tables generated by performing a light aggregation operation on the account portrait data in the account portrait set and the content information in the content information set. The DWM layer is used to improve the reusability of public metrics and reduce repeated processing; the other is the message queue of the Data WareHouse Servce (DWS) layer. The DWS layer stores the first aggregated data obtained by account dimension association and the second aggregated data obtained by content dimension association. The data therein can but is not limited to include content identifiers, B-side data, and C-side data; it can be seen that the traffic of the DWS layer is further reduced to the order of one hundred thousand per second, and the aggregated data of the content identifiers is even in the order of ten thousand per second, and the format is clearer and the dimension information is richer. Finally, a data query function can be provided through the above real-time data monitoring and display service; where the B-side data refers to the data or information associated with the pushed content created by the content producer, which can but is not limited to include the content information of the pushed content (also called content meta-information); the C-side data can include the basic content interaction data involved in the embodiments of the present application.

[0203] Next, the real-time distribution statistics access layer, the real-time storage engine, the real-time computing layer, the implementation of the real-time data aggregation model, and the real-time data monitoring and display service will be further described:

[0204] Among them, the real-time distribution statistics access layer is mainly responsible for accessing basic content interaction data, realizing the real-time access of a large amount of basic content interaction data per second, and performing extremely low-latency association with the association information set; at the same time, the real-time distribution statistics access layer interacts with the real-time storage engine to support high-concurrency writing, highly available distributed, and high-performance data query indexing. This is the key to reducing the delay of data from the hour level to the minute level. The logical relationship between the real-time distribution statistics access layer, the real-time computing layer, and the real-time storage engine can be seen in Figure 9 ;

[0205] Such as Figure 10As shown, the key to the real-time distribution statistics access layer is to push the basic content interaction data in the original message queue to the message queues associated with the operation types of the interaction operations that generate the basic content interaction data respectively. As shown in the figure, the basic content interaction data in the original message queue is split into micro queues and micro processes are deployed, which can speed up the efficiency of associating the basic content interaction data in each message queue at the account dimension and the content dimension; the real-time distribution statistics access layer externally has several message queues associated with the operation types of the above interaction operations, and different message queues store basic content interaction data with different aggregation granularities, including content identifiers (such as content IDs), account identifiers (such as account IDs), C-side data, B-side data, and account portrait data, etc.; the real-time storage engine stores the aggregated data output by the real-time computing layer above (which can include but is not limited to at least one of the above first aggregated data and second aggregated data), and saves the aggregated data output by the real-time computing layer to the message queue in the DWS layer to provide downstream multi-account reuse; thus, before using the method provided by the embodiments of the present application. It is necessary to perform complex data cleaning on the basic content interaction data obtained from the original message queue of tens of millions per second, and then perform account-level association and information-level association to obtain aggregated data in the required format, and the processing efficiency is very low. However, after adopting the method provided by the embodiments of the present application, it is possible to directly apply for the aggregated data stored in the message queue of the DWS layer based on the content identifier, and improve the speed of responding to data query requests for the content to be pushed for query.

[0206] Please continue to refer to Figure 10 , the real-time storage engine needs to have dimension indexes, support high concurrency, pre-aggregation, and high-performance real-time multi-dimensional OLAP queries; in the embodiments of the present application, the real-time storage engine can achieve the functional requirements of distributed high availability and horizontal scalability, and can also write massive aggregated data into the corresponding disk shards. The process of writing the aggregated data into the disk shards can be referred to the above content and will not be repeated here; the real-time storage engine can also perform high-performance queries, such as constructing the above data query indexes, and obtaining query data materialized view construction based on the data query indexes, etc. The relevant content of the data query indexes can be referred to the above description and will not be repeated here.

[0207] As an embodiment, in the embodiments of the present application, the real-time storage engine can be but is not limited to being divided into a real-time writing layer, an OLAP storage layer, and a background interface layer; among them, the real-time writing layer is mainly responsible for writing the aggregated data into the corresponding disk shards through Hash routing; the OLAP storage layer uses the MPP storage engine to design indexes and views that conform to the business and efficiently store massive aggregated data; the background interface layer is used for direct query and retrieval, and is used as a data service interface to access data, providing an efficient multi-dimensional real-time query interface.

[0208] As an embodiment, the real-time computing layer and the real-time data aggregation model are used to perform account dimension association and content dimension association on the basic content interaction data in each message queue. Specifically, refer to Figure 11 , which shows the relationship between the real-time computing layer and the real-time data aggregation model, as well as the schematic diagram of the main processes for performing account dimension association and content dimension; as shown in the figure, during the process of the real-time computing layer, based on the real-time dimension table (the above-mentioned account portrait set or content information set), first, the basic content interaction data obtained from each message queue is window-aggregated according to a preset time window (i.e., the above-mentioned first preset time window or second preset time window) to obtain an interaction data set (i.e., the above-mentioned first interaction data set or second interaction data set), and the basic content interaction data in the same interaction data set is aggregated to obtain aggregated data (i.e., the above-mentioned first aggregated data or second aggregated data). The specific process of obtaining the aggregated data can be referred to the above description and will not be repeated here.

[0209] Please continue to refer to Figure 11 , as an embodiment, in the processing of the real-time computing layer and the real-time data aggregation model, Redis cache is used as the first key-value pair database for storing the account portrait set, and Redis cache is used as the second key-value pair database for storing content information, which can accelerate the access speed of account portrait data and content information; at the same time, the cache consistency is ensured by listening to write HBase Proxy. The detailed content of the account portrait set, content information set, first key-value pair database, and second key-value pair database can be referred to the above description and will not be repeated here.

[0210] The above real-time data monitoring and display service is used to respond to the above data query request, create data query index information corresponding to the data query request, obtain the corresponding aggregated data based on the above data index information, perform data processing on the obtained aggregated data according to the data processing rules associated with the target business requirements, obtain the query data, and return the above query data. The specific process can be referred to the above content and will not be repeated here.

[0211] Next, the functions of each module in the above content push system will be introduced:

[0212] 1) Content production end and content consumption end

[0213] PGC, UGC, or PUGC are content producers of a Multi-Channel Network (MCN). Through a mobile terminal or a backend interface API system, they provide graphic and text content or video content provided by a local editing system or a web publishing system as push content. The video content includes short videos and mini-videos, which are the main content sources for recommended distribution. The above MCN is a product form of a multi-channel network that combines PGC content. With the strong support of capital, it ensures the continuous output of content, and ultimately realizes stable commercial monetization.

[0214] Among them, the content production side obtains the interface address of the upstream and downstream content interface server through communication with it, and uploads graphic and text content or video content through the interface address. The source of graphic and text content is usually a lightweight publishing end and an editing content entry. The release of video content is usually an image acquisition device. During the shooting process, the local video content can be paired with music, filter templates, and video beautification functions, etc.

[0215] The content consumption side communicates with the upstream and downstream content interface server to obtain the index information of the push content, and the index information is displayed in the form of a Feeds stream. When the content consumption side sends a request message for specific graphic and text content or video content, the content consumption side communicates with the content distribution outlet service to obtain the corresponding graphic and text content or video content in the index information.

[0216] In addition, the content consumption side can also report the basic content interaction data determined by the interaction operations triggered by the target account for the push content (such as but not limited to including information such as comments, likes, forwards, collections, views, bounces, plays, exposures, etc. of the target account for the push content) to the statistical reporting interface server in real time for statistical analysis. For example, lag, loading time, play clicks, etc.

[0217] 2) Upstream and downstream content interface server and content distribution outlet service

[0218] The upstream and downstream content interface server communicates directly with the content production side, stores the content meta-information of the push content submitted by the content production side in the content database, and synchronizes the push content submitted by the content production side to the scheduling center server for processing and circulation of the push content. Among them, the description of the content meta-information can be referred to the above content and will not be repeated here.

[0219] The content distribution outlet service sends the obtained push content to the content consumption side in the form of Feeds. The content distribution outlet service is usually a group of access services deployed nearby geographically close to the user.

[0220] 3) Content database

[0221] The content database is the core database for pushing content. All content source information of the pushed content published by the content production side is stored in the content database. That is, the content database in the embodiments of the present application can store the content meta-information of the pushed content generated by the content production side. For the description of the content meta-information, please refer to the above content and will not be repeated here.

[0222] Content processing mainly includes machine processing and manual review processing. For content feature modeling services, it is necessary to obtain the content source information of the pushed content from the content database. According to different content tags, the content database is divided into different content pools. Recommendation ranking services, duplicate elimination services, etc. all need to obtain the content information of the pushed content from the content database. For example, the duplicate elimination service will load the pushed content that has been warehoused and enabled in the past period (such as one week) according to business requirements. For the pushed content that is re-warehoused repeatedly, a filtering mark will be added and it will no longer be provided to the content distribution outlet service for display to users. Among them, for graphic and text-based pushed content, both the graphic and text duplicate elimination service and the high-quality content recognition service are machine processing processes, and the processing results are stored in the content database.

[0223] As an embodiment, during the manual review process, the content meta-information of the pushed content can be read from the content database, and at the same time, the results and status of the manual review will also be transmitted back and saved in the content database. The manual review results are also an important basis for measuring the efficiency of the subsequent algorithm filtering model.

[0224] 4) Scheduling Center Service

[0225] The scheduling center service is responsible for the entire scheduling process of the pushed content circulation, controlling the upstream and downstream content interface servers to receive the uploaded pushed content, and obtaining the content meta-information of the pushed content from the content database; scheduling the duplicate elimination service to mark and filter the pushed content that is warehoused repeatedly;

[0226] The scheduling center service can also, for the content that cannot be processed by machines, such as content with sensitive and security issues that requires manual review, call the manual review system for manual review processing; finally, the pushed content enabled by the manual review system is distributed through the content export service, usually provided to the terminal content consumer through the recommended engine or search engine or the direct display page of the operation.

[0227] As an embodiment, the scheduling center service can also communicate with the real-time data monitoring and display service to obtain the real-time statistics and monitoring data of the end distribution for adjusting the scheduling strategy. For example, the category data with good consumption data posterior enters the head of the review scheduling queue first.

[0228] 5) Manual Review System

[0229] The manual review system is the carrier of manual service capabilities. It is mainly used to review and filter sensitive, pornographic, and content that cannot be determined by machines, such as those not allowed by law. At the same time, it also performs label annotation and secondary confirmation of the content. The content to be reviewed comes from what is published by the self-media application and obtained from the public network. The results of the manual review are written into the content database through the dispatching center service. Since the push content of the graphic and text type is not yet fully mature through machine learning such as deep learning, secondary manual review processing is carried out on the machines processed by the machine. Through human-machine collaboration, the accuracy and efficiency of the annotation of the graphic and text itself are improved.

[0230] As an embodiment, the manual review system can also receive, while receiving the review tasks synchronized by the dispatching center, the suspicious push content whose data changes are monitored abnormally by the data operation system during synchronization. After rechecking and passing this part of the push content, it can be directly taken off the shelf or continue to be distributed, etc.

[0231] 6) Duplicate removal service

[0232] The duplicate removal service provides duplicate removal services for push content of graphic and text types, videos, and picture sets. It mainly vectorizes the graphics, pictures, and videos, then builds an index of the vectors, and then determines the similarity by comparing the distances between the vectors. Specifically, the duplicate removal service communicates with the dispatching center service for title duplicate removal, picture duplicate removal of the cover image, content text duplicate removal, and video fingerprint and audio fingerprint duplicate removal, etc. Simhash or BERT can be used to remove duplicates of the text vectors and picture vectors. For video content, video fingerprints and audio fingerprints can be extracted to construct vectors, and then the distances between the vectors (such as Euclidean distance) are calculated to determine whether there are duplicates. The specific duplicate removal method is not elaborated in the embodiments of this application.

[0233] 7) Real-time distribution statistics access layer

[0234] The real-time distribution statistics access layer communicates with the content consumption end. For example, information such as comments, likes, forwards, collections, views, bounces, plays, and exposures of the content is reported through the real-time statistics interface service. According to the functions of the real-time distribution statistics access layer and the data access strategy, real-time data access and preprocessing are realized. The functions of the time distribution statistics access layer can be referred to the above description.

[0235] 8) Real-time calculation layer and data aggregation model

[0236] The core of the real-time calculation layer and the data aggregation model layer is to handle the relationship between the real-time associated information set and the real-time data warehouse. According to the detailed strategies and solutions described above, the relationship between the access layer and the data storage engine is processed to improve the efficiency of data processing and calculation, and reduce the resource consumption of data processing and calculation, etc.

[0237] 9) Real-time storage engine layer

[0238] The real-time storage engine layer can achieve distributed and highly available horizontal expansion, and can store the aggregated data obtained by the real-time computing layer and the data aggregation model into the corresponding disk slices. For specific content, please refer to the above description and will not be repeated here.

[0239] 10) Real-time data monitoring and display service

[0240] Implementing the data monitoring and display service can store the obtained aggregated data, service-ify the calculation results of the aggregated data, provide real-time data display and external services; and in response to a data query request, create a corresponding data query index and return the query result. For the specific process, please refer to the above content and will not be repeated here.

[0241] The real-time data monitoring and display service can also monitor the basic content interaction data of the push content by the content consumption side, which can but is not limited to monitoring the statistical data of the push content with abnormal performance on the C side. For the description of the statistical data, please refer to the above content and will not be repeated here.

[0242] It should be noted that the above application scenarios are only examples and do not constitute a limitation on the protection scope of this application.

[0243] In the embodiments of this application, on the one hand, it can process the real-time obtained basic content interaction data, improving the timeliness of processing the basic content interaction data; on the other hand, in the process of querying data related to the push content in the embodiments of this application, a corresponding data query index can be established for the data query request, and relevant aggregated data can be directly obtained based on the data query index for processing, reducing the data query latency. Moreover, the aggregated data in the embodiments of this application can be data after user dimension association and content dimension association. In the process of processing the obtained aggregated data, the consumption of computing power can be reduced, thereby improving the efficiency of obtaining query data based on the aggregated data; it can discover the push content with abnormal performance at the content consumption end in the first time. For the real-time data analysis scenario, the speed of data query response is significantly improved, and the latency of returning query data is significantly reduced.

[0244] Please refer to Figure 12 , based on the same inventive concept, the embodiments of this application provide a data processing device 1200 for content push, including:

[0245] A data acquisition unit 1201, configured to obtain basic content interaction data in response to interaction operations triggered by each target account for the push content obtained by each of them;

[0246] The data splitting unit 1202 is configured to, based on the operation type of the above interaction operation, push the above basic content interaction data to at least one message queue associated with the operation type in real time;

[0247] The data aggregation unit 1203 is configured to, for each of the at least one message queue, perform data association operations respectively in the following manner to obtain corresponding aggregated data: based on at least one set of association information, perform an association information set association operation on a message queue to obtain corresponding aggregated data; wherein, one association information set association operation includes: based on a set of association information, convert the basic content interaction data obtained by the above message queue within a preset time window into aggregated data, and the number of the converted aggregated data is not greater than the number of the basic content interaction data received within the preset time window.

[0248] As an embodiment, each basic content interaction data includes an account identifier associated with the target account that triggers the above interaction operation and a content identifier associated with the push content that triggers the above interaction operation. The data aggregation unit 1203 is specifically configured to perform any one or a combination of the following operations:

[0249] If the above set of association information includes a set of account portraits, determine the basic content interaction data obtained by the above message queue within the first preset time window as a first interaction data set, and for each account identifier included in the basic content interaction data in the first interaction data set, perform the following operations respectively: based on the account portrait data of each target account recorded in the above set of account portraits, obtain the account portrait data of the target account associated with an account identifier, and through the obtained account portrait data, aggregate the basic content interaction data including the above account identifier in the first interaction data set to obtain a first aggregated data;

[0250] If the above set of association information includes a set of content information, determine the basic content interaction data obtained by the above message queue within the second preset time window as a second interaction data set, and for each content identifier included in the basic content interaction data in the second interaction data set, perform the following operations respectively: based on the content information of each push content recorded in the above set of content information, obtain the content information of the push content associated with a content identifier, and through the obtained content information, aggregate the basic content interaction data including the above content identifier in the second interaction data set to obtain a second aggregated data.

[0251] As an embodiment, the data aggregation unit 1203 is specifically configured to:

[0252] Determine the basic content interaction data including the above account identifier in the first interaction data set;

[0253] Obtain the content identifiers included in the determined basic content interaction data, and generate a set of content identifiers;

[0254] Associate the obtained account portrait data, the above-mentioned one account identifier, and the above-mentioned set of content identifiers to obtain a corresponding first aggregated data.

[0255] As an embodiment, the data aggregation unit 1203 is specifically configured to:

[0256] Determine the basic content interaction data including the above-mentioned one content identifier in the above-mentioned second interaction data set;

[0257] Obtain the account identifiers included in the determined basic content interaction data, and generate a set of account identifiers;

[0258] Associate the obtained content information, the above-mentioned one content identifier, and the above-mentioned set of account identifiers to obtain a corresponding second aggregated data.

[0259] As an embodiment, the above-mentioned set of account portraits is stored in a first key-value pair database, and the first key-value pair database is updated periodically based on a first period;

[0260] The above-mentioned set of content information is stored in a second key-value pair database, and the second key-value pair database is obtained by periodically backing up the content database based on a second period. The content database is used to record the content information of the above-mentioned push content in real time.

[0261] As an embodiment, the data aggregation unit 1203 is further configured to:

[0262] For each message queue in the above-mentioned at least one message queue, respectively perform data association operations in the following manner. After obtaining the corresponding aggregated data, respectively perform data storage operations on the obtained aggregated data in the following manner, and store the above-mentioned respective aggregated data into the corresponding disk slices: Store an aggregated data into the disk slice mapped to the push content associated with the above-mentioned one aggregated data, where one disk slice is used to store the aggregated data associated with one push content; and

[0263] In response to a data query request for a push content to be queried, obtain the aggregated data associated with the push content to be queried from the disk slice mapped to the push content to be queried;

[0264] Perform data processing on the obtained aggregated data according to the data processing rules associated with the target business requirements to obtain query data, and return the above-mentioned query data.

[0265] As an embodiment, the content identifier of the to-be-query pushed content and the query time period are carried in the above data query request. Specifically, the data aggregation unit 1203 is configured to:

[0266] Determine the disk shard mapped by the to-be-query pushed content as the target disk shard based on the above content identifier; and

[0267] Divide the above query time period into at least one sub-time period by a preset time granularity;

[0268] Determine the data query index information corresponding to the above data query request according to the storage addresses of the aggregated data mapped by each sub-time period in the above target disk shard in the aggregated data associated with the to-be-query pushed content;

[0269] Based on the above data index information, obtain the aggregated data mapped by each sub-time period from the storage addresses of the aggregated data mapped by each sub-time period in the above target disk shard. As an embodiment, Figure 12 The device in can be used to implement any one of the data processing methods based on content push described above.

[0270] Based on the same inventive concept as the above method embodiment, an embodiment of the present application also provides a computer device. This computer device can be used for data processing based on pushed content. In one embodiment, the computer device can be a server, such as Figure 1 the data processing server 130 shown. In this embodiment, the structure of the computer device can be as Figure 13 shown, including a memory 1301, a communication module 1303, and one or more processors 1302.

[0271] The memory 1301 is used to store the computer program executed by the processor 1302. The memory 1301 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and programs required to run the instant messaging function, etc.; the data storage area can store various instant messaging information and operation instruction sets, etc.

[0272] The memory 1301 can be a volatile memory, such as a random-access memory (RAM); the memory 1301 can also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1301 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1301 can be a combination of the above memories.

[0273] The processor 1302 can include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1302 is used to implement the above-described content push-based data processing method when calling the computer program stored in the memory 1301.

[0274] The communication module 1303 is used to communicate with the terminal device and other servers.

[0275] In the embodiments of the present application, the specific connection medium between the above-mentioned memory 1301, communication module 1303, and processor 1302 is not limited. In the embodiments of the present disclosure Figure 13 it is shown that the memory 1301 and the processor 1302 are connected through a bus 1304, and the bus 1304 is shown as a thick line in Figure 13 The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus 1304 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 13 only a thick line is shown in

[0276] The memory 1301 stores a computer storage medium, and the computer storage medium stores computer-executable instructions for implementing the content recommendation method of the embodiments of the present application. The processor 1302 is used to execute the above-described content push-based data processing method.

[0277] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned storage medium includes: removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc., all kinds of media that can store program codes.

[0278] Alternatively, if the above-mentioned integrated unit is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention, in essence, or the part that makes contributions to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the above methods of the various embodiments of the present invention. And the aforementioned storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical disks, etc., all kinds of media that can store program codes.

[0279] Based on the same inventive concept, an embodiment of the present application also provides a computer-readable storage medium that stores computer instructions. When the above computer instructions run on a computer, the computer is enabled to execute the content push-based data processing method described above.

[0280] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0281] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A data processing method based on content push, characterized in that Including: In response to interaction operations triggered by respective target accounts for the push content obtained by each of them, obtaining corresponding basic content interaction data; wherein, the basic content interaction data includes: an account identifier associated with the target account that triggers the corresponding interaction operation and a content identifier associated with the push content that triggers the interaction operation; Based on the operation type of the interaction operation, pushing the corresponding basic content interaction data to at least one message queue associated with the operation type in real time; For each of the at least one message queue, respectively performing data association operations in the following manner to obtain corresponding aggregated data: Based on at least one set of association information, performing an association information set association operation on a message queue to obtain corresponding aggregated data; wherein, the aggregated data includes first aggregated data, and one association information set association operation includes: If a set of association information includes an account profile set, determining the basic content interaction data obtained by the message queue within a first preset time window as a first interaction data set; wherein, the account profile set contains account profile data of each target account, and the account profile data is determined by the interaction operations triggered by the corresponding target accounts for the push content; For each account identifier included in the basic content interaction data in the first interaction data set, respectively performing the following operations: obtaining the account profile data of the target account associated with an account identifier from the account profile set, and determining the basic content interaction data in the first interaction data set that contains the account identifier, generating a content identifier set from the content identifiers included in the determined basic content interaction data, and associating the obtained account profile data, the account identifier, and the content identifier set to obtain a first aggregated data; Wherein, the number of the obtained first aggregated data is not greater than the number of the basic content interaction data received within the first preset time window.

2. The method according to claim 1, characterized in that, The aggregated data further includes second aggregated data, and the one association information set association operation further includes: If a set of association information includes a content information set, determining the basic content interaction data obtained by the message queue within a second preset time window as a second interaction data set, and for each content identifier included in the basic content interaction data in the second interaction data set, respectively performing the following operations: based on the content information of each push content recorded in the content information set, obtaining the content information of the push content associated with a content identifier, and aggregating the basic content interaction data in the second interaction data set that contains the content identifier through the obtained content information to obtain a second aggregated data; Wherein, the number of the obtained second aggregated data is not greater than the number of the basic content interaction data received within the second preset time window.

3. The method according to claim 2, wherein The aggregating the basic content interaction data in the second interaction data set that contains the content identifier through the obtained content information to obtain a second aggregated data includes: Determine each basic content interaction data in the second interaction data set that contains the one content identifier; Obtain each account identifier included in the determined basic content interaction data, and generate an account identifier set; Associate the obtained content information, the one content identifier, and the account identifier set to obtain a corresponding second aggregated data.

4. The method according to claim 2, wherein The account profile set is stored in a first key-value database, and the first key-value database is periodically updated based on a first period; The content information set is stored in a second key-value database, and the second key-value database is obtained by periodically backing up a content database based on a second period, and the content database is used to record the content information of the push content in real time.

5. The method according to claim 1, wherein For each message queue in the at least one message queue, respectively perform a data association operation in the following manner. After obtaining the corresponding aggregated data, it further includes: For each obtained aggregated data, respectively perform a data storage operation in the following manner, and store each aggregated data into a corresponding disk slice: Store an aggregated data into the disk slice mapped by the push content associated with the one aggregated data, where one disk slice is used to store the aggregated data associated with one push content; And In response to a data query request for a to-be-query push content, obtain the aggregated data associated with the to-be-query push content from the disk slice mapped by the to-be-query push content; Perform data processing on the obtained aggregated data according to the data processing rules associated with the target business requirements to obtain query data, and return the query data.

6. The method according to claim 5, wherein The content identifier and query time period of the to-be-query push content are carried in the data query request. The obtaining of the aggregated data associated with the to-be-query push content from the disk slice mapped by the to-be-query push content includes: Based on the content identifier, determine the disk slice mapped by the to-be-query push content as the target disk slice; and Divide the query time period into at least one sub-time period by a preset time granularity; According to the storage addresses of the aggregated data mapped by each sub-time period in the target disk slice in the aggregated data associated with the to-be-query push content, determine the data query index information corresponding to the data query request; Based on the data query index information, obtain the aggregated data mapped by each sub-time period from the storage addresses of the aggregated data mapped by each sub-time period in the target disk slice.

7. A data processing device based on content push, characterized in that, Include: A data collection unit, configured to obtain corresponding basic content interaction data in response to interaction operations triggered by each target account for the push content obtained by each; wherein, the basic content interaction data includes: the account identifier associated with the target account that triggers the corresponding interaction operation and the content identifier associated with the push content that triggers the interaction operation; A data splitting unit, configured to based on the operation type of the interaction operation, push the corresponding basic content interaction data to at least one message queue associated with the operation type in real time; A data aggregation unit is configured to perform data association operations on each of the at least one message queues respectively in the following manner to obtain corresponding aggregated data: Based on at least one set of association information, perform an association information set association operation on a message queue to obtain corresponding aggregated data; wherein, one association information set association operation includes: If a set of association information includes an account profile set, determine each piece of basic content interaction data obtained by the message queue within a first preset time window as a first interaction data set; wherein, the account profile set contains account profile data of each target account, and the account profile data is determined by interaction operations triggered by the corresponding target account for push content; For each account identifier included in each piece of basic content interaction data in the first interaction data set, perform the following operations respectively: Obtain the account profile data of the target account associated with an account identifier from the account profile set, and determine each piece of basic content interaction data in the first interaction data set that contains the account identifier. Generate a content identifier set from the content identifiers included in the determined pieces of basic content interaction data. Associate the obtained account profile data, the account identifier, and the content identifier set to obtain a first aggregated data; Wherein, the number of the obtained first aggregated data is not greater than the number of the basic content interaction data received within the first preset time window.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions run on a computer, the computer is caused to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Content aggregation method, server, client and system

    CN104683370A

  • Time series data aggregate query method and apparatus, computer device and readable medium

    CN108268589A

  • Service object pushing method and device

    CN111090822A

  • Information pushing method and device, electronic equipment and storage medium

    CN111259246A