Information pushing method based on converged media data

By adaptively extracting and standardizing converged media data, generating a thematic database and performing load analysis, the problems of large storage and transmission costs and high latency in traditional information push methods are solved, enabling rapid information push and system optimization in high-concurrency scenarios.

CN121858758APending Publication Date: 2026-04-14BEIJING ZHONGKEZHI MEDIA TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202610314174.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-16
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional information push methods lack forward-looking prediction and fine-grained scheduling of system load in high-concurrency, multi-source heterogeneous media data scenarios, resulting in large storage and transmission consumption, high latency, and even system lag and overload.

Method used

By adaptively extracting converged media data, generating a basic database table set, performing standardized processing, and then fusing thematic data to generate a thematic database, information is pushed using traffic shaping and priority queuing strategies based on user profiles and load analysis results.

Benefits of technology

In high-concurrency, multi-source heterogeneous media convergence scenarios, reduce storage and transmission overhead, lower computing resource consumption, and improve system response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858758A_ABST
    Figure CN121858758A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an information pushing method based on converged media data. A specific embodiment of the method comprises the following steps: in response to a detected access operation of a target user on target convergence media data and the number of access requests of the target convergence media data exceeds a preset access threshold, extracting the target convergence media data to obtain a basic library table set; standardizing the basic library table set to obtain a standard library table set; generating a thematic database based on the historical access log; generating a user portrait based on the basic data and the behavior data of the target user; generating a target content set based on the user portrait; and generating a load analysis result based on the time sequence prediction model and the historical access log, and pushing the target content set to a display device of the target user based on the network state information and the load analysis result. According to the embodiment, the technical effects of reducing storage and transmission occupancy, reducing computing resource consumption and increasing system response speed are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and specifically to an information push method based on converged media data. Background Technology

[0002] Converged media data is characterized by its multi-source heterogeneity and continuous massive growth. Traditional information push methods suffer from insufficient real-time performance and low system resource utilization when dealing with high-concurrency, multi-source heterogeneous converged media data scenarios. Currently, when pushing information based on converged media data, the common approach is to set push priorities based on fixed attributes of the content (such as source and publication time) or simple popularity indicators (such as click-through rate), lacking dynamic response to real-time system load and network conditions.

[0003] However, when using the above methods to push information based on converged media data, the following technical problems often arise: The system suffers from high storage and transmission costs, high latency, and may even cause system lag and overload. When dealing with high-concurrency, multi-source, heterogeneous converged media data scenarios, existing methods lack forward-looking prediction and fine-grained scheduling of the overall load. They fail to combine network status information (such as bandwidth and latency) with load analysis results to dynamically adjust traffic shaping strategies and priority queues, resulting in high system storage and transmission costs, high latency, and may even cause system lag and overload.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure propose information push methods, apparatuses, and devices based on converged media data to solve one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide an information push method based on converged media data. The method includes: in response to detecting that a target user accesses target converged media data and the number of access requests for the target converged media data exceeds a preset access threshold, adaptively extracting the target converged media data to obtain a basic database table set; standardizing the basic database table set to obtain a standard database table set; performing thematic data fusion processing on the standard database table set based on historical access logs and incremental update methods to generate a thematic database; generating a user profile based on the target user's basic data and behavioral data; performing collaborative filtering on the thematic database based on the user profile to generate a target content set; performing load analysis on the thematic database and the user profile based on a time series prediction model and the historical access logs to obtain load analysis results; and pushing the target content set to the target user's display device using traffic shaping and priority queuing strategies based on network status information and the load analysis results.

[0008] Secondly, some embodiments of this disclosure provide an information push device based on converged media data. The device includes: a detection unit configured to, in response to detecting that a target user accesses target converged media data and the number of access requests for the target converged media data exceeds a preset access threshold, adaptively extract the target converged media data to obtain a basic database table set; a standardization processing unit configured to perform standardization processing on the basic database table set to obtain a standard database table set; a topic data fusion processing unit configured to perform topic data fusion processing on the standard database table set based on historical access logs and incremental update methods to generate a topic database; a generation unit configured to generate a user profile based on the target user's basic data and behavioral data; a collaborative filtering unit configured to perform collaborative filtering on the topic database based on the user profile to generate a target content set; and a load analysis unit configured to perform load analysis on the topic database and the user profile based on a time series prediction model and the historical access logs to obtain load analysis results, and push the target content set to the target user's display device based on network status information and the load analysis results, using traffic shaping and priority queuing strategies.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] The above embodiments of this disclosure have the following beneficial effects: By constructing a full-link optimization mechanism of "data perception-intelligent processing-dynamic scheduling" through the information push method based on converged media data in some embodiments of this disclosure, the storage and transmission occupancy is reduced, computing resource consumption is lowered, and system response speed is accelerated in high-concurrency, multi-source heterogeneous converged media scenarios. Specifically, the reason for high system storage and transmission occupancy, high latency, and even potential system lag and overload is that existing methods lack forward-looking prediction and fine-grained scheduling of the overall load when dealing with high-concurrency, multi-source heterogeneous converged media data scenarios. They fail to combine network status information (such as bandwidth and latency) with load analysis results to dynamically adjust traffic shaping strategies and priority queues, resulting in high system storage and transmission occupancy, high latency, and even potential system lag and overload. Based on this, the information push method based on converged media data in some embodiments of this disclosure firstly, in response to detecting that a target user's access operation to target converged media data, and that the number of access requests for the target converged media data exceeds a preset access threshold, adaptively extracts the target converged media data to obtain a basic database table set. This allows for the rapid filtering of high-value, highly relevant data subsets from massive multi-source data, avoiding scanning the entire dataset from the source and significantly reducing the data scale for subsequent processing. Secondly, the basic database table set is standardized to obtain a standard database table set. This eliminates inconsistencies in the original data's dimensions, formats, and encoding, providing unified and well-organized high-quality data input for subsequent fusion analysis and reducing the computational complexity of data cleaning and transformation processes. Then, based on historical access logs and incremental update methods, the standard database table set undergoes thematic data fusion processing to generate a thematic database. This dynamically integrates related data, forming a data aggregation view with distinct themes, avoiding repeated related queries on the same data, and improving data retrieval and utilization efficiency. Subsequently, user profiles are generated based on the target user's basic data and behavioral data. This allows for the construction of accurate user interest models, providing a basis for personalized recommendations and reducing the distribution of invalid content and the waste of computational resources. Next, based on the aforementioned user profiles, collaborative filtering is performed on the aforementioned thematic database to generate a target content set. This enables rapid and accurate screening and sorting of candidate content, significantly narrowing the data scope for refined sorting operations and reducing the real-time processing pressure on the system. Finally, based on the time-series prediction model and the aforementioned historical access logs, load analysis is performed on the aforementioned thematic database and the aforementioned user profiles. The load analysis results, along with network status information and the aforementioned load analysis results, are used to push the aforementioned target content set to the display devices of the aforementioned target users using traffic shaping and priority queuing strategies. This allows for proactive prediction and adaptive control of system load and network conditions, ensuring the supply of resources for critical tasks, optimizing overall throughput, and reducing response latency.This implementation achieves the technical effects of reducing storage and transmission overhead, lowering computing resource consumption, and accelerating system response speed in high-concurrency, multi-source heterogeneous converged media scenarios. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 This is a flowchart of some embodiments of the information push method based on converged media data disclosed herein; Figure 2 This is a schematic diagram of the structure of some embodiments of the information push device based on converged media data according to the present disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] Figure 1 A flowchart 100 is shown, illustrating some embodiments of the information push method based on converged media data according to this disclosure. This information push method based on converged media data includes the following steps: Step 101: In response to detecting that the target user has accessed the target converged media data and that the number of access requests for the target converged media data exceeds the preset access threshold, adaptive extraction of the target converged media data is performed to obtain the basic library table set.

[0021] In some embodiments, the execution entity (e.g., a server) of the information push method based on converged media data can adaptively extract the target converged media data to obtain a basic database table set in response to detecting a target user's access operation to target converged media data and the number of access requests for the target converged media data exceeding a preset access threshold. The target converged media data can be a data set collected from multiple heterogeneous sources and containing various media formats (such as text, images, audio, and video). For example, the target converged media data may include, but is not limited to: news information, social media posts (Weibo, Toutiao), short video content, blog articles, local government news, local public service notices, and regional policy and regulatory interpretations. The preset access threshold can be 1000 access requests per minute.

[0022] In some optional implementations of certain embodiments, the aforementioned execution entity may, in response to detecting that a target user's access operation to the target converged media data is exceeded and the number of access requests to the target converged media data exceeds a preset access threshold, adaptively extract the target converged media data to obtain a basic library table set: The first step involves parsing the target media data to obtain a multimodal dataset in response to the detection of a target user's access operation to the target media data, where the number of access requests exceeds a preset access threshold. In practice, the executing entity, upon detecting a target user's access operation to the target media data and the number of access requests exceeding the preset access threshold, firstly performs entity recognition on the text included in the target media data using Named Entity Recognition (NER) technology to obtain a set of named entities (such as personal names, place names, and organization names). Then, based on Convolutional Neural Networks (CNNs), it extracts image features from the images included in the target media data to obtain a set of image feature vectors. Secondly, based on Recurrent Neural Networks (RNNs), it extracts audio features from the audio included in the target media data to obtain a set of audio feature sequences. Next, the video included in the target media data is treated as a combination of image sequences and audio tracks. The CNN is applied to the image sequences, and the RNN is applied to the audio tracks to obtain a video multimodal feature set. Finally, the aforementioned named entity set, image feature vector set, audio feature sequence set, and video multimodal feature set are defined as the multimodal dataset. As an example, the aforementioned convolutional neural network can be ResNet or YOLO. The aforementioned recurrent neural network can be a Long Short-Term Memory (LSTM) network or a Gate Recurrent Unit (GRU).

[0023] The second step is to clean the multimodal dataset to obtain a cleaned dataset. In practice, the executing entity can perform data cleaning on the multimodal dataset, including missing value handling, outlier handling, duplicate value handling, and data format standardization, to obtain a cleaned dataset. Missing value handling can be done by filling in missing values ​​with the average value. Outlier handling can be done using techniques such as box plot analysis, Z-Score methods, or DBSCAN clustering algorithms to identify outliers (data points that significantly deviate from the normal range and may be due to errors), and then delete outliers or correct them using the average value. Duplicate value handling can be done using fuzzy matching algorithms to identify and delete duplicate data. Data format standardization can include: standardizing dates to YYYY-MM-DD format, and standardizing the representation of mobile phone numbers and addresses, etc.

[0024] The third step is to integrate the cleaned dataset into the base database table set. In practice, the execution entity can use an ETL (Extraction Transformation Loading) tool to integrate the cleaned dataset into the base database table set.

[0025] Step 102: Standardize the basic library table set to obtain the standard library table set.

[0026] In some embodiments, the aforementioned execution entity may standardize the aforementioned basic library table set to obtain a standard library table set.

[0027] In some optional implementations of certain embodiments, the aforementioned execution entity may standardize the aforementioned basic library table set through the following steps to obtain a standard library table set: The first step involves smoothing outliers from the numeric fields in each base table of the aforementioned base table set, and standardizing the enumerated values ​​of the categorical fields in each base table set to obtain an initial table set. In practice, the execution entity can use the Z-Score smoothing method to smooth outliers from the numeric fields in each base table set, and use the one-hot encoding method to convert the categorical fields in each base table set into an encoding format that the execution entity can process, thus obtaining the initial table set.

[0028] The second step involves data anonymization of the personal information fields contained in each initial table in the aforementioned initial database set, resulting in a standard database set. In practice, the executing entity can use hash algorithms (such as MD5 or SHA-256) to replace or mask the personal information fields contained in each initial table in the aforementioned initial database set, thus obtaining the standard database set.

[0029] Step 103: Based on historical access logs and incremental update methods, perform thematic data fusion processing on the standard library table set to generate a thematic database.

[0030] In some embodiments, the aforementioned execution entity can perform thematic data fusion processing on the aforementioned standard library tables based on historical access logs and incremental updates to generate a thematic database. The aforementioned historical access logs can be text files automatically generated by the web server. These historical access logs record all requests initiated to the server in chronological order.

[0031] In some optional implementations of certain embodiments, the aforementioned execution entity may perform thematic data fusion processing on the aforementioned standard library tables based on historical access logs and incremental update methods to generate a thematic database through the following steps: The first step involves analyzing the historical access logs and standard library tables mentioned above using an association rule algorithm to obtain a high-frequency access data table set. The association rule algorithm can be either the Apriori algorithm or the FP-Growth algorithm.

[0032] The second step is to sort the aforementioned high-frequency access data tables by access popularity, resulting in a high-frequency access data table sequence. In practice, the executing entity can sort the aforementioned high-frequency access data tables in descending order of access popularity based on access frequency, thus obtaining a high-frequency access data table sequence. The access frequency can be calculated as (number of accesses) / (time elapsed since publication + 2)^G. Here, G represents a constant used to control the rate at which popularity decreases over time. For example, G can be 1.2 or 1.5.

[0033] The third step involves generating a set of topic definition files based on the preset domain knowledge graph and the aforementioned high-frequency access data table sequence. In practice, the executing entity first matches the high-frequency access data table sequence with the preset domain knowledge graph and identifies the topic name, basic table set, association key, and priority corresponding to the high-frequency access data table sequence to generate the topic definition file set. The topic definition files in the set include: topic name (e.g., "elderly care service subsidy policy"), basic table set (e.g., "policy basic information table," "elderly basic information table," "community service facility table"), association key (e.g., "policy number"), and priority. The preset domain knowledge graph can be a pre-defined structured semantic network that contains concepts (e.g., "elderly care services," "social assistance"), attributes (e.g., "publication date," "document number," "applicable objects," "subsidy standards"), and relationships (e.g., "policy-beneficiaries," "service institutions-providing-elderly care services") within a specific domain (e.g., public services).

[0034] The fourth step involves performing join queries and statistical aggregations on the standard library tables based on the aforementioned thematic definition file set to generate a wide table dataset. In practice, the execution entity can perform join queries (such as SQL JOIN operations) and statistical aggregations (such as SQL GROUP BY clauses and aggregate functions like SUM(), COUNT(), and AVG()) on the standard library tables based on the aforementioned thematic definition file set to generate a wide table dataset.

[0035] The fifth step is to perform feature vectorization on the wide table dataset to obtain the feature dataset. In practice, the execution entity can use methods such as TF-IDF, Word2Vec, or BERT to convert the text fields included in the wide table dataset into numerical vectors to obtain the feature dataset.

[0036] Step 6: Based on the unsupervised machine learning algorithm and the aforementioned feature dataset, determine the business label set of the feature dataset to obtain the label-enriched dataset. The unsupervised machine learning algorithm can be either K-Means or DBSCAN clustering algorithm.

[0037] Step 7: Analyze the aforementioned tag-enriched dataset and the aforementioned historical access logs to obtain the analysis results. In practice, the executing entity can determine the field types, data volume, and null value ratio of the aforementioned tag-enriched dataset, and analyze the ratio between point queries and analytical aggregation queries in the aforementioned historical access logs to obtain the analysis results.

[0038] Step 8: Based on the above analysis results, determine the storage format. In practice, the execution entity can determine the storage format based on a preset decision table and the above analysis results. The storage formats include row-based storage (such as Avro format, suitable for scenarios involving frequent whole-row read / write and updates) and column-based storage (such as Parquet and ORC formats, suitable for analysis scenarios involving large-scale scanning and aggregation of a few columns). For example, the preset decision table could be: if the ratio of point queries to analytical aggregation queries is >70%, the storage format is determined to be row-based; if the ratio is <30%, the storage format is determined to be column-based. If the ratio of point queries to analytical aggregation queries is >30% and <70%, and the data volume of the above-mentioned tag-enriched dataset exceeds 1TB, the storage format is determined to be column-based.

[0039] Step nine involves generating and executing a structured query statement based on the aforementioned storage format to create a physical table structure on the storage device. This structured query statement can be a SQL data definition language (DDL), such as a CREATE TABLE statement.

[0040] Step 10: Based on the data lake toolchain, incrementally update the aforementioned labeled enriched dataset into the aforementioned physical table structure to generate the thematic database. The data lake toolchain can be a suite of software for data ingestion, management, and analysis, including Apache Spark, Flink (for distributed processing), and Apache Atlas (for metadata management). The incremental updates can process only newly added or changed data.

[0041] Step 104: Generate user profiles based on the target users' basic data and behavioral data.

[0042] In some embodiments, the aforementioned executing entity may generate a user profile based on the target user's basic data and behavioral data. The aforementioned basic data may include, but is not limited to, user name, gender, age, residential address, employer, and work history. The aforementioned behavioral data may include, for example, historical clickstream data, search keywords, dwell time, and interaction feedback data.

[0043] In addressing the aforementioned technical challenges by employing technical solutions, and considering the application scenario—where a large number of users' attention is constantly changing within a short period (e.g., e-commerce platforms' "618" and "Double 11" promotions)—the following technical issues often arise: in a multi-source, heterogeneous, converged media data environment, the user profile construction process suffers from high data redundancy, complex feature extraction, and significant model training overhead, leading to high system computational resource consumption, heavy storage pressure, and significant response latency. To meet the following requirements for this application scenario: adaptability to dynamic updates and low redundancy, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned executing entity may generate a user profile based on the target user's basic data and behavioral data through the following steps: The first step is to standardize the acquired basic and behavioral data of the target users to obtain target user data. In practice, the aforementioned implementing entity can use Z-score standardization (or Min-Max standardization) to standardize the numerical data included in the basic and behavioral data of the target users, and use one-hot encoding to encode the categorical fields (such as product category, activity type, etc.) included in the basic and behavioral data of the target users to obtain target user data.

[0044] The second step is to determine the subject number included in the target user data as the primary key. In practice, the execution entity can extract fields that can uniquely identify each user (subject) from the target user data (such as user ID, national ID number) and define them as the primary key.

[0045] The third step is to build an index for the primary key in the storage device, thus obtaining the primary key index. In practice, the execution entity can create an index for the primary key in a database (such as MySQL or PostgreSQL) or a big data storage system (such as HBase) in the storage device, thus obtaining the primary key index.

[0046] The fourth step involves performing a join operation on the target user data based on the primary key and primary key index mentioned above to obtain preliminary joined data. In practice, the execution entity can perform a join operation (SQL JOIN operation) on the target user data based on the primary key and primary key index mentioned above to connect data scattered in different data tables (such as user basic information table and behavior record table) to obtain preliminary joined data.

[0047] The fifth step involves performing statistical aggregation on the numerical data included in the initial linked data, based on the primary key, to obtain statistical feature data. In practice, the executing entity can use the SQL GROUP BY clause and aggregate functions (SUM(), AVG(), MAX(), and COUNT()) based on the primary key to process the numerical data included in the initial linked data and obtain statistical feature data.

[0048] Step 6: Based on the aforementioned primary key, vectorize the textual and categorical data included in the preliminary associated data to obtain vectorized feature data. In practice, the executing entity can use word embedding techniques to vectorize the textual data (such as comment content) included in the preliminary associated data, and use one-hot encoding to vectorize the categorical data (such as product categories, activity types, etc.) included in the preliminary associated data, to obtain vectorized feature data. The word embedding techniques mentioned above can be Bag of Words, TF-IDF, BERT, etc.

[0049] Step 7: Integrate the aforementioned statistical feature data and vectorized feature data into a wide table, serving as the main behavioral data wide table. In practice, the aforementioned execution entity can use the aforementioned primary key and an SQL JOIN operation to integrate the aforementioned statistical feature data and vectorized feature data into a wide table, serving as a user profile.

[0050] The aforementioned steps one through seven and related content constitute an inventive point of this disclosure, and in conjunction with step "106," solve the technical problem of "high system computational resource consumption, high storage pressure, and significant response latency during user profile construction due to high data redundancy, complex feature extraction, and high model training overhead in a multi-source heterogeneous converged media data environment." Factors leading to high computational resource consumption, high storage pressure, and significant response latency often include: a large number of users' attention-grabbing content changing rapidly within a short period (e.g., e-commerce platforms' "618" and "Double 11" promotions). In a multi-source heterogeneous converged media data environment, the high data redundancy, complex feature extraction, and high model training overhead during user profile construction result in high system computational resource consumption, high storage pressure, and significant response latency. Solving these factors can reduce computational resource consumption, storage space occupation, and system response latency during user profile construction. To achieve this effect, firstly, the acquired basic data and behavioral data of the target users are standardized to obtain target user data. This eliminates differences in the units, formats, and distribution of the original data, resolving data inconsistency issues and providing high-quality, well-organized input for subsequent feature fusion and model calculations. Second, the subject IDs included in the target user data are determined as the primary key. This creates a globally unique identifier for each user, resolving user identity fragmentation and providing a solid foundation for cross-data source and cross-behavioral cycle user information association. Third, an index is built for the primary key in the storage device, resulting in a primary key index. This significantly improves the speed of data querying and association operations based on user identifiers. Fourth, based on the primary key and the primary key index, association operations are performed on the target user data to obtain preliminary associated data. This allows user information scattered across different data sources such as basic information tables, behavior log tables, and transaction record tables to be efficiently joined (JOIN) through primary and foreign key relationships, forming a complete data view centered around a single user and integrating multi-dimensional information. Fifth, based on the primary key, statistical aggregation processing is performed on the numerical data included in the preliminary associated data to obtain statistical feature data. This allows quantifiable key indicators to be extracted from the user's historical behavior. Sixth, based on the aforementioned primary key, the textual and categorical data included in the preliminary association data are vectorized to obtain vectorized feature data. This transforms discrete symbols that are difficult to compute directly (such as comments, search keywords, and interest tags) into dense numerical vectors that can be processed by machine learning models. Seventh, the aforementioned statistical feature data and vectorized feature data are integrated into a wide table as a user profile. This allows features from different sources and of different types to be horizontally concatenated according to the user's primary key, forming a flat, feature-complete dataset.Finally, in conjunction with "Step 106", based on the time series prediction model and the aforementioned historical access logs, load analysis is performed on the aforementioned thematic database and the aforementioned user profile to obtain the load analysis results. Based on network status information and the aforementioned load analysis results, traffic shaping and priority queuing strategies are adopted to push the aforementioned target content set to the display device of the aforementioned target user. This can reduce the consumption of computing resources, storage space occupation and system response latency during the user profile construction process.

[0051] Step 105: Based on user profiles, perform collaborative filtering on the thematic database to generate a target content set.

[0052] In some embodiments, the aforementioned execution entity may perform collaborative filtering on the aforementioned topic database based on the aforementioned user profile in order to generate a target content set.

[0053] In addressing the aforementioned technical challenges by employing technical solutions, and considering the application scenario—high-concurrency access to multi-source heterogeneous converged media data (e.g., during live broadcasts and content distribution of international sports events)—the following technical issues often arise: converged media data comes from diverse sources, has varying structures, and differs in dimensions and formats, making it difficult to directly use for model calculations. Furthermore, full-data feature engineering and model training consume significant resources and experience slow response times. Given the specific requirements of this application scenario—adaptability to multimodal data processing, low latency, and high throughput—we have decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned executing entity may perform collaborative filtering on the aforementioned topic database based on the aforementioned user profile to generate a target content set through the following steps: The first step is to obtain the core metadata dataset of the aforementioned thematic database. In practice, the executing entity can query the database system tables to obtain the core metadata dataset of the thematic database. The core metadata in this dataset includes publication time, source credibility, domain, and geographic region. As an example, the database system tables could be MySQL's INFORMATION_SCHEMA.COLUMNS or TABLES.

[0054] The second step involves preprocessing the core metadata dataset to obtain a candidate content set. In practice, firstly, the execution entity can use regular expressions to remove noise such as HTML tags and special characters from the core metadata dataset. Then, the SQL-based COALESCE function is used to fill the null values ​​in the core metadata dataset with default values ​​or statistical values ​​(mean, median). Finally, the numerical features in the core metadata dataset are standardized using Min-Max (or Z-score) to obtain the candidate content set. As an example, the regular expression can include, but is not limited to, "<[^>". > is used to remove HTML tags. "[^ws]" is used to remove all non-alphanumeric, non-underscore, and non-whitespace characters (i.e., special characters).

[0055] The third step involves encoding the metadata corresponding to the aforementioned candidate content set to obtain a structured feature set, and converting the text content included in the aforementioned candidate content set into numerical feature vectors to obtain a text feature vector set. In practice, the execution entity can use one-hot encoding to encode the metadata corresponding to the aforementioned candidate content set to obtain a structured feature set, and use TF-IDF (or Word2Vec, BERT, etc.) methods to convert the text content included in the aforementioned candidate content set into numerical feature vectors to obtain a text feature vector set.

[0056] The fourth step involves combining the structured feature set and the text feature vector set to obtain the candidate content feature set. In practice, the execution entity can use vector concatenation to combine the structured feature set and the text feature vector set to obtain the candidate content feature set. As an example, the vector concatenation method can be NumPy's np.hstack().

[0057] The fifth step involves generating a content recommendation list based on the content collaborative filtering method, the aforementioned user profile, and the aforementioned candidate content feature set. In practice, the executing entity can determine the similarity (such as cosine similarity or Jaccard similarity) between the aforementioned user profile and the aforementioned candidate content feature set, and filter out the candidate content feature set whose similarity is greater than a preset similarity threshold, as the content recommendation list. For example, the preset similarity threshold could be 0.75 or 0.85, etc.

[0058] Step 6: Based on the user collaborative filtering method, the aforementioned user profiles, and the aforementioned candidate content feature set, a user recommendation list is generated. In practice, the executing entity can determine the similarity (e.g., cosine similarity) between the aforementioned user profiles and each historical user profile in the historical user profile set, and use the candidate content feature set corresponding to the historical user profiles with similarity greater than the aforementioned preset similarity threshold as the user recommendation list. The aforementioned historical user profile set is the user profile set stored in the aforementioned historical access logs.

[0059] Step 7: Weight and merge the above content recommendation list and the above user recommendation list to obtain an initial fusion score. In practice, the above-mentioned implementing entity can use the formula: Initial Fusion Score = Content Recommendation List W1+ User Recommendation List W2 determines the initial fusion score. W1 and W2 can be preset values. For example, W1 = 0.5, W2 = 0.5.

[0060] Step 8: Based on preset business rules, adjust the initial fusion score to obtain the final recommendation score. In practice, the executing entity can adjust the business rules using the following formula: Final Recommendation Score = Initial Fusion Score × (1 + Timeliness Bonus Coefficient). Wherein, the timeliness bonus coefficient = 1 / (1 + Decay Rate × (Current Time - Content Publication Time)). This timeliness bonus coefficient represents the timeliness of each candidate content in the candidate content set. The newer the candidate content, the closer the coefficient is to 1, and the greater the bonus; as time goes on, the coefficient approaches 0, and the bonus weakens.

[0061] Step nine: Based on the final recommendation scores, sort the candidate content set in descending order to obtain the target content set. In practice, the executing entity can sort the candidate content set in descending order based on the final recommendation scores and select the top N content items as the target content set. Here, N can be a pre-defined positive integer. For example, N=10.

[0062] The aforementioned steps one through nine and related content serve as an inventive point of this disclosure, and, in conjunction with step "106," solve the technical problem of "diverse sources and inconsistent structures of converged media data, with differences in dimensions and formats, making it difficult to directly use for model calculations, and resulting in high resource consumption and slow response for full-data feature engineering and model training." Factors leading to high computational resource consumption and system response latency often include: in high-concurrency access scenarios of multi-source heterogeneous converged media data (e.g., during international sports event live broadcasts and content distribution), the diverse sources and inconsistent structures of converged media data, with differences in dimensions and formats, make it difficult to directly use for model calculations, and result in high resource consumption and slow response for full-data feature engineering and model training. Solving these factors can reduce computational resource consumption and system response latency. To achieve this effect, firstly, the core metadata dataset of the aforementioned thematic database is obtained. This provides high-quality, standardized data input for subsequent processing, avoiding scanning the entire dataset, reducing the amount of data to be processed from the source, and laying the foundation for accelerating information generation. Secondly, the core metadata dataset is preprocessed to obtain a candidate content set. Therefore, invalid, redundant, and low-quality data can be eliminated, reducing the unnecessary consumption of computing resources caused by data noise. Third, the metadata corresponding to the above candidate content set is encoded to obtain a structured feature set, and the text content included in the above candidate content set is converted into numerical feature vectors to obtain a text feature vector set. This transforms discrete, high-dimensional raw data (such as text and categories) into low-dimensional, dense numerical vectors, greatly compressing the scale of data representation. Fourth, the above structured feature set and the above text feature vector set are combined to obtain a candidate content feature set. This enables deep fusion of multimodal features, avoiding the need for subsequent algorithms to perform multiple alignment operations on features from different sources, simplifying the model structure, reducing redundant computation, and improving processing efficiency. Fifth, based on the content collaborative filtering method, the above user profile, and the above candidate content feature set, a content recommendation list is generated. This allows for the rapid identification of items similar to the user's historical preferences, completing the first round of rapid filtering of a large-scale candidate set, greatly reducing the computational pressure of subsequent fine-grained ranking steps. Sixth, based on the user collaborative filtering method, the above user profile, and the above candidate content feature set, a user recommendation list is generated. This allows us to discover users' potential but not yet explicitly expressed interests, effectively complementing content collaborative filtering. Seventh, the aforementioned content recommendation list and user recommendation list are weighted and merged to obtain an initial fusion score. This allows for flexible adjustment of the contribution of the two strategies—content similarity-based and user group similarity-based—using configurable weights (such as linear weighting). Eighth, based on preset business rules, the initial fusion score is adjusted to obtain the final recommendation score.Therefore, the key business logic of timeliness can be introduced with extremely low computational overhead (such as simple addition or subtraction of scores or threshold judgment). Ninth, based on the final recommendation score, the candidate content set is sorted in descending order to obtain the target content set. This ensures that the content that best meets user needs and business objectives is presented to the user first. Finally, combining with step "106", based on the time series prediction model and the historical access logs, load analysis is performed on the topic database and user profile to obtain the load analysis results. Based on network status information and the load analysis results, traffic shaping and priority queuing strategies are used to push the target content set to the display devices of the target users. This achieves rapid information generation and push under extremely high data throughput, while reducing computational resource consumption and system response latency.

[0063] In addressing the technical challenges mentioned above, and considering the application scenario—high-concurrency access to multi-source heterogeneous converged media data (e.g., during live broadcasts and content distribution of international sporting events)—the following technical issues often arise: Under high concurrency, the massive original volume, high feature dimensionality, and highly unpredictable access patterns of multimedia data lead to high data storage and transmission costs, slow query response times, and system throughput bottlenecks under static resource allocation. Given the specific requirements of this application scenario—flexible storage and transmission, and adaptability to accessing large amounts of data in a short period—we have decided to adopt the following solution: Optionally, after step 105 above, the method further includes: The first step is to analyze the data characteristics of the aforementioned target converged media data, the aforementioned thematic database, and the aforementioned user profiles to generate a compression strategy configuration table. In practice, the implementing entity can match the most suitable compression algorithm and parameters based on the data characteristics and predefined rules of the aforementioned target converged media data, the aforementioned thematic database, and the aforementioned user profiles. The aforementioned data characteristics can include: data type (e.g., text, numerical, image), data distribution (e.g., repetition, sparsity), and access patterns (e.g., read / write ratio, hot / cold data). As an example, the aforementioned predefined rules can include, but are not limited to: selecting dictionary encoding algorithms (e.g., Zstandard) for text data with high repetition, selecting columnar compression (e.g., RLE) for numerical data, using high compression ratio algorithms (e.g., Zstandard or Gzip) for data with low access frequency, and using fast compression algorithms (e.g., Snappy or LZ4) for data with high access frequency.

[0064] The second step involves applying multi-strategy compression to the target converged media data, the thematic database, and the user profiles, based on the aforementioned compression strategy configuration table, to obtain compressed data and a compression information table. In practice, the executing entity can utilize the native compression functionality provided by the database (such as MySQL's ROW_FORMAT=COMPRESSED, Oracle table compression, or AnalyticDB's Beam engine) and the aforementioned compression strategy configuration table to perform multi-strategy compression on the target converged media data, the thematic database, and the user profiles, obtaining compressed data and a compression information table. The compression information table can be used to record metadata such as the compression algorithm, original size, and compressed size for each data block.

[0065] The third step is to store the compressed data in a storage device and the compression information table in a metadata database. In practice, the executing entity can store the compressed data blocks in a distributed file system (such as HDFS) or object storage (such as S3) in the storage device, and store the compression information table in a metadata database (such as MySQL or Elasticsearch).

[0066] The fourth step involves generating a target data access path based on the aforementioned compression information table and the preset interceptor in response to a detected data access request. In practice, the executing entity can, upon detecting a data access request, query the corresponding data information (location, compression format, and verification information) in the compression information table based on the preset interceptor, and generate a target data access path that directly locates the compressed data block. The preset interceptor can be a Servlet Filter in Java EE or a HandlerInterceptor in the Spring framework.

[0067] The fifth step involves implementing real-time decompression based on streaming decompression technology and the aforementioned target data access path to obtain the target data. The target data is then cached according to the compression strategy configuration table. The streaming decompression technology can be lz4-java, zlib library streaming decompression, or Alibaba ZIP streaming decompression. The caching process can utilize Redis, Memcached, or Guava Cache.

[0068] Step 6: Based on the load balancer, determine the real-time load metrics of the server nodes to obtain a real-time load view. These real-time load metrics may include: CPU utilization, memory usage, network I / O, disk I / O, and the number of current concurrent tasks. For example, the load balancer could be Nginx, HAProxy, etc.

[0069] Step 7: Dynamically allocate tasks based on the scheduling algorithm and the aforementioned real-time load view. In practice, the execution entity can use a distributed resource scheduler to dynamically allocate tasks based on the scheduling algorithm and the aforementioned real-time load view. The scheduling algorithm can be round-robin, weighted round-robin, least-connections, or response-time-based scheduling algorithms, etc. As an example, the distributed resource scheduler can be YARN, Kubernetes Scheduler, or Apache Mesos, etc.

[0070] Step 8: Implement elastic provisioning of computing resources according to the auto-scaling rules. As an example, the auto-scaling rules mentioned above may include, but are not limited to: rules based on monitoring metrics (e.g., adding one instance in response to the average CPU utilization of all instances exceeding 70% for 3 consecutive minutes, and reducing one instance in response to the average CPU utilization being below 20% for 10 consecutive minutes), and rules based on scheduled tasks (e.g., adjusting the number of instances to 20 at 9:00 AM on each workday; and reducing the number of instances to 5 at 6:00 PM on each workday). These auto-scaling rules can be AWS Auto Scaling Groups, Google Cloud Managed Instance Groups, or Alibaba Cloud Elastic Scaling Service.

[0071] The first to eighth steps and related content described above, as an inventive point of this disclosure, solve the technical problem of "high data storage and transmission costs, slow query response, and system throughput bottlenecks under static resource allocation mode due to the large original volume, high feature dimensionality, and sudden access patterns of multimedia data under high concurrency access." Factors leading to high data storage and network transmission bandwidth consumption and long system response times often include: in high-concurrency access scenarios of multi-source heterogeneous converged media data (e.g., during international sports event live broadcasts and content distribution), the large original volume, high feature dimensionality, and sudden access patterns of multimedia data result in high data storage and transmission costs, slow query response, and system throughput bottlenecks under static resource allocation mode. Solving these factors can reduce data storage and network transmission bandwidth and reduce system response latency. To achieve this effect, firstly, the data characteristics of the aforementioned target converged media data, the aforementioned thematic database, and the aforementioned user profile are analyzed to generate a compression strategy configuration table. This allows for accurate identification of data redundancy patterns and access hotspots, providing a basis for subsequent differentiated compression, avoiding invalid calculations, and reducing data processing volume from the source. Second, based on the aforementioned compression strategy configuration table, the target converged media data, the aforementioned thematic database, and the aforementioned user profiles are compressed using multiple strategies to obtain compressed data and a compressed information table. This significantly reduces data volume, saves storage space and transmission bandwidth, while retaining key semantic information, laying the foundation for rapid retrieval. Third, the compressed data is stored in a storage device, and the compressed information table is stored in a metadata database. This enables structured organization of compressed data and efficient management of metadata, supporting rapid location and on-demand decompression, reducing unnecessary global scanning. Fourth, in response to detected data access requests, a target data access path is generated based on the aforementioned compressed information table and preset interceptors. This accurately filters irrelevant data requests, directly locking the target data block, reducing I / O overhead and query latency. Fifth, based on streaming decompression technology and the aforementioned target data access path, real-time decompression processing is implemented to obtain the target data, and the target data is cached according to the aforementioned compression strategy configuration table. This enables decompression and processing simultaneously, avoiding the resource consumption of full data decompression, and accelerating subsequent access by caching high-frequency data. Sixth, based on the load balancer, determine the real-time load metrics of server nodes to obtain a real-time load view. This allows for dynamic perception of system load distribution, providing a basis for intelligent resource scheduling and preventing single-point overload. Seventh, based on the scheduling algorithm and the aforementioned real-time load view, perform dynamic task allocation. This allows for the priority allocation of computing tasks to idle nodes, balancing cluster load, improving overall processing throughput, and reducing task queuing time. Eighth, based on automatic scaling rules, achieve elastic provisioning of computing resources.This allows for on-demand scaling of computing instances, enabling rapid expansion to ensure performance during peak periods and cost savings during off-peak periods, thus optimizing resource utilization. Ultimately, this reduces data storage and network transmission bandwidth consumption, and decreases system response latency.

[0072] Step 106: Based on the time series prediction model and historical access logs, perform load analysis on the topic database and user profiles to obtain load analysis results. Based on network status information and load analysis results, use traffic shaping and priority queuing strategies to push the target content set to the target user's display device.

[0073] In some embodiments, the execution entity may perform load analysis on the topic database and the user profile based on the time series prediction model and the historical access logs to obtain load analysis results, and based on network status information and the load analysis results, use traffic shaping and priority queuing strategies to push the target content set to the display device of the target user.

[0074] In some optional implementations of certain embodiments, the aforementioned execution entity may perform load analysis on the aforementioned topic database and the aforementioned user profile based on a time series prediction model and the aforementioned historical access logs to obtain load analysis results, and based on network status information and the aforementioned load analysis results, push the aforementioned target content set to the display device of the aforementioned target user using traffic shaping and priority queuing strategies: The first step is to obtain the historical access frequency time-series data corresponding to the aforementioned thematic database and the concurrent access volume time-series data corresponding to the aforementioned user profiles, based on the aforementioned historical access logs. In practice, the aforementioned execution entity can use log analysis tools (such as Logstash in the ELK stack) to obtain the historical access frequency time-series data corresponding to the aforementioned thematic database, and use application performance management (APM) tools (such as SkyWalking and Pinpoint) to obtain the concurrent access volume time-series data corresponding to the aforementioned user profiles. The aforementioned historical access frequency time-series data can be extracted access records sorted by time. The aforementioned concurrent access volume time-series data can be records of concurrent connections reflecting user activity patterns.

[0075] The second step involves performing predictive analysis on the historical access frequency time-series data and concurrent access volume time-series data based on the aforementioned time-series prediction model to obtain load analysis results. The aforementioned time-series prediction model can be a Long Short-Term Memory (LSTM) network model comprising an input layer (inputting time-series data with a time window length of 10 time units), an LSTM layer (with 50 hidden neurons), and an output layer (a fully connected layer that outputs the predicted load value for the next time point). This time-series prediction model can be pre-trained using the aforementioned historical access frequency time-series data, employing the Adam optimizer with mean squared error (MSE) as the loss function, a batch size of 32, and an initial learning rate of 0.001. The aforementioned load analysis results can include: access volume during specific time periods, peak concurrent user counts, etc.

[0076] The third step is to generate protocol priorities based on the load analysis results and the target content set described above. In practice, the execution entity can generate protocol priorities based on a preset business rule engine, taking into account the load analysis results and the target content set. This preset business rule engine could be Drools. As an example, the protocol priorities could include, but are not limited to: real-time news = high priority, historical articles = normal priority.

[0077] The fourth step is to monitor the network status information of the target user's display device in real time. In practice, the execution entity can monitor the network status information of the target user's display device in real time through a network status API (such as Navigator.connection). This network status information can include network bandwidth utilization, latency, packet loss rate, etc.

[0078] The fifth step involves dynamically adjusting the traffic shaping strategy based on the aforementioned network status information and protocol priorities, and assigning priority queue indices to each target content in the target content set, resulting in a priority queue index set. In practice, the execution entity can dynamically adjust the parameters of the traffic shaper based on the aforementioned network status information. For example, in response to tight network bandwidth, the token bucket generation rate can be reduced by the traffic shaper to limit the overall output bandwidth; in response to good network conditions, the output bandwidth limitation can be relaxed. Then, based on the aforementioned protocol priorities, priority queue indices (e.g., 0 - highest, 9 - lowest) are assigned to each target content in the target content set, resulting in a priority queue index set. This priority queue index set can be used to determine the position of each target content in the target content set within the sending queue. The priority queue index set can be an ordered list to ensure that the sender agent prioritizes processing the content with the smallest index value (i.e., the highest priority). The traffic shaper adjustment can be implemented using the Linux TC (Traffic Control) tool.

[0079] The sixth step involves pushing the target content set to the target user's display device according to the priority queue index set. In practice, the execution entity can first establish a stable connection with the target user's display device (such as a WebSocket long connection), and then encapsulate the target content in the target content set into network data packets and send them to the target user's display device according to the priority order of the priority queue index set.

[0080] The above embodiments of this disclosure have the following beneficial effects: By constructing a full-link optimization mechanism of "data perception-intelligent processing-dynamic scheduling" through the information push method based on converged media data in some embodiments of this disclosure, the storage and transmission occupancy is reduced, computing resource consumption is lowered, and system response speed is accelerated in high-concurrency, multi-source heterogeneous converged media scenarios. Specifically, the reason for high system storage and transmission occupancy, high latency, and even potential system lag and overload is that existing methods lack forward-looking prediction and fine-grained scheduling of the overall load when dealing with high-concurrency, multi-source heterogeneous converged media data scenarios. They fail to combine network status information (such as bandwidth and latency) with load analysis results to dynamically adjust traffic shaping strategies and priority queues, resulting in high system storage and transmission occupancy, high latency, and even potential system lag and overload. Based on this, the information push method based on converged media data in some embodiments of this disclosure firstly, in response to detecting that a target user's access operation to target converged media data, and that the number of access requests for the target converged media data exceeds a preset access threshold, adaptively extracts the target converged media data to obtain a basic database table set. This allows for the rapid filtering of high-value, highly relevant data subsets from massive multi-source data, avoiding scanning the entire dataset from the source and significantly reducing the data scale for subsequent processing. Secondly, the basic database table set is standardized to obtain a standard database table set. This eliminates inconsistencies in the original data's dimensions, formats, and encoding, providing unified and well-organized high-quality data input for subsequent fusion analysis and reducing the computational complexity of data cleaning and transformation processes. Then, based on historical access logs and incremental update methods, the standard database table set undergoes thematic data fusion processing to generate a thematic database. This dynamically integrates related data, forming a data aggregation view with distinct themes, avoiding repeated related queries on the same data, and improving data retrieval and utilization efficiency. Subsequently, user profiles are generated based on the target user's basic data and behavioral data. This allows for the construction of accurate user interest models, providing a basis for personalized recommendations and reducing the distribution of invalid content and the waste of computational resources. Next, based on the aforementioned user profiles, collaborative filtering is performed on the aforementioned thematic database to generate a target content set. This enables rapid and accurate screening and sorting of candidate content, significantly narrowing the data scope for refined sorting operations and reducing the real-time processing pressure on the system. Finally, based on the time-series prediction model and the aforementioned historical access logs, load analysis is performed on the aforementioned thematic database and the aforementioned user profiles. The load analysis results, along with network status information and the aforementioned load analysis results, are used to push the aforementioned target content set to the display devices of the aforementioned target users using traffic shaping and priority queuing strategies. This allows for proactive prediction and adaptive control of system load and network conditions, ensuring the supply of resources for critical tasks, optimizing overall throughput, and reducing response latency.This implementation achieves the technical effects of reducing storage and transmission overhead, lowering computing resource consumption, and accelerating system response speed in high-concurrency, multi-source heterogeneous converged media scenarios.

[0081] Continue to refer to Figure 2 As a response to the above Figure 1 The implementation of the method shown in this disclosure provides some embodiments of an information push device based on converged media data. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0082] like Figure 2 As shown, some embodiments of the information push device 200 based on converged media data include: a detection unit 201, a standardization processing unit 202, a topic data fusion processing unit 203, a generation unit 204, a collaborative filtering unit 205, and a load analysis unit 206. The detection unit 201 is configured to adaptively extract the target media data to obtain a basic database table set in response to the detection of a target user's access operation to the target media data, and the number of access requests to the target media data exceeding a preset access threshold. The standardization processing unit 202 is configured to standardize the basic database table set to obtain a standard database table set. Thematic data fusion processing unit 203 is configured to perform thematic data fusion processing on the standard database table set based on historical access logs and incremental update methods to generate a thematic database. The generation unit 204 is configured to generate a user profile based on the target user's basic data and behavioral data. The collaborative filtering unit 205 is configured to perform collaborative filtering on the thematic database based on the user profile to generate a target content set. The load analysis unit 206 is configured to perform load analysis on the thematic database and the user profile based on a time series prediction model and the historical access logs to obtain load analysis results. Based on network status information and the load analysis results, the target content set is pushed to the target user's display device using traffic shaping and priority queuing strategies.

[0083] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.

[0084] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0085] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0086] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0087] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.

[0088] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0089] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0090] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: in response to detecting that a target user's access operation to target converged media data, and that the number of access requests for the target converged media data exceeds a preset access threshold, adaptively extract the target converged media data to obtain a basic library table set; standardize the aforementioned basic library table set to obtain a standard library table set; perform thematic data fusion processing on the aforementioned standard library table set based on historical access logs and incremental update methods to generate a thematic database; generate a user profile based on the aforementioned target user's basic data and behavioral data; perform collaborative filtering on the aforementioned thematic database based on the aforementioned user profile to generate a target content set; perform load analysis on the aforementioned thematic database and the aforementioned user profile based on a time series prediction model and the aforementioned historical access logs to obtain load analysis results; and, based on network status information and the aforementioned load analysis results, push the aforementioned target content set to the aforementioned target user's display device using traffic shaping and priority queuing strategies.

[0091] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0093] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a detection unit, a standardization processing unit, a topic data fusion processing unit, a generation unit, a collaborative filtering unit, and a load analysis unit. The names of these units do not necessarily limit the specific unit itself. For example, the detection unit may also be described as "a unit that, in response to detecting a target user's access operation to target converged media data, and the number of access requests to the target converged media data exceeding a preset access threshold, adaptively extracts the target converged media data to obtain a basic database table set."

[0094] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0095] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. An information push method based on converged media data, comprising: In response to detecting that a target user accesses target converged media data and that the number of access requests for the target converged media data exceeds a preset access threshold, the target converged media data is adaptively extracted to obtain a basic library table set. The basic library table set is standardized to obtain the standard library table set; Based on historical access logs and incremental update methods, the standard library table set is subjected to thematic data fusion processing to generate a thematic database; Based on the target users' basic and behavioral data, a user profile is generated. Based on the user profile, the topic database is collaboratively filtered to generate a target content set; Based on the time series prediction model and the historical access logs, load analysis is performed on the topic database and the user profile to obtain load analysis results. Based on network status information and the load analysis results, traffic shaping and priority queuing strategies are used to push the target content set to the target user's display device.

2. The method according to claim 1, wherein, In response to detecting a target user's access operation to target converged media data, and the number of access requests for the target converged media data exceeding a preset access threshold, adaptive extraction is performed on the target converged media data to obtain a basic library table set, including: In response to detecting that a target user accesses target converged media data and that the number of access requests for the target converged media data exceeds a preset access threshold, the target converged media data is parsed to obtain a multimodal dataset. The multimodal dataset is cleaned to obtain a cleaned dataset; The cleaned dataset is integrated into the base library table set.

3. The method according to claim 1, wherein, The standardization process of the base library table set to obtain the standard library table set includes: Outlier smoothing is performed on the numeric fields contained in each basic database table in the basic database table set, and enumerated values ​​are standardized on the categorical fields contained in each basic database table in the basic database table set to obtain the initial database table set. Data anonymization processing is performed on the personal information fields contained in each initial table in the initial table set to obtain the standard table set.

4. The method according to claim 1, wherein, The behavioral data includes: historical clickstream, search keywords, dwell time, and interaction feedback data.

5. The method according to claim 1, wherein, The process of performing load analysis on the topic database and the user profile based on the time series prediction model and the historical access logs to obtain load analysis results, and then pushing the target content set to the target user's display device using traffic shaping and priority queuing strategies based on network status information and the load analysis results, includes: Based on the historical access logs, obtain the historical access frequency time-series data corresponding to the topic database and the concurrent access volume time-series data corresponding to the user profile; Based on the time series prediction model, predictive analysis is performed on the historical access frequency time series data and concurrent access volume time series data to obtain load analysis results; Based on the load analysis results and the target content set, a protocol priority is generated; Real-time monitoring of the network status information of the target user's display device; Based on the network status information and the protocol priority, the traffic shaping strategy is dynamically adjusted, and priority queue indexes are assigned to each target content in the target content set to obtain a priority queue index set; Based on the priority queue index set, the target content set is pushed to the target user's display device.

6. The method according to claim 1, wherein, The process of performing thematic data fusion processing on the standard library tables based on historical access logs and incremental update methods to generate a thematic database includes: Based on the association rule algorithm, the historical access logs and the standard library table set are analyzed to obtain a high-frequency access data table set; The high-frequency access data table set is sorted by access popularity to obtain a high-frequency access data table sequence; Based on the preset domain knowledge graph and the high-frequency access data table sequence, a topic definition file set is generated, wherein the topic definition files in the topic definition file set include: topic name, basic table set, association key and priority; Based on the aforementioned topic definition file set, perform relational queries and statistical aggregations on the standard library table set to generate a wide table dataset; The wide table dataset is processed into feature vectors to obtain the feature dataset; Based on the unsupervised machine learning algorithm and the feature dataset, the business label set of the feature dataset is determined to obtain the label-enriched dataset; The analysis results are obtained by analyzing the tag-enriched dataset and the historical access logs; Based on the analysis results, a storage format is determined, wherein the storage format includes row-based storage and column-based storage; Based on the storage format, a structured query statement is generated, and the structured query statement is executed to generate a physical table structure in the storage device; Based on the data lake toolchain, the tag-enriched dataset is incrementally updated into the physical table structure to generate a thematic database.

7. An information push device based on converged media data, comprising: The detection unit is configured to adaptively extract the target converged media data in response to the detection of a target user's access operation to the target converged media data and the number of access requests to the target converged media data exceeding a preset access threshold, thereby obtaining a basic library table set. The standardization processing unit is configured to perform standardization processing on the base library table set to obtain a standard library table set. Thematic data fusion processing unit is configured to perform thematic data fusion processing on the standard library table set based on historical access logs and incremental update method to generate a thematic database; The generation unit is configured to generate a user profile based on the target user's basic data and behavioral data. The collaborative filtering unit is configured to perform collaborative filtering on the topic database based on the user profile to generate a target content set; The load analysis unit is configured to perform load analysis on the topic database and the user profile based on the time series prediction model and the historical access logs, obtain load analysis results, and push the target content set to the target user's display device based on network status information and the load analysis results, using traffic shaping and priority queuing strategies.

8. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 6.

9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Integrated convergence media cloud production and release system and method

    CN106101176A

  • Personalized content recommendation and behavior analysis method and system for convergence media user

    CN119862327A

  • Multi-platform media content optimization recommendation method and system based on AI

    CN120045795A

  • Big data-oriented multi-model fusion analysis processing method, apparatus and device, and medium

    CN121658465A

  • Methods and apparatus for coordination of network traffic between wireless network devices and computing platforms

    US20200359265A1