A data collection and analysis system for new media operation management
By using a closed-loop collaborative design for end-to-end data collection and multimodal intelligent acquisition and analysis, the problems of incomplete data collection, poor real-time performance, insufficient multimodal mining, and weak security in new media operation systems have been solved. This has enabled efficient and secure data collection, processing, and decision support, thereby improving the overall efficiency and security of new media operations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA WEST NORMAL UNIVERSITY
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-21
Smart Images

Figure CN122437683A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of new media operation technology, specifically to a data collection and analysis system for new media operation management that integrates multi-source fusion acquisition, intelligent refined processing, deep predictive analysis, and end-to-end security protection. It is applicable to the full-process management of operational data across various new media platforms such as short videos, social media, news, and live streaming, and can be applied to scenarios such as new media operations for SMEs, multi-account management for MCN agencies, and new media matrix operations for brands. This system is based on the core concept of end-to-end data closed-loop collaboration, integrating core technologies such as multi-source heterogeneous data fusion acquisition, multimodal intelligent analysis, Transformer-LSTM hybrid prediction, and quantum encryption end-to-end security protection. It addresses industry pain points in existing new media operations, such as data dispersion, delayed analysis, insufficient multimodal data value mining, weak security, and blind decision-making. It achieves intelligent and refined management of operational data, improving operational efficiency and the scientific nature of decision-making, and belongs to the interdisciplinary technical field of data processing and new media operation. Background Technology
[0002] With the rapid development of digital media technology, new media has become a core carrier for corporate brand communication, user reach, and product conversion. The level of refinement in operation and management directly determines operational effectiveness and market competitiveness. A core requirement of new media operations is to collect and analyze data throughout the entire process to uncover user behavior patterns, content dissemination characteristics, and operational shortcomings, providing data support for decision-making. However, current new media operations involve multiple types of data, including content, users, operations, and competitors, scattered across different platforms such as Douyin, Xiaohongshu, WeChat Official Accounts, Weibo, and Bilibili. Each platform has different interface protocols, encryption rules, and anti-scraping mechanisms, making data integration difficult. Existing new media operation data collection and analysis solutions still have many shortcomings and cannot meet the needs of refined operations, as detailed below:
[0003] First, data collection suffers from limitations, lacking completeness and real-time accuracy. Existing systems often rely on single API connections or simple web scraping, failing to capture core, in-depth platform data such as video completion rates, deeper user interaction patterns, and real-time emotional fluctuations in live stream comments. Furthermore, existing collection technologies have weak anti-scraping capabilities, easily leading to data interruptions and gaps when faced with frequently updated platform verification mechanisms (such as slider verification, device fingerprint detection, and weekly signature algorithm changes). Additionally, traditional collection methods often employ scheduled tasks, resulting in delays exceeding 30 minutes, failing to capture real-time data changes during live streams or sudden trending events. Moreover, existing solutions neglect comprehensive multimodal data collection, often only collecting text and basic numerical values, failing to reflect the full operational picture and resulting in a lack of material for subsequent analysis.
[0004] Secondly, data processing capabilities are weak, and the value of multimodal data is insufficiently explored. Existing cleaning techniques only perform basic deduplication filtering, resulting in low accuracy in recognizing semantically repetitive but textually different content. Noise removal is incomplete, and inconsistent data definitions across different platforms (such as different definitions of "like") lead to incomparable data. More importantly, multimodal data such as images, videos, and audio are only subject to simple format conversion or storage, without semantic fusion and feature extraction. For example, it cannot identify the emotional tone in video footage or the frequency of brand logos in images, thus failing to fully explore the data value. The purity of the processed data only improves by about 30%, which is insufficient to support in-depth analysis.
[0005] Secondly, data analysis lacks depth and focus, resulting in insufficient decision support capabilities. Existing systems primarily rely on basic data statistics, such as readership and follower counts, with limited analytical dimensions and a lack of integrated analysis of users, content, operations, and competitors. Traditional analytical models perform poorly in semantic understanding and trend prediction, failing to grasp popular online terms or implied emotions, and lack personalized decision-making pushes. Operational decisions depend on experience, exhibiting strong bias and failing to answer questions like "Why is this content popular?" or "What should we publish next?", leading to wasted resources.
[0006] Fourth, the data security protection system is inadequate, posing multi-dimensional risks. The existing system relies solely on basic account passwords and TLS encryption, making data transmission vulnerable to quantum computing attacks. Core and general data are not stored in a hierarchical manner, and access control and anomaly detection are lacking. Once the database is compromised, sensitive user information is highly susceptible to leakage, with a security incident rate as high as 15% and data breaches accounting for over 60%, posing significant legal and reputational risks to the enterprise.
[0007] Finally, the system suffers from poor adaptability and scalability, and insufficient module collaboration. Existing systems mostly employ fixed architectures, failing to simultaneously meet the needs of operational scenarios of varying scales and lacking custom configuration support. Each module operates independently, with a lack of collaborative design in the data collection, processing, analysis, and decision-making processes, resulting in data silos. For example, analysis results cannot be fed back to the data collection module to optimize collection strategies, and decision-making suggestions cannot be directly linked to content production tools, leading to low overall operational efficiency.
[0008] In summary, the current system suffers from defects such as incomplete data collection, inaccurate processing, insufficient multimodal mining, superficial analysis, lack of security, and poor module collaboration. There is an urgent need for a new media operation and management data collection and analysis system with full-link closed-loop collaboration as its core, possessing completeness, real-time performance, multimodal parsing capabilities, high security, and high adaptability. This invention proposes a solution to the above problems, making up for all the shortcomings of the existing technology. Summary of the Invention
[0009] The purpose of this invention is to overcome the aforementioned deficiencies of existing technologies and provide a data collection and analysis system for new media operation and management. This invention is based on a full-link data closed-loop collaboration concept, achieving real-time collection of multi-platform, multi-type, and multi-modal data throughout the entire process; improving data quality and value mining capabilities through intelligent and refined processing and multi-modal semantic fusion; strengthening decision support by constructing a multi-dimensional in-depth analysis and dynamic, accurate prediction system; creating a full-link, multi-layered security protection system to prevent data risks; and simultaneously improving system adaptability and module synergy, achieving seamless collaboration across the entire link of collection, processing, storage, analysis, decision-making, and security protection. This provides accurate, efficient, and secure support for new media operation decisions, reduces operating costs, and improves operational effectiveness.
[0010] To achieve the aforementioned objectives, this invention adopts the following technical solution: A data collection and analysis system for new media operation management, comprising a data collection module, a data preprocessing module, a data storage module, a data analysis module, a decision-making and push module, a security protection module, and a system management module. Each module operates with a core of end-to-end data closed-loop collaboration, interconnecting and seamlessly transferring data to achieve end-to-end closed-loop management of new media operation data collection, processing, storage, analysis, decision-making, and security protection. The functional design of each module retains core innovations while strengthening inter-module collaboration, and simultaneously compensates for the shortcomings of insufficient depth in multimodal data processing, as detailed below:
[0011] 2.1 Data Acquisition Module
[0012] The data acquisition module is the core of the system's data input. It utilizes a fusion acquisition solution combining API, distributed crawler, and multimodal intelligent acquisition, along with anti-crawling adaptation and intelligent scheduling. This enables real-time and incremental acquisition of multi-platform, multi-type, and multimodal data, addressing issues such as incomplete acquisition, poor real-time performance, weak anti-crawling capabilities, and shallow multimodal acquisition. Specifically, it includes an API integration unit, a distributed crawler unit, a multimodal intelligent acquisition unit, an anti-crawling adaptation unit, and an acquisition scheduling unit. These units collaborate to complete acquisition tasks, and the acquired data is synchronized to the data preprocessing module in real time, achieving seamless integration between acquisition and processing.
[0013] API Integration Unit: Adopting a RESTful architecture, this unit supports OAuth 2.0 authentication and dynamic adaptive adjustment of rate limits. It pre-connects to open APIs of mainstream platforms such as Douyin, Xiaohongshu, and WeChat Official Accounts to obtain basic and some in-depth authorization data. It also reserves API extension interfaces to support rapid integration with new platforms without requiring modifications to the system core, thus improving adaptability. This unit has a built-in token refresh mechanism that automatically requests a refresh when the AccessToken is about to expire, ensuring continuous data collection.
[0014] Distributed crawler unit: Utilizing a distributed cluster architecture, multiple data collection nodes are deployed. A single node can collect up to 500,000 records per day. It supports dynamic webpage rendering (Headless Browser) and collects in-depth data from APIs not publicly available on the platform. An incremental collection and data pre-validation strategy is employed, collecting only newly added / updated data while performing basic format validation on the collected data to reduce invalid data flowing into subsequent stages and lower system load. Crawler nodes share a task queue via Redis to avoid duplicate collection.
[0015] Multimodal Intelligent Acquisition Unit: This is the core optimized unit of this module. It overcomes the shortcomings of traditional multimodal acquisition which only collects data, and realizes the integration of acquisition and preliminary feature extraction. It collects multimodal data such as text, images, videos, bullet comments, and emoticons. It extracts text from images using OCR technology and transcribes video / audio speech using ASR technology. At the same time, it uses a lightweight CNN model to extract visual features of images (such as color histograms and object categories), transforming unstructured data into semi-structured feature data, providing a foundation for subsequent analysis.
[0016] Anti-scraping adaptation unit: It has a built-in dynamic protocol adaptation mechanism to simulate the request characteristics (User-Agent, device fingerprint) of different user terminals, and automatically adjusts the collection frequency, request header parameters and parsing rules according to the response status code and page structure changes of the target platform; it maintains a proxy IP pool and automatically switches when an IP is detected to ensure that the collection task is not interrupted.
[0017] Data Acquisition and Scheduling Unit: Dynamically allocates acquisition resources based on task priority and system load; supports multiple modes including scheduled acquisition, event-triggered acquisition (such as keyword monitoring), and manual acquisition, ensuring that important data is acquired first.
[0018] 2.2 Data Preprocessing Module
[0019] The data preprocessing module is responsible for cleaning, transforming, and fusing the collected raw data to improve data quality. Specifically, it includes a data cleaning unit, an entity alignment unit, a multimodal fusion unit, and a data standardization unit.
[0020] Data cleaning unit: Removes duplicate data, filters advertising noise, and corrects outliers; uses semantic similarity algorithms to identify comments with duplicate content but different text, improving deduplication accuracy.
[0021] Entity alignment unit: Identifies the identity of the same entity (such as the same brand or the same influencer) on different platforms, establishes a globally unique ID mapping, and solves the problem of data silos.
[0022] Multimodal fusion unit: Associates text, image, and video features, such as associating video frame features with bullet screen text at the corresponding time to form multimodal data records.
[0023] Data standardization unit: unify the data standards of different platforms, such as unifying "likes" from different platforms into a standard interaction indicator, and unifying timestamps into UTC standard time.
[0024] 2.3 Data Storage Module
[0025] The data storage module employs a tiered storage strategy to ensure data access efficiency and security, specifically including a hot data storage unit, a cold data storage unit, and an encrypted storage unit.
[0026] Hot data storage unit: Uses Redis and Elasticsearch cluster to store real-time interactive data for the past 7 days, supporting millisecond-level queries.
[0027] Cold data storage unit: Utilizes object storage and relational databases to store historical archived data, reducing costs.
[0028] Encrypted storage unit: Core sensitive data is encrypted and stored using quantum encryption algorithms to ensure that the data cannot be cracked even if quantum computers become widespread in the future.
[0029] 2.4 Data Analysis Module
[0030] The data analysis module is the core intelligent engine of the system, employing deep learning models for in-depth data mining. Specifically, it includes a sentiment analysis unit, a propagation path tracing unit, a user profile building unit, and a trend prediction unit.
[0031] Sentiment Computing Unit: Based on the BERT-BiLSTM-CRF combined model, it performs fine-grained sentiment polarity judgment and emotion category recognition on text comments and video bullet comments, and supports aspect-level sentiment analysis.
[0032] Propagation Path Tracking Unit: Based on graph database technology, it reconstructs the node diffusion path of information in social networks and identifies key propagation nodes.
[0033] User profile building unit: Based on user behavior sequences, tags are assigned and interest vectors are calculated to achieve user segmentation.
[0034] Trend prediction unit: Employs a Transformer-LSTM hybrid model, combining historical data with external indices, to predict the trend of topic popularity and provide confidence intervals.
[0035] 2.5 Decision Push Module
[0036] The decision-making push module transforms analysis results into actionable suggestions, specifically including a strategy generation unit, an automated reporting unit, a risk warning unit, and a feedback loop unit.
[0037] Strategy Generation Unit: Based on the analysis results, it generates suggestions for content selection, optimization of publication time, and placement strategies.
[0038] Automated Reporting Unit: Supports custom templates and automatically generates daily, weekly, and monthly reports.
[0039] Risk warning unit: Set thresholds and issue alerts through multiple channels once negative public opinion or data anomalies are triggered.
[0040] Feedback closed-loop unit: Records the effect of decision execution, back-optimizes the analysis model, and forms a closed loop.
[0041] 2.6 Security Protection Module
[0042] The security protection module runs through the entire system chain and specifically includes a transmission encryption unit, an access control unit, and an audit log unit.
[0043] Transmission encryption unit: Employs national cryptographic algorithms and quantum key distribution technology to ensure data transmission security.
[0044] Access control unit: Based on the RBAC model, providing fine-grained access control.
[0045] Audit log unit: Records all operations to ensure traceability.
[0046] 2.7 System Management Module
[0047] The system management module is responsible for system configuration and monitoring, specifically including a user management unit, a task configuration unit, and a system monitoring unit, to ensure stable system operation.
[0048] The beneficial effects of this invention are as follows:
[0049] 1. End-to-End Closed-Loop Collaboration for Enhanced Operational Efficiency: This invention breaks down barriers between data collection, processing, analysis, and decision-making processes through an end-to-end closed-loop collaborative data design, achieving seamless data flow and feedback. Decision results can directly inform collection strategies, and analysis models can self-optimize based on execution performance, significantly improving the overall efficiency and response speed of new media operations and resolving the issue of poor module collaboration in existing systems.
[0050] 2. Deep Multimodal Fusion for Data Value Extraction: This invention breaks through the limitations of traditional methods that only process text data. Through a multimodal intelligent acquisition unit and a multimodal fusion unit, it achieves semantic-layer fusion analysis of text, images, videos, and audio. It can identify deeper information such as emotion in video footage and brand elements in images, significantly improving the depth of data value extraction and solving the problem of insufficient multimodal data value extraction.
[0051] 3. Deep Intelligent Analysis, Enhanced Decision Support: This invention introduces a Transformer-LSTM hybrid prediction model and aspect-level sentiment analysis technology, achieving a leap from descriptive analysis to predictive analysis. It can accurately predict trending topics and identify user sentiment tendencies with fine granularity, providing operators with forward-looking and actionable decision-making suggestions and reducing the uncertainty of decision-making.
[0052] 4. Quantum encryption protection to ensure data security: This invention is the first to introduce a quantum encryption full-link security protection system into the new media operation system. It adopts encryption algorithms that are resistant to quantum computing attacks for core sensitive data, and combines hierarchical storage and anomaly detection to significantly reduce the risk of data leakage, meet increasingly stringent data compliance requirements, and solve the problem of weak security.
[0053] 5. High adaptability and scalability, reducing operation and maintenance costs: This invention adopts a modular design and distributed architecture, supporting rapid access to multiple platforms and custom configurations, and can adapt to the operational needs of enterprises of different sizes. Dynamic anti-crawling adaptation and intelligent scheduling mechanisms reduce data collection and maintenance costs, and improve system stability and availability. Attached Figure Description
[0054] Figure 1 This is a schematic diagram of the overall architecture of the system of the present invention, showing the connection relationship and data flow between the modules;
[0055] Figure 2 The flowchart shows the process of the multimodal intelligent acquisition unit in the data acquisition module, illustrating the specific steps of OCR, ASR, and feature extraction.
[0056] Figure 3 This is a schematic diagram of the Transformer-LSTM hybrid prediction model in the data analysis module, showing the logic of the input layer, encoding layer and prediction layer;
[0057] Figure 4 This is a schematic diagram of the quantum encryption end-to-end protection mechanism in the security protection module, illustrating the encryption process of key generation, transmission, and storage.
[0058] Figure 5 The flowchart shows the feedback closed-loop unit in the decision push module, illustrating the closed-loop logic from decision execution to model optimization. Detailed Implementation
[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0060] Example 1: System Overall Architecture and Hardware Deployment
[0061] like Figure 1 As shown, the data collection and analysis system for new media operation and management provided by this invention adopts a cloud-based distributed architecture in its physical deployment to support high-concurrency data processing and elastic expansion. The overall system architecture is divided into an infrastructure layer, a data layer, a service layer, and an application layer, with each layer interconnected through a high-speed internal network.
[0062] The infrastructure layer forms the physical foundation of the system, including compute node clusters, storage resource pools, and network load balancing equipment. The compute node clusters are managed using Kubernetes containerized orchestration, comprising a master node and worker nodes. The master node is responsible for task scheduling, resource monitoring, and fault recovery; worker nodes run specific data acquisition, processing, and analysis services. For CPU-intensive tasks (such as video transcoding, image recognition, and model inference), high-performance computing instances equipped with GPU accelerator cards are configured; for I / O-intensive tasks (such as database read / write and file storage), high-disk-throughput instances and NVMe SSD storage are configured. The storage resource pool employs a tiered storage strategy, storing hot data in an in-memory database and cold data in object storage. Network load balancing deploys an LVS+Nginx load balancer at the system entry point to evenly distribute external API requests and internal microservice calls to backend service instances, avoiding single-point overload. Simultaneously, a WAF (Web Application Firewall) is configured to defend against SQL injection, XSS, and DDoS attacks, ensuring the security of the system entry point. In addition, the introduction of service mesh (such as Istio) technology can manage communication traffic between services, enabling circuit breaking, rate limiting, and tracing, thereby enhancing the stability of the microservice architecture.
[0063] The data layer comprises a data acquisition module, a data preprocessing module, and a data storage module. The data acquisition module is deployed in a separate, secure DMZ, interacting with external new media platforms through a proxy IP pool to prevent external attacks from directly accessing the internal network. The data preprocessing and data storage modules are deployed in the core area of the internal network to ensure the security and privacy of the data processing process. The message queue uses a Kafka cluster with multiple partitions to ensure high throughput, and consumer groups dynamically scale according to business logic.
[0064] The service layer comprises a data analysis module, a decision-making and push module, and a system management module. This layer adopts a microservice architecture, with each functional unit deployed independently and communicating via gRPC or RESTful API. This design makes the system highly adaptable, allowing for flexible expansion of the computing resources of a particular module based on business needs. For example, during peak sales periods, the computing nodes of the trend prediction unit can be expanded independently. Service registration and discovery utilize Consul or Nacos to ensure dynamic awareness of service instances.
[0065] The application layer includes a security protection module and a user interface. The security protection module is deployed alongside each service in a sidecar pattern, monitoring traffic and user behavior in real time. The user interface provides web dashboards, mobile apps, and API interfaces to meet the needs of different users. The front-end uses Vue.js or React frameworks, supporting component-based development for rapid iteration.
[0066] Each module operates around a closed-loop data collaboration model, working in tandem with the others. For example, when the decision-making push module detects that a certain type of content is performing poorly, it automatically triggers the data collection module to increase the frequency of collecting competitor data for that type of content in order to analyze the reasons and form a closed loop. The system monitoring unit collects various metrics through Prometheus and displays them through Grafana. Once CPU or memory usage exceeds a threshold, it automatically triggers an elastic scaling strategy.
[0067] Example 2: Detailed Implementation of the Data Acquisition Module
[0068] like Figure 2 As shown in the illustration, this embodiment details the workflow of the multimodal intelligent acquisition unit in the data acquisition module. This unit is crucial for the system to acquire high-quality data, overcoming the limitations of traditional acquisition methods that only acquire text.
[0069] Step S1: Task Initialization and Multimodal Recognition. The system receives the data collection task and parses the target URL or content ID. It identifies the content type: if it's text and image content, it marks that text and images need to be collected; if it's video content, it marks that video stream, audio stream, cover image, and bullet comments need to be collected. The task metadata is written to the MySQL task table, and the status is marked as "Pending Execution".
[0070] Step S2: Basic Data Crawling. Obtain the content's metadata (title, author, publication time, number of likes, number of comments, etc.) and the original file link through the API integration unit or distributed crawler unit. For pages requiring login, retrieve valid session credentials from the Cookie pool. The Cookie pool uses the LRU (Least Recently Used) algorithm for management, periodically checking its liveness status, automatically removing expired cookies and triggering a re-login process.
[0071] Step S3: Multimodal Feature Extraction. This is the core innovative step of this unit. For image data: Download the original image and first perform quality inspection to remove blurry or damaged images. Then, call an OCR engine (such as PaddleOCR) to extract text information from the image, such as promotional slogans and product prices in posters, and convert the image text into text fields for storage. At the same time, use a lightweight CNN model (such as MobileNet) to extract visual feature vectors, including color histograms (to determine the style of the image), object detection boxes (to determine whether it contains people, products, logos), and scene classification (to determine whether it is indoor, outdoor, studio, etc.). These features will serve as the basis for subsequent analysis of the image's attractiveness. For video data: Download video files or streaming media clips. First, call an ASR engine (such as Whisper) to transcribe the speech in the video into text and add timestamps to form subtitle text. This makes the video content searchable by search engines and also allows for sentiment analysis. Second, extract keyframes from the video (1 frame per second or scene transition frames), and perform the same OCR and CNN feature extraction on the keyframes as on the images. Furthermore, the spectral characteristics of the audio were analyzed to identify the rhythm (fast / slow), mood (excited / soothing), and volume changes of the background music; these auditory features are highly correlated with the completion rate. For text data: in addition to collecting the main text, comment text and bullet screen text were also collected. Bullet screen text was time-aligned to ensure that the content corresponded to the video footage at all times.
[0072] Step S4: Data Pre-validation and Encapsulation. The extracted multimodal features are associated with the original data and encapsulated into a unified multimodal data object (MultimodalDataObject). Basic format validation is performed to ensure field integrity. If validation fails, a retry mechanism is triggered or the data is marked as abnormal. Validation rules include: non-empty checks for required fields, data type checks, and numeric range checks.
[0073] Step S5: Synchronize to the preprocessing module. The validated data is synchronized to the data preprocessing module in real time via a Kafka message queue to ensure data timeliness. The message body includes the data content and source identifier for easy traceability later.
[0074] Through the above steps, the system not only collects data but also completes preliminary intelligent processing, laying the foundation for subsequent in-depth analysis and significantly improving the usability of multimodal data. The anti-scraping adaptation unit monitors response codes in real time during this process. If a 403 error or CAPTCHA is encountered, it automatically switches the IP address or adjusts the request header fingerprint (such as Canvas fingerprint or WebGL fingerprint) to ensure that the data collection success rate remains above 95%.
[0075] Example 3: Detailed Implementation of Data Preprocessing and Storage
[0076] This embodiment details the collaborative work between the data preprocessing module and the data storage module, focusing on entity alignment and quantum encrypted storage.
[0077] In the data preprocessing stage, the entity alignment unit employs an entity parsing algorithm based on multi-feature fusion. Specifically, this includes:
[0078] Step A: Extract the name, brief introduction, avatar hash value, authentication information, and associated backlinks of each platform entity.
[0079] Step B: Calculate the edit distance similarity and semantic vector similarity of the names. Use the BERT model to convert the names into vectors, calculate the cosine similarity, and solve the aliasing problem (e.g., "XX Company" and "XX Group").
[0080] Step C: Calculate the perceptual hash (pHash) distance of the avatars to identify visual identity.
[0081] Step D: Construct a weighted similarity scoring formula. When the overall score exceeds a preset threshold, the entity is identified as the same entity and assigned a globally unique identifier (GlobalUID). For example, Douyin account A and WeChat account B can be associated with the same brand UID to achieve cross-platform data aggregation. The weighting coefficient can be dynamically adjusted according to the authentication type, with enterprise authentication having a higher weight than personal authentication.
[0082] During the data storage phase, the encrypted storage unit implements end-to-end quantum encryption protection. Considering the future threat of quantum computing to traditional RSA / ECC algorithms, this system employs a quantum-resistant cryptographic algorithm (Post-Quantum Cryptography, PQC), such as a lattice-based cryptography algorithm (specifically, either Kyber or Dilithium algorithms). The specific process is as follows:
[0083] 1. Key Generation: The system has a built-in quantum key generator (or simulated quantum random number generator) to generate key pairs with high entropy values. The key pair consists of a public key and a private key, and the private key must not be released outside the repository.
[0084] 2. Data Classification: Based on data sensitivity, data is classified into public, internal, and core levels. Core-level data (such as user phone numbers, identity information, and core operational strategies) must be encrypted using quantum encryption algorithms.
[0085] 3. Encrypted Storage: Before data is written to the database, the encryption service unit encrypts it using the PQC algorithm. The ciphertext is stored in the database, while the key is stored in a separate Hardware Security Module (HSM), physically isolated from the database. The database employs a sharding strategy, using either UserID or a timestamp as the sharding key to ensure even data distribution.
[0086] 4. Access Decryption: When an authorized user requests data, the system verifies permissions and obtains the key from the HSM through a secure channel for decryption. The plaintext key is not stored in memory. The decryption process takes place in a secure enclave environment to prevent memory dump attacks.
[0087] In addition, the hot data storage unit uses a Redis cluster with an expiration policy to automatically clean up hot data older than 7 days and move it to cold storage. The cold data storage unit uses compression storage technology (such as ZSTD compression) to reduce storage costs. All storage operations are logged to ensure data flow is traceable. For index building, inverted indexes are built for text data, and HNSW indexes are built for vector data to support mixed retrieval.
[0088] Example 4: Model Building for the Data Analysis Module
[0089] like Figure 3 As shown in this embodiment, the construction of the Transformer-LSTM hybrid model in the trend prediction unit of the data analysis module is described in detail. This model aims to solve the problem that traditional time series models cannot capture long-distance dependencies and external interference.
[0090] Model input layer: Receives multi-source feature data, including historical readership sequences, interaction sequences, external search indices (such as Baidu Index), holiday factors, and content feature vectors (provided by the multimodal acquisition unit). The input data is first normalized to eliminate the influence of units.
[0091] The Transformer Encoder layer utilizes the Transformer's self-attention mechanism to capture global dependencies within the input sequence. For example, it identifies the long-term association between "Friday" and "high readership," unaffected by intermediate data noise. Transformer layers can be computed in parallel, improving training efficiency. The multi-head attention mechanism allows the model to focus on different information locations in different representation subspaces.
[0092] Temporal Layer (LSTM): The context vector output by the Transformer is input into the LSTM (Long Short-Term Memory) layer. LSTM excels at capturing local fluctuations and trend inertia in time series. Information flow is controlled through forget gates, input gates, and output gates, remembering long-term trends and forgetting short-term noise. The number of LSTM units is set to 128 or 256, and the number of layers is 2 to prevent overfitting.
[0093] Output layer: Outputs the predicted popularity value for a specific future time period (e.g., the next 24 hours) through a fully connected layer, along with confidence intervals (e.g., upper and lower limits at 95% confidence). The activation function is either ReLU or Sigmoid, depending on the range of the output values.
[0094] Training process: Supervised learning is performed using historical data, and the loss function is a combination of mean squared error (MSE) and confidence interval penalty. The optimizer is AdamW, and the learning rate uses a Warmup strategy. The model is incrementally trained periodically (e.g., weekly) using new data to adapt to changes in the platform algorithm. A Dropout layer (scale 0.2) is introduced to enhance generalization ability.
[0095] Furthermore, the sentiment computing unit employs a BERT-BiLSTM-CRF combined model. BERT extracts semantic vectors, BiLSTM captures contextual dependencies, and CRF performs sequence labeling to identify the evaluation objects (such as "logistics" and "packaging") and their corresponding sentiment tendencies (positive / negative) in the reviews. This enables the system to not only know "users are dissatisfied" but also "users are dissatisfied with logistics," providing accurate basis for decision-making. The model is fine-tuned using a labeled industry corpus to improve domain adaptability.
[0096] Example 5: Decision Push and Feedback Closed Loop
[0097] like Figure 5 As shown in the figure, this embodiment elaborates on the workflow of the feedback closed-loop unit in the decision push module, which is the key to realizing system intelligence.
[0098] Step 1: Strategy Generation. Based on the analysis results, the system generates an initial strategy library. For example, if the trend prediction module shows "high traffic on weekends," the strategy generation unit suggests "increasing the posting frequency on weekends"; if sentiment analysis shows "users like lotteries," it suggests "increasing interactive benefits." Strategy generation uses a multi-armed bandit algorithm to balance exploration and exploitation, recommending known efficient strategies while also trying new ones.
[0099] Step 2: Strategy Execution and Data Collection. Operations personnel adopt the suggestions (or partially adopt them) and publish the content. The system automatically tracks the entire data chain of this publication (exposure, clicks, conversions, retention). The data collection module captures this post-execution data in real time and tags it with the strategy for easy attribution later.
[0100] Step 3: Effect Evaluation and Error Calculation. Compare the actual data with the system's predicted data. Calculate accuracy metrics (such as MAPE). If the actual effect is better than the prediction, analyze which features were effective (e.g., capitalizing on a sudden trending topic); if it is lower than the prediction, analyze whether it is a content quality issue or an external environmental issue. Introduce an A / B testing framework to compare the differences in effects between different strategy groups to ensure the scientific nature of the evaluation.
[0101] Step 4: Model Backpropagation and Optimization. Utilizing reinforcement learning, the operational effectiveness is used as the reward function. If the strategy is effective, the weight of that strategy's feature is increased; if it fails, the weight is decreased. For example, if the "weekend release" strategy performs poorly three times consecutively, the system automatically lowers the priority of this suggestion and tries recommending "weekday release." The reward function design considers long-term value (such as user retention) rather than just short-term clicks.
[0102] Step 5: Personalized Recommendations. Based on the optimized model, the system pushes differentiated suggestions to operations personnel in different roles. Title suggestions are pushed to content editors, and budget suggestions are pushed to campaign managers. Push channels include system in-app messages, emails, and WeChat webhooks.
[0103] Through this closed loop, the system is no longer a static tool, but an intelligent assistant that continuously evolves with the operation process, becoming smarter the more it is used. After each closed loop iteration, the model version is updated, while the old version is retained for rollback.
[0104] Example 6: Implementation Details of the Security Protection Module
[0105] This embodiment details the specific implementation of the security protection module to ensure that the system meets the highest security standards.
[0106] 1. Encrypted Transmission: All inter-module communication and external API calls use the HTTPS / TLS 1.3 protocol. For core data transmission, an application-layer quantum encryption layer is added, ensuring data security even if TLS is cracked. The key negotiation process uses the ECDHE algorithm, providing forward security.
[0107] 2. Access Control: Based on the RBAC (Role-Based Access Control) model, providing fine-grained access control. For example, regular operations staff can only view data, not export it; administrators can configure the system, but cannot view sensitive user information in plaintext. Multi-factor authentication (MFA) is supported, requiring a password and mobile verification code for login. Permission changes take effect immediately without requiring a service restart.
[0108] 3. Anomaly Detection: Deploy a User Behavior Analysis (UEBA) system to monitor abnormal operations. For example, if an account queries a large amount of sensitive data within a short period of time, or logs in outside of working hours, the system will automatically trigger an alert and freeze the account. The detection algorithm uses Isolation Forest to identify anomalies.
[0109] 4. Data Anonymization: Sensitive information is dynamically anonymized at the presentation layer. For example, a mobile phone number may be displayed as 138****1234. This anonymization can only be restored under authorized audit scenarios. Anonymization rules are configurable to meet different compliance requirements.
[0110] 5. Audit Logs: All operation logs are stored in WORM (WriteOnceReadMany) storage media to prevent tampering, and are retained for no less than 6 months to meet the requirements of the Cybersecurity Law. Blockchain technology is introduced to store log hashes, ensuring that the logs are immutable and verifiable.
[0111] 6. Privacy-preserving computation: Supports federated learning mode, which allows for joint training of models with other institutions without sharing the original data, thereby further improving model accuracy while protecting data privacy.
[0112] Example 7: Workflow in a Typical Application Scenario
[0113] To further illustrate the practicality of the present invention, two typical application scenarios are described below.
[0114] Scenario 1: Promotion of a new product launch by a certain brand.
[0115] 1. Pre-launch Period (T-7 days): The system, through the trend prediction module, detected an increase in the popularity of the topic "summer sun protection." The strategy recommendation unit suggested: "Launching new products in conjunction with the sun protection topic, with keywords including 'refreshing' and 'non-greasy.'" The operations staff adopted the suggestion and created content. The data collection module began monitoring the volume of related topics from competitors.
[0116] 2. Peak Period (Day T): New product content is simultaneously released on multiple platforms. A real-time dashboard displays traffic inflow across each platform. The risk warning unit detects a small number of negative "allergy" comments on a certain platform. The system automatically flags these comments and notifies customer service to intervene. The propagation path tracking module discovers that a KOL's repost brought in 30% of the traffic; the system automatically records the KOL's weight, providing a basis for future collaborations.
[0117] 3. Duration (T+7 days): The sentiment analysis module aggregates reviews from across the internet to generate a "Product Reputation Report." It shows that "Packaging" has a high approval rating, while "Price" is the subject of much controversy. The strategy recommendation unit suggests: "For price controversies, issue limited-time coupons or emphasize cost-effectiveness in copywriting." The entity alignment module aggregates sales leads collected from various channels into the CRM system and calculates the final ROI.
[0118] 4. Review Period (T+30 days): The system automatically generates a case closure report, comparing predicted data with actual data. A feedback loop mechanism updates model parameters and optimizes the prediction accuracy for the next activity.
[0119] Scenario 2: Management of sudden public opinion crisis.
[0120] 1. Monitoring Period: The system monitors brand keywords 24 / 7. At 2 AM one day, it detected a surge in the number of negative posts on a certain social media platform, with the growth rate exceeding the threshold.
[0121] 2. Warning Period: The risk warning unit immediately triggers a red alert, notifying the brand's public relations manager via phone and SMS. The system automatically generates a public opinion briefing, including the source of the negative information, core viewpoints, and dissemination path.
[0122] 3. Handling Period: The decision-making unit pushes suggestions on response strategies, such as "official statement" or "third-party endorsement." The system monitors the trend of public opinion after the response in real time. If negative sentiment decreases, the handling is marked as effective; if it continues to rise, it is recommended to upgrade the handling plan.
[0123] 4. Recovery period: After the public opinion subsides, the system generates a crisis review report, analyzes the causes of the crisis, and suggests optimizing product or service processes to prevent similar incidents from happening again.
[0124] Summarize
[0125] This invention constructs a complete data collection and analysis system for new media operation and management, achieving end-to-end intelligent processing from data collection, processing, and analysis to decision feedback. Its core advantages lie in its dynamic and adaptive data collection capabilities, deep fusion and analysis of multimodal data, high-precision prediction capabilities using Transformer-LSTM, and high security through quantum encryption. This system can significantly reduce the labor costs of new media operations, improve the scientific rigor and timeliness of decision-making, effectively mitigate public opinion risks, and possesses extremely high commercial value and application prospects. The system supports private deployment and SaaS service models, flexibly adapting to different customer needs.
[0126] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A data collection and analysis system for new media operation and management, characterized in that: With end-to-end data closed-loop collaboration as its core concept, the invention includes a data acquisition module, a data preprocessing module, a data storage module, a data analysis module, a decision push module, a security protection module, and a system management module. The modules are interconnected and the data flows seamlessly, realizing end-to-end closed-loop management of new media operation data collection, processing, storage, analysis, decision-making, and security protection. It is applicable to the full-process management of operation data for short video, social, news, and live streaming new media platforms.
2. The data collection and analysis system for new media operation and management according to claim 1, characterized in that: The data acquisition module adopts a fusion acquisition scheme of API + distributed crawler + multimodal intelligent acquisition, including an API interface unit, a distributed crawler unit, a multimodal intelligent acquisition unit, an anti-crawling adaptation unit, and an acquisition scheduling unit; the multimodal intelligent acquisition unit realizes the integration of data acquisition and preliminary feature extraction, and collects multimodal data such as text, images, videos, bullet comments, and emoticons.
3. The data collection and analysis system for new media operation and management according to claim 2, characterized in that: The multimodal intelligent acquisition unit extracts text from images using OCR technology, transcribes video / audio speech using ASR technology, and uses a lightweight CNN model to extract visual features from images, transforming unstructured data into semi-structured feature data.
4. The data collection and analysis system for new media operation and management according to claim 1, characterized in that: The data preprocessing module includes a data cleaning unit, an entity alignment unit, a multimodal fusion unit, and a data standardization unit. The entity alignment unit establishes a globally unique ID mapping for the same entity on different platforms, and the multimodal fusion unit realizes the association and fusion of text, image, and video features.
5. The data collection and analysis system for new media operation and management according to claim 1, characterized in that: The data storage module adopts a hierarchical storage strategy, including a hot data storage unit, a cold data storage unit, and an encrypted storage unit. The encrypted storage unit uses a quantum computing-resistant encryption algorithm to encrypt and store core sensitive data, which is stored hierarchically from ordinary data.
6. The data collection and analysis system for new media operation and management according to claim 1, characterized in that: The data analysis module includes a sentiment computing unit, a propagation path tracking unit, a user profile construction unit, and a trend prediction unit. The trend prediction unit uses a Transformer-LSTM hybrid model, combining historical data and external indices to predict the trend of topic popularity and provides a confidence interval for the prediction results.
7. A data collection and analysis system for new media operation and management according to claim 6, characterized in that: The sentiment computing unit uses a BERT-BiLSTM-CRF combined model to achieve fine-grained sentiment polarity judgment and emotion category recognition for text comments and video bullet comments, while also supporting aspect-level sentiment analysis.
8. The data collection and analysis system for new media operation and management according to claim 1, characterized in that: The decision push module includes a strategy generation unit, an automated report unit, a risk warning unit, and a feedback closed-loop unit. The feedback closed-loop unit records the decision execution effect and inputs the effect data back to the data analysis module to optimize the analysis model, forming a decision-execution-optimization closed loop.
9. A data collection and analysis system for new media operation and management according to claim 1, characterized in that: The security protection module runs through the entire system chain and includes a transmission encryption unit that uses national cryptographic algorithms and quantum key distribution technology, a fine-grained access control unit based on the RBAC model, and an audit log unit that enables traceable operations. It also has data desensitization, anomaly detection, and federated learning privacy computing capabilities.
10. A data collection and analysis system for new media operation and management according to any one of claims 1-9, characterized in that: The system adopts a cloud-based distributed architecture, which is divided into an infrastructure layer, a data layer, a service layer, and an application layer. The system supports private deployment and SaaS service mode, and is suitable for application scenarios such as new media operation of small and medium-sized enterprises, multi-account management of MCN agencies, and new media matrix operation of brands.