A data processing system for network audiovisual new media
By applying multimodal fusion recognition and a dynamic compliance library, the problem of incomplete multimodal data parsing in online audiovisual new media systems has been solved, enabling efficient compliance review and personalized recommendations, and improving the overall processing capabilities of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA WEST NORMAL UNIVERSITY
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-21
Smart Images

Figure CN122437962A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more particularly to a data processing system for online audiovisual new media. Background Technology
[0002] With the rapid development of the online audiovisual new media industry, audiovisual content is showing trends of diversification, massive volume, and multimodality. Users are constantly raising their requirements for content quality, transmission efficiency, personalized experience, and compliance. Currently, online audiovisual new media data processing mostly adopts a traditional single-module architecture, with each link operating independently and poor data interoperability. This makes it difficult to adapt to the efficient processing needs of massive and diverse audiovisual data, and has gradually exposed many shortcomings.
[0003] At the semantic analysis level, most existing new media data processing systems can only independently analyze single text or audio / video modalities, failing to achieve collaborative fusion analysis of multimodal data such as audio tracks, images, and text. This results in insufficient and in-depth analysis of audiovisual content, making it easy for extracted keywords, thematic tags, and sentiment information to be biased. It is difficult to accurately capture the core semantics and deeper connotations of the content, which in turn affects the effectiveness of subsequent content distribution, personalized recommendations, and other processes.
[0004] At the compliance review level, the existing system's compliance database is mostly statically set up and lacks a dynamic update mechanism. It cannot receive the latest policies, regulations and review standards issued by external regulatory authorities in real time, resulting in the review standards lagging behind industry regulatory requirements. At the same time, the review method still relies heavily on manual review, which is not only inefficient and unable to meet the rapid review needs of massive audiovisual content, but also prone to omissions and misreviews due to differences in human subjective judgment, bringing potential compliance risks to the platform.
[0005] Application content
[0006] This application aims to address, at least to some extent, the technical problems in the related art.
[0007] To achieve the above objectives, this application proposes a data processing system for online audiovisual new media, comprising: a data acquisition module that collects raw audio-visual, text, and bullet-screen audiovisual media data from multiple data sources, including new media platforms, live streaming sources, on-demand libraries, and user upload terminals; a data preprocessing module, communicatively connected to the data acquisition module, that performs deduplication, noise reduction, format unification, and format conversion on the raw audiovisual media data to generate preprocessed audiovisual data; a semantic analysis engine, communicatively connected to the data preprocessing module, that performs audio-visual transcription, deep semantic analysis of text, and image content recognition on the preprocessed audiovisual data to extract keywords, topic tags, classification information, and sentiment information of the audiovisual content; and a content compliance review module, communicatively connected to the semantic analysis engine, that combines the semantic analysis results with preset parameters. The compliance library automatically identifies and marks sensitive, illegal, and vulgar content in audiovisual content. An adaptive streaming media encoder, communicating with both the data preprocessing module and the content compliance review module, performs multi-level adaptive transcoding on the preprocessed audiovisual data based on review approval instructions, semantic tags, and real-time network bandwidth and terminal type information, generating multi-bitrate, multi-protocol streaming media streams adapted to different terminals and network environments. A distributed media storage module, communicating with the adaptive streaming media encoder, performs distributed storage and indexing management of the multi-bitrate streaming media streams, semantic tags, and review results. An intelligent content distribution module, communicating with the distributed media storage module, schedules the corresponding audiovisual data stream to the nearest node for distribution based on user location, network status, and terminal performance.
[0008] In addition, the application may also include the following additional technical features:
[0009] Specifically, it also includes a personalized recommendation engine, which is communicatively connected to the distributed media storage module and the semantic analysis engine, respectively. Based on the user's viewing history, interaction behavior, and the semantic tags and sentiment information, it constructs a user interest profile, recalls matching audiovisual content from the distributed media storage module, generates a personalized recommendation list, and performs online learning and model updates based on user feedback.
[0010] Specifically, the semantic analysis engine also includes a multimodal fusion recognition submodule, which is used to simultaneously analyze audio tracks, on-screen text, bullet screen text, and keyframe images in the same audiovisual data, and generate a comprehensive semantic description of the audiovisual data content through cross-modal feature alignment and fusion.
[0011] Specifically, the content compliance review module includes a dynamic compliance database update submodule, which communicates with external regulatory data sources and manual review platforms to receive the latest policies, regulations, and review standards in real time or periodically, and automatically updates the sensitive word database, image feature database, and audio feature database in the preset compliance database accordingly, while incorporating the feedback results of manual review into the update mechanism.
[0012] Specifically, the adaptive streaming media encoder includes a dynamic decision-making submodule for encoding parameters, which is communicatively connected to the semantic analysis engine. It is used to dynamically adjust the bitrate control algorithm, keyframe interval, and encoding complexity based on the content type tags output by the semantic analysis engine, so as to optimize the encoding efficiency and image quality of specific types of content.
[0013] Specifically, it also includes an interactive bullet screen processing module, which is communicatively connected to the data acquisition module and the distributed media storage module. It is used to receive, filter, and review bullet screen data posted by users in real time, and associate the reviewed bullet screen data with the corresponding audiovisual data timestamp for synchronous presentation on the playback terminal.
[0014] Specifically, the distributed media storage module adopts IPFS or a content-addressable distributed storage architecture to generate a unique hash identifier for the content of the audiovisual data itself, which is used for data deduplication storage, integrity verification and P2P accelerated distribution.
[0015] Specifically, the intelligent content distribution module includes an edge node pre-deployment sub-module, which is used to predict future hot topics based on the hot topic tags extracted by the semantic analysis engine and the user interest profile, and to pre-cache the bitstream of the hot topics to edge nodes close to the user.
[0016] In summary, the beneficial effects of the data processing system for online audiovisual new media proposed in this application are as follows:
[0017] 1. It can achieve collaborative fusion analysis of multimodal data such as audio tracks, images, and text. Through cross-modal feature alignment and fusion technology, it can perform comprehensive analysis of audiovisual content, effectively avoiding the bias caused by single-modal analysis. It can accurately extract keywords, theme tags, and sentiment information of the content, accurately capture the core semantics and deep connotations of audiovisual content, provide accurate data support for subsequent stages, and significantly improve the processing effect of each subsequent stage.
[0018] 2. It can receive review standards issued by external regulatory authorities in real time or periodically, realize automatic dynamic updates of the compliance database, ensure that the review standards are in sync with industry regulatory requirements, and realize the automation of compliance review by relying on the accurate semantic results output by the semantic analysis engine. This greatly reduces the reliance on manual review, significantly improves the review efficiency of massive audio-visual content, meets the needs of rapid review, and avoids the problems of missed or incorrect reviews caused by differences in human subjective judgment, effectively reducing the compliance risks faced by the platform. Attached Figure Description
[0019] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0020] Figure 1 This is a flowchart of a data processing system for online audiovisual new media according to this application;
[0021] Figure 2 This is a schematic diagram of a sub-module of a data processing system for new online audiovisual media according to this application;
[0022] Figure 3 This is a flowchart of a personalized recommendation engine for a data processing system for online audiovisual new media, as described in this application. Detailed Implementation
[0023] To make the technical means, inventive features, objectives, and effects of this application readily understandable, this application is further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0024] The present application will now be described in further detail with reference to the accompanying drawings.
[0025] like Figure 1 As shown in the figure, an embodiment of the present application provides a data processing system for network audiovisual new media, which includes a data acquisition module that collects raw audio-visual, text, and bullet screen audiovisual media data from multiple data sources, including new media platforms, live streaming sources, on-demand libraries, and user upload terminals.
[0026] Specifically, a multi-source heterogeneous data access architecture is adopted, which connects to mainstream new media platforms (such as short video and live streaming platforms) through API interfaces. Live streaming source data is captured in real time based on protocols such as RTMP and HTTP-FLV. On-demand inventory content is obtained through direct database connection or file reading. At the same time, a data receiving channel for user upload is built, which supports real-time acquisition of raw data in various formats (such as MP4, JPG, TXT, etc.). A data synchronization mechanism ensures the real-time performance and integrity of data from each data source, avoiding data omission or delay.
[0027] RTMP (Real-Time Messaging Protocol) is a dedicated streaming media transmission protocol based on TCP. Its core purpose is to achieve stable real-time transmission of audio and video streams, and it is one of the classic push and pull protocols in live streaming scenarios.
[0028] HTTP-FLV is a streaming media transmission method based on the HTTP protocol. Essentially, it transmits FLV format audio and video streams through the standard HTTP protocol. It can be understood as an HTTP-adapted version of the RTMP protocol, which can achieve streaming media transmission without relying on a dedicated transmission protocol.
[0029] This module boasts comprehensive data source coverage, encompassing platform-owned content, third-party collaborative content, and user-generated content, thus meeting the diverse content needs of online audiovisual new media. Its flexible acquisition methods adapt to the characteristics of different data sources, enabling real-time acquisition of live content, batch capture of on-demand content, and instant reception of user-uploaded content. It supports simultaneous acquisition of multiple data types, eliminating the need for separate acquisition channels for different data types, reducing system deployment costs, and improving data acquisition efficiency.
[0030] The data preprocessing module communicates with the data acquisition module to perform deduplication, noise reduction, format unification, and format conversion on the original audiovisual media data, generating preprocessed audiovisual data.
[0031] Specifically, digital signal processing technology is used to filter environmental noise and current noise in audio and video, and to repair blur and noise in text and image content. The format parsing engine identifies raw data of different formats and converts them into a system-compatible standard format (e.g., audio and video are uniformly converted to MP4 format, and text and images are uniformly converted to JPG format). At the same time, according to subsequent processing requirements, parameters such as resolution, frame rate, and bit rate are adjusted to ensure the consistency of data format.
[0032] This module effectively eliminates redundant data, reduces storage pressure, and avoids efficiency losses caused by duplicate data in subsequent processing; noise reduction improves the quality of audiovisual content and enhances the user viewing experience; image and text restoration ensures content clarity; format unification and conversion eliminate data format heterogeneity, providing standardized data input for subsequent processing, reducing the adaptation difficulty of each module, and improving the overall system processing efficiency.
[0033] The semantic analysis engine communicates with the data preprocessing module to perform audio-visual transcription, deep semantic analysis of text, and image content recognition on the preprocessed audio-visual data, extracting keywords, topic tags, classification information, and sentiment information from the audio-visual content.
[0034] Specifically, it integrates three major technologies: speech-to-text, natural language processing, and computer vision. Speech-to-text technology transcribes audio content from audio and video into text. Computer vision technology identifies key elements such as people, scenes, and objects in the video and converts them into text descriptions. Natural language processing technology then performs in-depth analysis on the transcribed text, image and text content, and bullet screen text, including word segmentation, part-of-speech tagging, entity recognition, and semantic association analysis, thereby extracting core keywords and topic tags. A sentiment dictionary and machine learning model determine the sentiment tendency (positive, negative, neutral) of the content, and the content is classified in conjunction with a pre-set content classification system.
[0035] This engine enables deep semantic mining of various types of audiovisual data, breaking through the limitations of single text parsing and comprehensively capturing core information from audio, video, images, and bullet comments; the extraction of keywords, topic tags, and classification information provides accurate data support for subsequent content distribution and personalized recommendations; sentiment analysis can quickly grasp user feedback trends on content, providing a basis for content operation decisions.
[0036] The content compliance review module communicates and connects with the semantic analysis engine. Combining the semantic analysis results with the preset compliance library, it automatically identifies and marks sensitive, illegal, and vulgar content in audiovisual content.
[0037] Specifically, a pre-defined compliance library is constructed, including a sensitive word library, a violation image feature library, and a vulgar content feature library. Keywords, topic tags, sentiment information output by the semantic analysis engine, as well as image recognition results and audio / video transcribed text, are compared and matched with the features in the compliance library in real time. The rule engine and machine learning model determine whether the content is sensitive, violates regulations, or is vulgar. Suspected violations are marked, and confirmed violations are graded (e.g., minor violations, major violations). At the same time, the reasons for violations and the characteristics of violations are recorded.
[0038] This module combines semantic analysis results to provide a more comprehensive review scope. It can identify not only text violations but also violations in images and audio, avoiding missed or incorrect reviews. The preset compliance library can be flexibly adjusted to quickly adapt to different regulatory requirements, ensuring consistency in review standards. Violation markers are clear, facilitating subsequent manual review and handling of violations, thus reducing risks.
[0039] The adaptive streaming media encoder communicates with the data preprocessing module and the content compliance review module respectively. Based on the review approval instructions, semantic tags, and real-time network bandwidth and terminal type information, it performs multi-level adaptive transcoding on the preprocessed audiovisual data to generate multi-bitrate and multi-protocol streaming media streams that are adapted to different terminals and network environments.
[0040] Specifically, it receives approval instructions from the content compliance review module, obtains content type tags (such as high-definition video, short video, and audio) output by the semantic analysis engine, and simultaneously collects information such as network bandwidth, terminal type, and screen resolution of user terminals in real time through the network monitoring module. Based on mainstream encoding standards, it dynamically adjusts encoding parameters, sets multiple bitrate levels (such as high-definition, standard-definition, and smooth), and adopts different streaming media protocols to adaptively transcode the pre-processed audiovisual data to ensure that the transcoded bitstream can adapt to the playback capabilities of different terminals and the transmission capabilities of different network environments.
[0041] This encoder achieves adaptive bitstream adaptation, solving the problems of stuttering and blurring in audiovisual content playback under different terminals and network environments, thus improving the user's viewing experience. Multiple bitstream levels can be automatically switched according to the user's network status, ensuring the viewing quality for high-definition users while also taking into account the basic playback needs of low-bandwidth users. It dynamically adjusts encoding parameters based on content type tags to optimize encoding efficiency, reducing bitstream size and storage and transmission costs while ensuring image quality.
[0042] The distributed media storage module communicates with the adaptive streaming media encoder to perform distributed storage and indexing management of multi-bitrate streaming media streams, semantic tags, and review results.
[0043] Specifically, large streaming media files are divided into multiple small segments using data sharding technology and stored on different nodes. At the same time, a unified indexing system is established to associate the streaming media bitstream with corresponding semantic tags and review results. Load balancing technology is used to distribute the load across storage nodes to avoid overloading a single node. A data redundancy backup mechanism ensures data security and recoverability and supports fast data querying and retrieval.
[0044] This module's storage capacity is flexibly expandable, capable of meeting the storage needs of massive amounts of audiovisual data; the distributed architecture improves data access speed, allowing users to retrieve data from the nearest storage node, reducing access latency; data redundancy backup ensures that data will not be lost due to the failure of a single node, improving the security and reliability of data storage; index management enables fast data association queries, facilitating rapid subsequent retrieval of related data and improving the overall system response speed.
[0045] The intelligent content distribution module communicates with the distributed media storage module and, based on the user's location, network status, and terminal performance, schedules the corresponding audio-visual data stream to the nearest node to complete the distribution.
[0046] Specifically, the system collects user location information, network bandwidth, terminal performance and other data in real time. Through scheduling algorithms, it analyzes the user's optimal access node and schedules the corresponding multi-bitrate streams in the distributed media storage module to edge nodes close to the user. At the same time, based on the user's network status and terminal performance, it automatically matches the most suitable stream for the user, realizing the local distribution of audio-visual content and reducing the latency and stuttering caused by cross-network transmission.
[0047] This module significantly reduces content transmission latency by distributing content based on proximity, improves user playback startup speed, and enhances user experience; it matches the optimal bitstream based on user terminal and network status to avoid resource waste while ensuring playback quality; and it improves the system's distribution capacity and stability by rationally allocating distribution pressure to each node through load balancing scheduling.
[0048] In one embodiment of this application, such as Figure 3 As shown, it also includes a personalized recommendation engine, which communicates with the distributed media storage module and the semantic analysis engine respectively. Based on the user's viewing history, interaction behavior, semantic tags, and sentiment information, it constructs a user interest profile, recalls matching audiovisual content from the distributed media storage module, generates a personalized recommendation list, and performs online learning and model updates based on user feedback.
[0049] It should be noted that, employing collaborative filtering algorithms and deep learning models, the system first collects user interaction data such as viewing history, likes, comments, favorites, and skips. This data is then combined with semantic tags and sentiment information output by the semantic analysis engine to extract user interest features and construct a multi-dimensional user interest profile (e.g., interest domains, preferred content types, and sentiment preferences). Based on this interest profile, a content retrieval algorithm filters audiovisual content from the distributed media storage module that matches the user's semantic tags and sentiment preferences. A ranking algorithm optimizes the order of the recommendation list. Simultaneously, user feedback on the recommended content is collected in real time (e.g., clicks, dwell time, unfollowing), and the recommendation model parameters are updated through an online learning mechanism to continuously optimize recommendation accuracy.
[0050] Personalized recommendations are highly accurate, precisely matching user interests and needs, increasing click-through rates and dwell time, and enhancing user engagement; the online learning mechanism can adapt to changes in user interests in real time, avoiding fixed recommendations and improving the freshness and effectiveness of recommendations; combined with semantic tags and sentiment analysis, it breaks through the limitations of traditional recommendations based on viewing history, can uncover users' potential interests, and expand users' content consumption scenarios.
[0051] In one embodiment of this application, such as Figure 2As shown, the semantic analysis engine also includes a multimodal fusion recognition submodule, which is used to simultaneously analyze audio tracks, on-screen text, bullet screen text, and keyframe images in the same audiovisual data. Through cross-modal feature alignment and fusion, it generates a comprehensive semantic description of the audiovisual data content.
[0052] It should be noted that multimodal fusion technology is used to extract features separately from the audio track (speech content), on-screen text (subtitles, text in the screen), bullet screen text, and keyframe images of the audiovisual data. Through feature mapping, the features of different modalities are transformed into the same feature space. An attention mechanism is used to achieve cross-modal feature alignment, eliminating semantic bias between different modal data. Then, the features of each modality are integrated through a fusion algorithm to generate a comprehensive semantic description containing multi-dimensional information such as speech, text, and images, thus fully capturing the core meaning of the audiovisual content.
[0053] In one embodiment of this application, such as Figure 2 As shown, the content compliance review module includes a dynamic compliance database update submodule, which communicates with external regulatory data sources and manual review platforms to receive the latest policies, regulations, and review standards in real time or periodically, and automatically updates the sensitive word database, image feature database, and audio feature database in the preset compliance database accordingly. At the same time, the feedback results of manual review are incorporated into the update mechanism.
[0054] It should be noted that a real-time connection channel is established with external regulatory data sources (such as policy documents and violation cases issued by regulatory authorities) to obtain the latest policies, regulations, and review standards in real time or periodically through API interfaces or file synchronization. At the same time, it connects with the manual review platform to collect cases of missed or incorrect reviews, as well as newly added violation characteristics, discovered during the manual review process. The latest information is analyzed through a rule engine to automatically update the sensitive words, violation image features, and audio features in the preset compliance database. The results of manual review feedback are transformed into the basis for updating the compliance database, forming a two-way update mechanism of regulatory updates and manual feedback to ensure that the compliance database is consistent with the latest review requirements.
[0055] In one embodiment of this application, such as Figure 2 As shown, the adaptive streaming media encoder includes a dynamic decision-making submodule for encoding parameters, which communicates with the semantic analysis engine. It is used to dynamically adjust the bitrate control algorithm, keyframe interval, and encoding complexity based on the content type tags output by the semantic analysis engine, in order to optimize the encoding efficiency and image quality of specific types of content.
[0056] It should be noted that the semantic analysis engine receives content type tags (such as action videos, static images and text, audio programs, and short videos), and presets corresponding encoding parameter adjustment rules based on the characteristics of different content types. For example, action videos have fast-changing scenes, requiring efficient bitrate control algorithms to shorten keyframe intervals and increase encoding complexity to ensure image quality; static images and text or audio programs have small scene changes, allowing for reduced encoding complexity and extended keyframe intervals to improve encoding efficiency while maintaining quality. The dynamic encoding parameter decision submodule automatically matches the corresponding adjustment rules based on the content type tags and adjusts the encoding parameters in real time to achieve personalized encoding.
[0057] In one embodiment of this application, such as Figure 1 As shown, it also includes an interactive bullet screen processing module, which communicates with the data acquisition module and the distributed media storage module. It is used to receive, filter, and review bullet screen data posted by users in real time, and associate the reviewed bullet screen data with the corresponding audiovisual data timestamp for storage, so that the playback terminal can present it synchronously.
[0058] It should be noted that a real-time bullet screen receiving channel is established to receive bullet screen text data posted by users. First, invalid bullet screens (such as blank bullet screens and duplicate bullet screens) are removed through a filtering algorithm. Then, the content compliance review module is connected to the compliance standards to review the bullet screen text for sensitive, illegal, and vulgar content. For bullet screen data that passes the review, the corresponding audiovisual content identifier and publication time are extracted and accurately associated with the timestamp of the audiovisual data. The associated bullet screen data is stored in a distributed media storage module. At the same time, an index relationship between bullet screens and audiovisual content is established so that the playback terminal can call the corresponding bullet screen data in real time according to the playback progress, realizing the synchronous presentation of bullet screens and audiovisual content.
[0059] Bullet comment filtering and compliance review prevent illegal bullet comments from affecting user viewing experience and platform compliance; timestamp association enables accurate synchronization between bullet comments and audiovisual content, improving the accuracy of bullet comment presentation; and integration with distributed storage modules ensures secure storage and rapid retrieval of bullet comment data, supporting concurrent processing of massive bullet comments and handling peak interaction scenarios.
[0060] In one embodiment of this application, such as Figure 2 As shown, the distributed media storage module adopts IPFS or a content-addressable distributed storage architecture to generate a unique hash identifier for the content of the audiovisual data itself, which is used for data deduplication storage, integrity verification and P2P accelerated distribution.
[0061] It should be noted that IPFS (InterPlanetary File System) is a decentralized, content-addressed, peer-to-peer distributed file storage and transmission protocol. Its goal is to build a more open, censorship-resistant, and highly available next-generation internet infrastructure. Using IPFS or content-addressed storage technologies, it no longer uses file paths as storage identifiers. Instead, it performs hash calculations on the content of the audiovisual data itself to generate a unique hash value as the content identifier. When storing new data, its hash value is first calculated and compared with the hash values of already stored data to deduplicate the data and avoid duplicate storage. The hash value is used to verify data integrity; if data is tampered with, the hash value will change, allowing for rapid identification of data anomalies. Simultaneously, P2P (peer-to-peer) accelerated distribution is achieved based on hash identifiers, allowing users to obtain data fragments from multiple storage nodes, improving distribution speed.
[0062] The IPFS architecture features decentralization, preventing single-node failures from affecting data access, improving the stability and reliability of the storage system, and adapting to the storage and distribution needs of massive amounts of audio-visual data; hash identifiers ensure data integrity, effectively preventing data tampering and improving data security.
[0063] In one embodiment of this application, such as Figure 2 As shown, the intelligent content distribution module includes an edge node pre-deployment sub-module, which is used to predict future hot topics based on hot topic tags and user interest profiles extracted by the semantic analysis engine, and pre-cache the bitstream of the hot topics to edge nodes close to users.
[0064] It should be noted that a machine learning prediction model is used, combined with information such as trending topic tags (e.g., current hot topics and popular content types) extracted by the semantic analysis engine, user interest profiles, historical distribution data, and user growth trends, to analyze and predict audiovisual content that may become trending in the near future. The edge node pre-deployment submodule obtains the multi-bitrate stream of trending content from the distributed media storage module based on the prediction results and caches it in advance to edge nodes close to the user group, thus completing the content pre-deployment. When a user requests access to trending content, the data can be obtained directly from the nearest edge node without scheduling from the central node.
[0065] It should be noted that, in this document, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0066] The present application and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present application. The actual structure is not limited to this. In conclusion, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the present application, such design should fall within the protection scope of the present application.
Claims
1. A data processing system for new online audiovisual media, characterized in that, include: The data acquisition module collects raw audio-visual media data, including images, text, and bullet comments, from multiple data sources such as new media platforms, live streaming sources, on-demand libraries, and user upload terminals. The data preprocessing module is communicatively connected to the data acquisition module and performs deduplication, noise reduction, format unification, and format conversion on the original audiovisual media data to generate preprocessed audiovisual data. The semantic analysis engine communicates with the data preprocessing module to perform audio-visual transcription, deep semantic analysis of text, and image content recognition on the preprocessed audio-visual data, extracting keywords, topic tags, classification information, and sentiment information of the audio-visual content. The content compliance review module communicates with the semantic analysis engine and, in conjunction with the semantic analysis results and the preset compliance library, automatically identifies and marks audiovisual content that is sensitive, illegal, or vulgar. An adaptive streaming media encoder is communicatively connected to the data preprocessing module and the content compliance review module, respectively. Based on the review approval instruction, semantic tags, and real-time network bandwidth and terminal type information, it performs multi-level adaptive transcoding on the preprocessed audiovisual data to generate multi-bitrate and multi-protocol streaming media streams that are adapted to different terminals and network environments. The distributed media storage module is communicatively connected to the adaptive streaming media encoder and performs distributed storage and indexing management of the multi-bitrate streaming media streams, semantic tags, and review results. The intelligent content distribution module communicates with the distributed media storage module and schedules the audio-visual data of the corresponding bitstream to the nearest node for distribution based on the user's location, network status, and terminal performance.
2. The data processing system for new network audiovisual media according to claim 1, characterized in that, It also includes a personalized recommendation engine, which is connected to the distributed media storage module and the semantic analysis engine respectively. Based on the user's viewing history, interaction behavior and the semantic tags and sentiment information, it constructs a user interest profile, recalls matching audiovisual content from the distributed media storage module, generates a personalized recommendation list, and performs online learning and model updates based on user feedback.
3. The data processing system for new network audiovisual media according to claim 1, characterized in that, The semantic analysis engine also includes a multimodal fusion recognition submodule, which is used to simultaneously analyze audio tracks, on-screen text, bullet screen text, and keyframe images in the same audiovisual data. Through cross-modal feature alignment and fusion, a comprehensive semantic description of the audiovisual data content is generated.
4. The data processing system for new network audiovisual media according to claim 1, characterized in that, The content compliance review module includes a dynamic compliance database update submodule, which communicates with external regulatory data sources and manual review platforms. It is used to receive the latest policies, regulations, and review standards in real time or periodically, and automatically update the sensitive word database, image feature database, and audio feature database in the preset compliance database accordingly. At the same time, the feedback results of manual review are incorporated into the update mechanism.
5. A data processing system for new network audiovisual media according to claim 1, characterized in that, The adaptive streaming media encoder includes a dynamic decision-making submodule for encoding parameters, which is communicatively connected to the semantic analysis engine. It is used to dynamically adjust the bitrate control algorithm, keyframe interval, and encoding complexity based on the content type tags output by the semantic analysis engine, so as to optimize the encoding efficiency and image quality of specific types of content.
6. The data processing system for new network audiovisual media according to claim 1, characterized in that, It also includes an interactive bullet screen processing module, which is communicatively connected to the data acquisition module and the distributed media storage module. It is used to receive, filter, and review bullet screen data posted by users in real time, and associate the approved bullet screen data with the corresponding audiovisual data timestamp for synchronous presentation on the playback terminal.
7. A data processing system for new network audiovisual media according to claim 1, characterized in that, The distributed media storage module adopts IPFS or a content-addressable distributed storage architecture, and generates a unique hash identifier for the content of the audiovisual data itself, which is used for data deduplication storage, integrity verification and P2P accelerated distribution.
8. A data processing system for new network audiovisual media according to claim 1, characterized in that, The intelligent content distribution module includes an edge node pre-deployment sub-module, which is used to predict future hot topics based on the hot topic tags extracted by the semantic analysis engine and the user interest profile, and to pre-cache the bitstream of the hot topics to edge nodes close to the user.