Audio and video content identification method and device, medium and program product
By bypassing the capture of audio and video streams and using synchronization source identifiers and target microservices for separation and processing, the latency and device heating issues of multi-track real-time audio and video analysis are resolved, and efficient multi-track audio and video content recognition is achieved, meeting real-time requirements and improving the accuracy of violation detection.
Patent Information
- Application Number
- CN202511070466.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-10-03
AI Technical Summary
Existing real-time audio and video content analysis technology cannot achieve real-time content analysis of multiple tracks and multiple scenarios while maintaining low latency. In particular, in the end-side analysis mode, device heating affects the user experience, while MCU centralized analysis has the problems of increased encoding and decoding delay and loss of original track information.
By bypassing the server to grab the original audio and video streams, using the synchronization source identifier to separate the data streams of different tracks, calling the target microservice for processing, and aligning the processing results through cross-modal timestamps, multi-user multi-track parallel processing is achieved.
It realizes the real-time requirements of multi-track audio and video data, supports compound rule query indexing, expands applicable scenarios and applicable needs, improves the accuracy of tracing violators, and reduces device power consumption and latency.
Smart Images

Figure CN120751195A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio and video content recognition, and in particular to an audio and video content recognition method, device, medium and program product. Background Art
[0002] Real-time audio and video content analysis uses artificial intelligence and big data technologies to instantly process live audio and video streams, extracting key information (such as speech-to-text, object recognition, and behavioral analysis) and enabling applications such as content review, quality assessment, and interaction optimization. Its core characteristic is that latency must be controlled to the millisecond level (e.g., below 300ms) to meet real-time requirements. 12 This technology is widely used in scenarios such as video conferencing, live interactive broadcasts, and security monitoring, achieving dynamic adjustments and instant feedback through a streaming data processing architecture and machine learning algorithms.
[0003] Current real-time audio and video content analysis suffers from two major technical flaws. The end-side analysis model relies on the computing power of the terminal device for content quality inspection, which causes high-power consumption devices to heat up and affect the user experience (such as in mobile live broadcast scenarios), and is unable to obtain multi-party interactive content from a global perspective (such as multiple people violating regulations in video conferences). MCU centralized analysis, on the other hand, uses server-decoded mixed streams for analysis, which increases encoding and decoding latency by 200-500ms, making it difficult to meet real-time quality inspection requirements. The original track information is lost after mixing, making it difficult to locate violations (such as being unable to distinguish between violators in multiple video sources). The current industry pain point is that existing technologies cannot achieve real-time content analysis across multiple tracks and scenarios while maintaining the low latency advantage of the SFU. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, medium, and program product for audio and video content recognition, which can detect global defects of bills while accurately detecting fine-grained microscopic defects.
[0005] According to one aspect of the present application, a method for identifying audio and video content is provided, the method comprising:
[0006] The original audio and video streams are captured from the server in a bypassed manner, and the original audio and video streams of different tracks are separated by the synchronization source identifier to determine the data stream to be processed for each track corresponding to each user;
[0007] For each data stream to be processed, the corresponding target microservice is called based on the business requirements for processing the data stream to be processed, and the data stream to be processed is processed based on the target microservice to obtain a processing result;
[0008] The corresponding processing results of the processed data streams of different tracks are aligned across modal timestamps to obtain the audio and video content recognition results.
[0009] According to one aspect of the present application, a device for identifying audio and video content is provided, the device comprising:
[0010] The separation module is used to capture the original audio and video streams from the server in a bypass manner, separate the original audio and video streams of different tracks by using the synchronization source identifier, and determine the data stream to be processed for each track corresponding to each user;
[0011] A processing module is used to call a corresponding target microservice for each data stream to be processed based on the business requirements for processing the data stream to be processed, and process the data stream to be processed based on the target microservice to obtain a processing result;
[0012] The processing result alignment module is used to align the corresponding processing results of the to-be-processed data streams of different tracks across modal timestamps to obtain the audio and video content recognition results.
[0013] According to another aspect of the present application, an electronic device is provided, the electronic device comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by at least one processor. The computer program is executed by at least one processor so that the at least one processor can execute the audio and video content recognition method of any embodiment of the present application.
[0017] According to another aspect of the present application, a computer-readable storage medium is provided, which stores computer instructions, and the computer instructions are used to enable a processor to implement the audio and video content recognition method of any embodiment of the present application when executed.
[0018] According to another aspect of the present application, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the audio and video content recognition method of any embodiment of the present application is implemented.
[0019] The technical solution of the embodiment of the present application is to capture the original audio and video stream from the server in a bypass manner, separate the original audio and video streams of different tracks through the synchronization source identifier, and determine the data stream to be processed of each track corresponding to each user; for each data stream to be processed, the corresponding target microservice is called based on the business needs of processing the data stream to be processed, and the data stream to be processed is processed based on the target microservice to obtain the processing result; the corresponding processing results of the processed data streams of different tracks are aligned across modal timestamps to obtain the audio and video content recognition result. The above solution can realize parallel processing of audio and video data of multiple tracks of multiple users, meet the real-time requirements of real audio and video content recognition, and support query indexing of composite rules, and can support subsequent reuse of audio and video content recognition results of each track, expanding the applicable scenarios and applicable needs.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1 A flowchart of a method for identifying audio and video content provided in an embodiment of the present application;
[0023] Figure 2 A flowchart of a method for identifying audio and video content provided in another embodiment of the present application;
[0024] Figure 3 A flowchart of a method for identifying audio and video content provided in another embodiment of the present application;
[0025] Figure 4 A schematic diagram of the structure of an audio and video content recognition device provided in an embodiment of the present application;
[0026] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0028] It should be noted that the terms "first", "second", "third", "fourth", "actual", "preset", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0029] Figure 1 This is a flowchart of an audio and video content recognition method provided in an embodiment of the present application. The embodiment of the present application is applicable to the case of performing content recognition and analysis on audio and video. The method can be executed by an audio and video content recognition device, which can be implemented in the form of hardware and / or software, and the audio and video content recognition device can be configured in an electronic device. Figure 1 As shown, the method includes:
[0030] S110 , capturing the original audio and video streams from the server in a bypass manner, separating the original audio and video streams of different tracks by using the synchronization source identifier, and determining the data stream to be processed for each track corresponding to each user.
[0031] The server can be a server that transmits the original audio and video streams, such as an SFU server. The synchronization source identifier is a 32-bit random identifier in the RTP protocol header that uniquely identifies a media stream source (such as a camera or microphone). When multiple users send audio and video streams simultaneously, the synchronization source identifier SSRC value for each stream is generated independently and does not conflict with each other. The receiving end can use the SSRC to distinguish data streams from different sources.
[0032] In an embodiment of the present application, the original audio and video stream transmitted from the server can be transmitted through an established protocol. The original audio and video stream is captured from the server in a bypass manner. The server usually performs routing, protocol conversion, or load balancing operations on the audio and video stream, while bypass capture skips these processes and directly obtains the original stream from the server's underlying interface or network transport layer. By bypassing the capture of the original audio and video stream, the transit links can be reduced and the end-to-end transmission delay can be reduced, which is particularly suitable for scenarios with high real-time requirements (such as interactive live broadcasts and video conferencing). Through a distributed processing architecture, efficient multi-channel concurrent processing can be achieved. Directly capturing the original stream can avoid quality loss caused by server transcoding or compression, retain the original audio and video encoding format (such as 4K resolution, high frame rate) and timestamp synchronization information, and ensure the original fidelity of image quality and sound quality.
[0033] In an embodiment of the present application, by bypassing the capture of the original audio and video stream, the original audio and video stream can carry the original timestamp, which is convenient for the subsequent alignment of each track data. The original audio and video stream transmitted by the server may be data of different modes transmitted by different users, for example, it may include audio data of user A and audio and video data of user B. When the original audio and video are transmitted, the protocol header includes a synchronization source identifier. The synchronization source identifier reflects the different users and different modes described in the audio and video. After obtaining the original audio and video data, the original audio and video streams of different tracks can be separated based on the synchronization source identifier to obtain the data to be processed corresponding to different tracks.
[0034] S120 : For each data stream to be processed, call a corresponding target microservice based on the business requirement for processing the data stream to be processed, and process the data stream to be processed based on the target microservice to obtain a processing result.
[0035] Target microservices are functional modules designed for different content recognition needs, including face recognition, object detection, optical character recognition (OCR), speech transcription, voiceprint recognition, and sentiment analysis. Each microservice forms a separate, heterogeneous branch, capable of implementing an independent function and supporting API calls.
[0036] In an embodiment of the present application, the business needs for processing the data stream to be processed can be determined, that is, the functions required for processing the data stream to be processed can be determined, and the target microservice to be called is determined based on the business needs. The target microservice is called, and the data stream to be processed is processed based on the target microservice. For each data stream to be processed, if the cloud service computing power meets the requirements, it can be processed in parallel to improve the real-time processing. In the case that all the data streams to be processed are not processed in parallel, the inference tasks can be allocated to the GPU cluster through a polling algorithm based on the graphics processor GPU computing power load and task priority. The computing power load includes, for example, video memory occupancy, CUDA core utilization, etc. The process of determining the character priority can be, for example, in a live broadcast scenario, where the real-time requirements are high, so the priority of live broadcast violation detection is higher than that of ordinary content analysis.
[0037] In the embodiment of the present application, the macro processes the data stream to be processed of each track, and obtains the processing results corresponding to the data to be processed of each track, such as the processing results corresponding to the audio of user A, the processing results corresponding to the audio of user B, and the processing results corresponding to the image of user B.
[0038] S130 , aligning the corresponding processing results of the to-be-processed data streams of different tracks with cross-modal timestamps to obtain audio and video content recognition results.
[0039] For example, in the process of processing the data streams to be processed in different tracks, the processing is performed separately through heterogeneous branches, and the processing results corresponding to each track exist separately after processing. The processing results can be sorted and combined to achieve effective integration of the processing results.
[0040] In the embodiment of the present application, since the original audio and video stream is obtained by bypass capture, the original audio and video stream carries a timestamp, and the data streams to be sorted out on each track after the original audio and video stream is separated also carry corresponding timestamps. Furthermore, the processing results after processing each data stream to be processed also carry corresponding timestamps. Based on the timestamps carried by each processing result, the processing results obtained after processing the data to be processed on different tracks can be timestamp aligned across modalities, that is, the processing results corresponding to the data to be processed of different users and different modalities of the same user are timestamp aligned to obtain audio and video content recognition results aligned in the time dimension, so as to facilitate subsequent queries on the audio and video recognition results.
[0041] The technical solution of the embodiment of the present application is to capture the original audio and video stream from the server in a bypass manner, separate the original audio and video streams of different tracks through the synchronization source identifier, and determine the data stream to be processed of each track corresponding to each user; for each data stream to be processed, the corresponding target microservice is called based on the business needs of processing the data stream to be processed, and the data stream to be processed is processed based on the target microservice to obtain the processing result; the corresponding processing results of the processed data streams of different tracks are aligned across modal timestamps to obtain the audio and video content recognition result. The above solution can realize parallel processing of audio and video data of multiple tracks of multiple users, meet the real-time requirements of real audio and video content recognition, and support query indexing of composite rules, and can support subsequent reuse of audio and video content recognition results of each track, expanding the applicable scenarios and applicable needs.
[0042] Before processing the data stream to be processed based on the target microservice and obtaining a processing result, the method further includes:
[0043] Predict the target microservices required in the business scenario based on the business scenario;
[0044] The target microservice is preloaded so as to process the data stream to be processed based on the preloaded target microservice in real time during the business scenario.
[0045] For example, in some scenarios, the content to be recognized is known in advance, and the target microservices required to process the audio and video streams in that scenario can be predicted based on the business scenario. Preloading the target microservices allows the data stream to be processed based on the preloaded target microservices while the business scenario is ongoing. This eliminates the time-consuming issue of loading the target microservices during the real-time business scenario, effectively improving the real-time performance of audio and video content recognition.
[0046] Figure 2 This is a flowchart of an audio and video content recognition method provided by another embodiment of the present application. The present embodiment is optimized based on the above embodiment. For solutions not fully described in the present embodiment, please refer to the above embodiment. Figure 2 As shown, the method of the embodiment of the present application specifically includes the following steps:
[0047] S210 , capturing the original audio and video streams from the server in a bypass manner, separating the original audio and video streams of different tracks by using the synchronization source identifier, and determining the data stream to be processed for each track corresponding to each user.
[0048] S220 : For the separated to-be-processed data streams of different tracks, align the to-be-processed data streams with timestamps based on a hardware accelerated decoder.
[0049] For example, since the original audio and video data captured from the server bypass carries a timestamp, the timestamps of the data streams to be processed on different tracks can be aligned based on the timestamps. For the data streams to be processed on different tracks after separation, the data streams to be processed are aligned with the timestamps based on the hardware accelerated decoder. The hardware accelerated decoder refers to a technology that replaces traditional CPU software decoding with dedicated hardware (such as GPU's NVENC / NVDEC, IntelQuick Sync, AMD VCE and other modules) and uses fixed-function circuits or parallel computing units to perform high-speed decoding of video streams (such as H.264 / AV1). It can significantly reduce power consumption (reduce CPU occupancy by more than 50%) and improve throughput (support 8K@60fps real-time decoding), and is widely used in scenarios such as video conferencing and 4K playback. Its core advantage lies in the hardware-level optimized instruction set and low latency characteristics. The real-time performance of data processing can be improved through the decoding and timestamp alignment processing of the hardware accelerated decoder.
[0050] S230 : storing the to-be-processed data streams of different tracks into a distributed cache queue according to different track identifiers and data types; wherein the cache window in the distributed cache queue is a window of preset time.
[0051] Exemplarily, for the data streams to be processed on different tracks, there are corresponding identifiers of different tracks, as well as the data types of the data streams to be processed, such as video or audio. The data streams to be processed on different tracks are stored in a distributed cache queue, so that the data streams to be processed are subsequently extracted based on the distributed cache queue for processing. The data streams to be processed on different tracks can be stored in different cache queues for parallel acquisition and processing. The cache window of the distributed cache queue can be a window of preset time, and the preset time size can be determined according to actual conditions. For example, it can be set to be greater than or equal to the time required to process the data stream to be processed in the cache window. For example, it can be 50ms.
[0052] S240: For each data stream to be processed, call a corresponding target microservice based on the business requirement for processing the data stream to be processed, and process the data stream to be processed based on the target microservice to obtain a processing result.
[0053] S250 . For a time period corresponding to a segment of audio, obtain image frame data corresponding to the time period.
[0054] For example, because audio generally occurs within a time period, audio that is too short or a single frame cannot be effectively recognized. During the time stamp alignment of the processed results obtained after processing the data stream to be processed, the image frame data corresponding to the time period of a segment of effectively recognizable audio can be obtained.
[0055] S260: Align the processing result corresponding to the audio segment, the processing result corresponding to the image frame data, and the corresponding time period to achieve cross-modal timestamp alignment of the processing results.
[0056] For example, the audio segment is recognized, and a processing result corresponding to the audio segment exists. The image frame data within the time period is processed, and a processing result corresponding to the image frame exists. The processing results can be uniformly output in JSONSchema format, including fields such as timestamp (accuracy ±5ms), spatial coordinates (Bounding Box), and confidence level. The processing results corresponding to the audio can be aligned with the processing results corresponding to the image frame data to achieve cross-modal timestamp alignment of the processing results.
[0057] In an embodiment of the present application, for the data streams to be processed of different tracks after separation, the data streams to be processed are aligned with the timestamps based on the hardware accelerated decoder; the data streams to be processed of different tracks are stored in a distributed cache queue according to different track identifiers and data types; wherein the cache window in the distributed cache queue is a window of preset time. The data to be processed is effectively cached by the cache mechanism, which enables fast extraction and processing, and does not occupy too much storage space, thus achieving instant processing and instant elimination. For the time period corresponding to a segment of audio, the image frame data corresponding to the time period is obtained; the processing results corresponding to the audio segment, the processing results corresponding to the image frame data, and the corresponding time period are aligned to achieve cross-modal timestamp alignment of the processing results, thereby achieving timestamp alignment of the processing results of audio and image, which facilitates subsequent rapid indexing of the processing results.
[0058] Figure 3 This is a flowchart of an audio and video content recognition method provided by another embodiment of the present application. The present embodiment is optimized based on the above embodiment. For solutions not fully described in the present embodiment, please refer to the above embodiment. Figure 3 As shown, the method of the embodiment of the present application specifically includes the following steps:
[0059] S310: grab the original audio and video stream from the server in a bypass mode, separate the original audio and video streams of different tracks by using the synchronization source identifier, and determine the data stream to be processed for each track corresponding to each user.
[0060] S320: For each data stream to be processed, call a corresponding target microservice based on the business requirement for processing the data stream to be processed, and process the data stream to be processed based on the target microservice to obtain a processing result.
[0061] S330: Perform cross-modal timestamp alignment on the corresponding processing results of the to-be-processed data streams of different tracks to obtain audio and video content recognition results.
[0062] S340: Build a distributed index library, and store the structured audio and video content recognition results in the distributed index library according to timestamp sharding.
[0063] For example, a distributed index library can be built in advance, or it can be built after the recognition of the first audio segment and the image frame corresponding to the time period is completed. The structured audio and video recognition content is stored in the distributed index library according to the recognition results corresponding to each time period after timestamp segmentation.
[0064] S350: Determine a time period to be queried, and determine a compensation time period according to a time difference between audio transmission and image frame transmission.
[0065] For example, in the subsequent indexing process according to time, the time period to be queried is determined. In the process of audio and image transmission, there may be a "sound before image" phenomenon - that is, the sound signals emitted by the characters (such as conversations, footsteps) are received before the visual images. Therefore, the audio and image frames of the same length retrieved during the retrieval process may not completely correspond when they are specifically presented, and there may be misalignment. The compensation time period can be determined based on the time difference in the transmission process of the audio and image frames. For example, under normal circumstances, audio transmission is 50ms earlier than image frame transmission, and the compensation time period can be set to 50ms.
[0066] S360: Add the compensation time period before the to-be-queried time period, and add the compensation time period after the to-be-queried time period to obtain a target time period.
[0067] Exemplarily, a compensation time period is added before the time period to be queried, and the compensation time period is added after the time period to be queried to obtain a target time period to compensate for the query time period, so that the processing results corresponding to the audio in the target time period to be queried and the processing results corresponding to the image frame data have complete processing results aligned with the timestamps corresponding to the time period to be queried.
[0068] S370: Query the audio and video content recognition results based on the target time period in the distributed index library.
[0069] Exemplarily, the audio and video content recognition results are queried in the distributed index library based on the target time period, that is, the recognition results of the audio corresponding to the target time period and the recognition results of the image frame data corresponding to the target time period are queried respectively, and the audio recognition results and the recognition results of the image frame data corresponding to the target time period are transmitted, so that the receiving end can extract the audio recognition results and the image frame data recognition results that are timestamp-aligned with the query time period from the target time period.
[0070] The solution of the embodiment of the present application is to construct a distributed index library, store structured audio and video content recognition results in the distributed index library according to timestamp sharding; determine the time period to be queried, and determine the compensation time period based on the time difference between audio transmission and image frame transmission; add the compensation time period before the time to be queried, and add the compensation time period after the time to be queried to obtain a target time period; query the audio and video content recognition results in the distributed index library based on the target time period, thereby solving the problem of time difference in audio data and image frame data transmission, and realizing compensation and alignment of absolute time.
[0071] In the embodiment of the present application, after performing cross-modal timestamp alignment on the corresponding processing results of the to-be-processed data streams of different tracks to obtain the audio and video content recognition results, the method further includes:
[0072] According to the user's compound query logic, respectively query the audio and video content recognition results for the audio and video data of each track corresponding to each query condition in the compound query logic;
[0073] The recognition results are combined to obtain a query result corresponding to the compound query logic.
[0074] Exemplarily, in the process of a user querying for identification results, the user's query rule may be a compound query rule. For example, instead of only querying whether the image contains illegal items, or only querying whether sensitive words appear in the audio, the query is that both the image contains illegal items and the audio contains sensitive words. Therefore, for the user's compound query logic, you can first query the audio and video content recognition results for the audio and video data of each track corresponding to each query condition in the compound query logic based on the separate query logics. For example, if the compound query logic is that the image contains illegal items and the audio contains sensitive words, then query whether the recognition results corresponding to the image frame data contain illegal items, and query whether the recognition results corresponding to the audio frame data contain sensitive words. Combine the query results corresponding to the data of different tracks to obtain the query results that conform to the query logic.
[0075] In an embodiment of the present application, the method further includes:
[0076] If the user has a new query logic, then query whether there is a content recognition result corresponding to the new query logic in the audio and video content recognition results;
[0077] If it exists, the content recognition result is returned;
[0078] If it does not exist, the target microservice corresponding to the newly added query logic is called to process the to-be-processed data stream corresponding to the newly added query logic.
[0079] For example, in many scenarios, users may have new query logic after the query. For example, after querying whether the image contains illegal items and the audio contains sensitive words, the user continues to query whether the image contains specific people. In this case, there is no need to identify the original audio and video again. The query result corresponding to the new query logic can be queried in the audio and video content recognition results. If the recognition result corresponding to the new query logic is found in the audio and video recognition results, the queried content recognition result will be returned to the user as the new query logic. If the content recognition result corresponding to the new query logic is not found in the audio and video recognition results, the target microservice corresponding to the new query logic is called to process the to-be-processed data stream corresponding to the new query logic to obtain the content recognition result corresponding to the new query logic.
[0080] The embodiments of this application provide specific implementation solutions for specific application scenarios, which are elaborated in detail below from three levels: system architecture, key modules, and workflow.
[0081] System architecture design:
[0082] The system adopts a three-level distributed architecture, with each layer independently decoupled and data flow realized through message queues:
[0083] 1. Streaming media interception layer:
[0084] SFU bypass capture module: Based on QUIC / RTP protocol parsing technology, it bypasses the SFU server to capture the original audio and video streams, and separates multi-user multi-track data (such as user A's video stream and user B's audio stream) through the SSRC identifier.
[0085] Lossless decoding engine: Uses hardware-accelerated decoder to convert raw streams into timestamp-aligned YUV video frames (resolution adaptive) and PCM audio frames (16kHz sampling rate).
[0086] Frame cache queue: Create a distributed Redis cache queue based on user ID and track type (Video / Audio) to store the original frame data to be analyzed (cache window 50ms).
[0087] 2. Atomic capability reasoning layer:
[0088] Dynamic task scheduler: Based on GPU computing load (video memory usage, CUDA core utilization) and task priority (for example, live broadcast violation detection priority > general content analysis), it allocates inference tasks to the GPU cluster through a polling algorithm.
[0089] Atomic Capability Library: Pre-built multi-domain AI model microservices, including:
[0090] Video: face recognition, object detection, OCR;
[0091] Audio: speech transcription, voiceprint recognition, and sentiment analysis.
[0092] Structured metadata generation: Model output is unified into JSON Schema format, including fields such as timestamp (accuracy ±5ms), spatial coordinates (Bounding Box), and confidence.
[0093] 3. Dynamic logic decision layer:
[0094] (1) Spatiotemporal indexing engine:
[0095] Timeline alignment service: Uses a sliding window compensation algorithm (dynamic window size adjustment, initial value 50ms) to align cross-modal timestamps of AI analysis results for audio and video frames.
[0096] Ultra-fast index storage: A distributed index library is built based on Elasticsearch, which stores structured metadata in shards by timestamp (with a sharding granularity of 1 second). It supports millisecond-level range queries (e.g., "query all face recognition results between 10:00:00.200 and 10:00:00.300").
[0097] (2) Rule Engine:
[0098] DSL syntax definer: allows users to write complex decision logic (example rules: face matching blacklist list && voice transcription contains sensitive words && prohibited items detected);
[0099] Dynamic loading mechanism: When adding new rules, there is no need to re-parse the original stream data, and the existing atomic capability results can be directly reused.
[0100] Technical implementation of key modules:
[0101] 1. Multi-track stream extraction and identification
[0102] SSRC dynamic mapping table: By parsing the SSRC field in the RTP packet header, a mapping relationship between user ID and media stream is established (for example: user A_video stream → SSRC = 0x1234, user A_audio stream → SSRC = 0x5678), solving the problem of multi-track streams without identification under the SFU architecture.
[0103] QUIC stream reassembly technology: In view of the multiplexing characteristics of the QUIC protocol, deep packet analysis (DPI) technology is used to separate the interleaved audio and video data packets and reassemble them into a complete media stream.
[0104] 2. Atomic Capacity Dynamic Scheduling Algorithm
[0105] Priority queue management:
[0106] High-priority tasks: Scenarios with high real-time requirements (such as live broadcast violation detection) are preferentially assigned to low-load GPU instances;
[0107] Batch task merging: Frame merging (e.g., 10-frame merging inference) is performed on non-real-time tasks (such as historical content review) to reduce GPU memory switching overhead.
[0108] Model preheating mechanism: Based on historical data analysis, traffic peaks are predicted and high-frequency atomic capability models are loaded into GPU memory in advance (such as preloading item detection models in e-commerce live broadcast scenarios).
[0109] Timeline indexing and rule triggering:
[0110] Sliding window compensation algorithm:
[0111] Audio-video alignment: Dynamically adjust the offset of video frames based on the audio timestamp (formula: Δt = (T_audio - T_video) × network jitter coefficient);
[0112] Cross-user alignment: Synchronize the time base of each user flow through the NTP server to eliminate clock deviation between devices.
[0113] Compound event judgment logic:
[0114] Time overlap determination: Calculate the time window intersection of multiple atomic capability results (e.g., face recognition results and speech sensitive words appear in the same 500ms window);
[0115] Spatial correlation analysis: Combines the bounding box coordinates of the video frame to determine the spatial relationship between the illegal items and specific people (for example, if the prohibited items appear in the host's hand area).
[0116] Typical workflow (taking live broadcast violation detection as an example):
[0117] participant SFU as SFU server
[0118] Participant interception layer as streaming media interception layer
[0119] Participant reasoning layer as atomic ability reasoning layer
[0120] Participant decision layer as dynamic logic decision layer
[0121] SFU->>Interception layer: forwarding multi-track audio and video streams
[0122] Interception layer->>Interception layer: Separate user A video stream / user B audio stream
[0123] Interception layer ->> Inference layer: Send decoded video frame (timestamp T1)
[0124] Interception layer ->> Inference layer: Send decoded audio frame (timestamp T2)
[0125] Inference layer->>Inference layer: call face recognition (video frame) → output blacklist ID
[0126] Inference layer ->> Inference layer: Call speech transcription (audio frame) → Output sensitive word list Inference layer ->> Decision layer: Write structured metadata to Elasticsearch
[0127] Decision layer ->>Decision layer: Timeline alignment (T1 and T2 deviation compensation)
[0128] Decision layer->>Decision layer: Execution rules: Blacklist ID && Sensitive words
[0129] Decision layer ->> Operation and maintenance platform: trigger violation alarm (marking time interval T1-T2+500ms)
[0130] The beneficial effects of the above scheme are:
[0131] Multi-track processing capability: supports simultaneous analysis of 16-channel 1080p video streams + 32-channel audio streams, with a latency of <80ms (measured data);
[0132] Atomic capability reuse rate: The AI inference results of the same video frame can be reused in more than 5 scenes, increasing GPU utilization by 58%;
[0133] Violation location accuracy: Based on the original track separation technology, the accuracy of tracing the violator is 99.3% (compared to 47% of the MCU solution).
[0134] Advantages of multi-track processing capabilities:
[0135] Causal Chain 1: From Streaming Media Interception Layer Design to Improving Violation Location Accuracy
[0136] A[SFU bypass lossless ripping]-->B{keep original multi-track data}
[0137] B-->C1 (Independent analysis of user A's video stream)
[0138] B-->C2 (Independent analysis of user B's audio stream)
[0139] C1+C2-->D[cross-user multimodal event association]
[0140] D-->E (accuracy rate of tracing the violating entity is 99.3%)
[0141] Derivation of the advantages of computing power reuse efficiency
[0142] Causal Chain 2: From Atomic Capability Decoupling Design to Boosting GPU Utilization
[0143] A[Dynamic scheduling of atomic capabilities] --> B{One-time inference for multiple scenarios}
[0144] B-->C1 (face recognition results of the same video frame)
[0145] C1-->D1 (for live broadcast real-name authentication)
[0146] C1-->D2 (for blacklist comparison)
[0147] C1-->D3 (for audience attention analysis)
[0148] D1+D2+D3-->E (reduce repeated reasoning by 3 times)
[0149] E-->F (GPU utilization increased by 58%)
[0150] Experimental verification:
[0151] In e-commerce live streaming scenarios, a single 1080p video stream must simultaneously run three atomic capabilities: face recognition, object detection, and OCR.
[0152] Traditional solution: Three capabilities are independently inferred, and the GPU memory occupies 9.2GB;
[0153] The present invention reuses YUV data of the same video frame, and reduces the video memory usage to 3.8 GB (reduced by 58.7%).
[0154] Real-time advantage derivation:
[0155] Causal Chain 3: From Timeline Indexing Engine to End-to-End Latency Compression
[0156] A1 [Bypass capture to avoid MCU encoding and decoding] --> B1 (reduce 300ms delay)
[0157] A2[Hardware accelerated decoding]-->B2(decoding delay <5ms)
[0158] A3 [sliding window time alignment] --> B3 (cross-modal error < 20ms)
[0159] B1+B2+B3-->C [end-to-end analysis latency <80ms]
[0160] Scalability advantage derivation:
[0161] Causal Chain 4: Decoupling the Rule Engine from Atomic Capabilities to Improving Scenario Expansion Efficiency
[0162] graph LR
[0163] A[rule engine is independent of video stream] --> B{no need to retrain the model for new scenarios}
[0164] B-->C1 (define DSL rule file)
[0165] C1-->D1 (load existing atomic capability results)
[0166] D1-->E (scenario launch cycle shortened from 7 days to 2 hours)
[0167] Figure 4 This is a structural diagram of an audio and video content recognition device provided in an embodiment of the present application. The device can execute the audio and video content recognition method provided in any embodiment of the present application and has the corresponding functional modules and beneficial effects of the execution method. Figure 4 As shown, the device includes:
[0168] Separation module 410, used to capture the original audio and video streams from the server in a bypass mode, separate the original audio and video streams of different tracks by using the synchronization source identifier, and determine the data stream to be processed for each track corresponding to each user;
[0169] The processing module 420 is configured to call a corresponding target microservice for each data stream to be processed based on the business requirement for processing the data stream to be processed, and process the data stream to be processed based on the target microservice to obtain a processing result;
[0170] The processing result alignment module 430 is used to align the corresponding processing results of the to-be-processed data streams of different tracks with respect to cross-modal timestamps to obtain audio and video content recognition results.
[0171] In the embodiment of the present application, the separation module 410 determines each data stream to be processed corresponding to each user, including:
[0172] For the separated data streams of different tracks, the data streams to be processed are aligned with the timestamps based on the hardware accelerated decoder;
[0173] According to different track identifiers and data types, the to-be-processed data streams of different tracks are stored in a distributed cache queue; wherein the cache window in the distributed cache queue is a window of preset time.
[0174] In the embodiment of the present application, the processing result alignment module 430 performs cross-modal timestamp alignment on the corresponding processing results of the to-be-processed data streams of different tracks after processing, including:
[0175] For a time period corresponding to an audio segment, obtain the image frame data corresponding to the time period;
[0176] The processing results corresponding to the audio segment, the processing results corresponding to the image frame data, and the corresponding time periods are aligned to achieve cross-modal timestamp alignment of the processing results.
[0177] In the embodiment of the present application, after performing cross-modal timestamp alignment on the corresponding processing results of the to-be-processed data streams of different tracks and obtaining the audio and video content recognition results, the apparatus further includes:
[0178] A storage module is used to build a distributed index library and store the structured audio and video content recognition results in the distributed index library according to timestamp sharding;
[0179] During the query process of the audio and video content recognition result, the device further includes:
[0180] A compensation time period determination module is used to determine the time period to be queried and determine the compensation time period based on the time difference between audio transmission and image frame transmission;
[0181] a target time period determination module, configured to add the compensation time period before the time to be queried, and add the compensation time period after the time to be queried, to obtain a target time period;
[0182] The query module is used to query the audio and video content recognition results based on the target time period in the distributed index library.
[0183] In an embodiment of the present application, before processing the data stream to be processed based on the target microservice and obtaining the processing result, the device further includes:
[0184] A target microservice prediction module is used to predict the target microservice required in the business scenario based on the business scenario;
[0185] A preloading module is used to preload the target microservice so as to process the data stream to be processed based on the preloaded target microservice in real time in the business scenario.
[0186] In the embodiment of the present application, after performing cross-modal timestamp alignment on the corresponding processing results of the to-be-processed data streams of different tracks and obtaining the audio and video content recognition results, the apparatus further includes:
[0187] A query module, configured to query, based on a user's compound query logic, from the audio and video content recognition results, the recognition results corresponding to the audio and video data of each track corresponding to each query condition in the compound query logic;
[0188] The combination module is used to combine the recognition results to obtain the query result corresponding to the compound query logic.
[0189] In an embodiment of the present application, the device further includes:
[0190] A new query module is added, which is used to query whether there is a content recognition result corresponding to the new query logic in the audio and video content recognition results if the user has added a new query logic;
[0191] A return module, configured to return the content recognition result if it exists;
[0192] A new processing module is added, which is used to call the target microservice corresponding to the newly added query logic to process the to-be-processed data stream corresponding to the newly added query logic if it does not exist.
[0193] An audio and video content recognition device provided in an embodiment of the present application can execute an audio and video content recognition method provided in any embodiment of the present application, and has functional modules and beneficial effects corresponding to the execution method.
[0194] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0195] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0196] Multiple components in electronic device 10 are connected to I / O interface 15, including an input unit 16, such as a keyboard and mouse; an output unit 17, such as various types of displays and speakers; a storage unit 18, such as a magnetic disk and optical disk; and a communication unit 19, such as a network card, a modem, a wireless audio and video content recognition transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0197] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any other suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the audio and video content recognition method.
[0198] In some embodiments, the audio and video content recognition method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the audio and video content recognition method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the audio and video content recognition method in any other appropriate manner (for example, by means of firmware).
[0199] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0200] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable audio and video content recognition device, so that when executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0201] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0202] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0203] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0204] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0205] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the audio and video content recognition method provided in any embodiment of the present application.
[0206] In the process of implementation, the computer program product can be written in one or more programming languages or a combination thereof to write computer program code for performing the operations of the present invention, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, through the Internet using an Internet service provider).
[0207] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired information of the technical solution of this application can be achieved. This document is not limited here.
[0208] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A method for identifying audio and video content, characterized in that: The method comprises: The original audio and video streams are captured from the server in a bypassed manner, and the original audio and video streams of different tracks are separated by the synchronization source identifier to determine the data stream to be processed for each track corresponding to each user; For each data stream to be processed, the corresponding target microservice is called based on the business requirements for processing the data stream to be processed, and the data stream to be processed is processed based on the target microservice to obtain a processing result; The corresponding processing results of the processed data streams of different tracks are aligned across modal timestamps to obtain the audio and video content recognition results.
2. The method according to claim 1, characterized in that Determine the data streams to be processed for each user, including: For the separated data streams of different tracks, the data streams to be processed are aligned with the timestamps based on the hardware accelerated decoder; According to different track identifiers and data types, the to-be-processed data streams of different tracks are stored in a distributed cache queue; wherein the cache window in the distributed cache queue is a window of preset time.
3. The method according to claim 1, characterized in that Align the cross-modal timestamps of the corresponding processing results of the processed data streams of different tracks, including: For a time period corresponding to an audio segment, obtain the image frame data corresponding to the time period; The processing results corresponding to the audio segment, the processing results corresponding to the image frame data, and the corresponding time periods are aligned to achieve cross-modal timestamp alignment of the processing results.
4. The method according to any one of claims 1 to 3, characterized in that After performing cross-modal timestamp alignment on the corresponding processing results of the to-be-processed data streams of different tracks to obtain the audio and video content recognition results, the method further includes: Constructing a distributed index library and storing the structured audio and video content recognition results in the distributed index library according to timestamp sharding; During the query process of the audio and video content recognition result, the method further includes: Determine the time period to be queried, and determine the compensation time period based on the time difference between audio transmission and image frame transmission; Adding the compensation time period before the time to be queried, and adding the compensation time period after the time to be queried, to obtain a target time period; The audio and video content recognition result is queried in the distributed index library based on the target time period.
5. The method according to claim 1, wherein Before processing the data stream to be processed based on the target microservice and obtaining a processing result, the method further includes: Predict the target microservices required in the business scenario based on the business scenario; The target microservice is preloaded so as to process the data stream to be processed based on the preloaded target microservice in real time during the business scenario.
6. The method according to claim 1, wherein After performing cross-modal timestamp alignment on the corresponding processing results of the to-be-processed data streams of different tracks to obtain the audio and video content recognition results, the method further includes: According to the user's compound query logic, respectively query the audio and video content recognition results for the audio and video data of each track corresponding to each query condition in the compound query logic; The recognition results are combined to obtain a query result corresponding to the compound query logic.
7. The method according to claim 6, characterized in that The method further comprises: If the user has a new query logic, then query whether there is a content recognition result corresponding to the new query logic in the audio and video content recognition results; If it exists, the content recognition result is sent back; If it does not exist, the target microservice corresponding to the newly added query logic is called to process the to-be-processed data stream corresponding to the newly added query logic.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the audio and video content recognition method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the audio and video content recognition method according to any one of claims 1 to 7 when executed.
10. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the audio and video content recognition method according to any one of claims 1 to 7.