A live broadcast goods carrying real-time detection method and system based on multi-modal fusion

By structuring live data streams into atomic event sequences and constructing dynamic semantic event graphs, combined with a dual-channel detection mechanism, the problem of missing cross-modal event associations is solved, enabling deep integration of live content and complex violation identification, thus improving the accuracy and flexibility of the detection system.

CN121330409BActive Publication Date: 2026-04-14浙江省市场监管发展研究中心(浙江省平台经济监测中心浙江省广告监测中心)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
浙江省市场监管发展研究中心(浙江省平台经济监测中心浙江省广告监测中心)
Filing Date
2025-12-12
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing live streaming real-time detection technologies cannot effectively capture and express the inherent logical connections and temporal context relationships between cross-modal events when performing multimodal fusion, resulting in limited accuracy and depth in identifying complex violations.

Method used

Multimodal live streaming data streams are uniformly abstracted into atomic event sequences containing timestamps, modal sources, subjects, predicates, and objects. A dynamic semantic event graph is constructed, and a dual-channel detection mechanism is used to perform subgraph matching of the violation pattern library and structural anomaly detection of graph structure entropy change rate, thereby achieving deep fusion and contextual association of cross-modal information.

Benefits of technology

It enhances the coverage and flexibility of real-time detection in live-streaming e-commerce and the ability to respond to unexpected risks, provides intuitive and traceable detection results, reduces the false positive rate, and improves the overall availability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330409B_ABST
    Figure CN121330409B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of multi-modal data processing, and discloses a live broadcast goods carrying real-time detection method based on multi-modal fusion, which comprises the following steps: S1, collecting a multi-modal live broadcast data stream in real time; S2, based on the atomic event sequence, incrementally constructing a dynamic semantic event graph; S3, on the dynamic semantic event graph; S4, according to the results of subgraph matching detection or structural anomaly detection. By uniformly abstracting multi-modal data streams such as video, audio and text into structured atomic events and incrementally constructing a dynamic semantic event graph capable of expressing the time, entity and logical association between events, the discrete and heterogeneous live broadcast information is deeply fused and contextually associated, the limitation that traditional methods can only perform shallow feature splicing is overcome, the internal connection and combination mode between different modal events can be understood from a global perspective, and a unified and semantically rich data basis is provided for subsequent complex behavior recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal data processing technology, specifically to a real-time detection method and system for live-streaming e-commerce based on multimodal fusion. Background Technology

[0002] With the development of internet technology, live-streaming e-commerce, as an emerging e-commerce model, has been widely adopted. To ensure the compliance of live-streaming content and maintain a healthy platform ecosystem, real-time content detection technologies for the live-streaming process have emerged. These technologies typically need to process multimodal data generated during the live stream, including video streams, audio streams, and user comment text streams, to identify any potential violations or abnormal behaviors.

[0003] In existing live streaming real-time detection technologies, a common approach is to employ a multimodal fusion model. The basic process is as follows: First, for different data modalities such as video, audio, and text, independent deep learning models are used to extract features, transforming the raw data of each modality into high-dimensional feature vectors. Then, these feature vectors extracted from different modalities are concatenated or weighted and fused using an attention mechanism to form a unified fused feature vector. Finally, this fused feature vector is input into a downstream classifier, which determines the overall state of the live streaming content.

[0004] However, existing technologies have inherent limitations in multimodal fusion: the aforementioned fusion methods based on direct splicing or weighting of features, while incorporating multimodal information to some extent, still essentially treat information from different sources as fragmented feature data. This makes it difficult to effectively capture and express the inherent logical connections and temporal context relationships between cross-modal events. For example, the system cannot structurally semantically associate a broadcaster's verbal description of a product in audio with their specific actions of displaying the product in video footage. Therefore, its accuracy and depth are limited when identifying complex combinations of behaviors that require a comprehensive understanding of multimodal information. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a real-time detection method and system for live-streaming e-commerce based on multimodal fusion, which solves the problems of missing cross-modal semantic associations and difficulty in identifying complex violations.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a real-time detection method for live-streaming e-commerce based on multimodal fusion, comprising the following steps:

[0007] S1. Unified Abstraction and Structure: The heterogeneous live streaming data streams, including video streams, audio streams, and text streams, which are collected in real time, are uniformly abstracted and structured into a series of atomic event sequences containing timestamps, modal sources, subjects, predicates, and objects. This step aims to transform the originally discrete and unstructured multimodal information into a unified and semantically rich data foundation, laying the foundation for subsequent deep fusion and analysis.

[0008] S2. Dynamic Semantic Event Graph Construction: Based on the atomic event sequence, a dynamic semantic event graph is incrementally constructed and dynamically maintained. This graph uses each atomic event constituting the atomic event sequence as a node and establishes pre-defined inter-event relationships as edges. This step enables deep fusion and contextual association of discrete, heterogeneous live broadcast information, allowing for a global perspective on the inherent connections and combination patterns between different modal events, significantly improving the depth of information fusion.

[0009] S3. Dual-channel collaborative detection: On the dynamic semantic event graph, two complementary detection mechanisms are executed in parallel: one is subgraph matching detection based on a violation pattern library, used to accurately identify known, well-defined violation behavior patterns; the other is structural anomaly detection based on graph structure entropy change rate, used to keenly perceive unknown, non-pre-defined, group collaborative anomaly behaviors. This dual-channel detection mechanism achieves collaborative discovery of known violations and unknown anomalies, greatly improving the coverage of the detection system and its flexibility in responding to sudden risks.

[0010] S4. Intelligent Decision Making and Evidence Solidification: Based on the detection results of either the subgraph matching detection or the structural anomaly detection mechanism, a real-time detection alarm for the live content is generated. At the same time, the present invention can also solidify an interpretable chain of evidence based on the detection source that triggers the alarm, providing intuitive and traceable proof for the detection results, and realizing the transparency of the detection process and the credible verification of the results.

[0011] Preferably, the step of structuring the real-time acquired multimodal live data stream into a sequence of atomic events specifically includes:

[0012] From the video stream in the multimodal live data stream, video atomic events are extracted and generated using object detection, behavior recognition, and optical character recognition technologies;

[0013] Audio atomic events are extracted and generated from the audio stream in the multimodal live data stream using automatic speech recognition technology;

[0014] From the text stream in the multimodal live data stream, natural language processing technology is used to extract and generate text atomic events;

[0015] The video atomic events, audio atomic events, and text atomic events together constitute an atomic event sequence.

[0016] Preferably, the subgraph matching detection based on the violation pattern library includes: matching the local structure of the dynamic semantic event graph with a predefined violation subgraph template in the violation pattern library.

[0017] Preferably, the structural anomaly detection based on the graph structure entropy change rate specifically includes:

[0018] Calculate the rate of change of the graph structure entropy of the dynamic semantic event graph over time;

[0019] A structural anomaly is determined to have occurred when the rate of change exceeds a dynamically updated baseline threshold.

[0020] Preferred options also include:

[0021] When generating the real-time detection alarm, a chain of evidence is established;

[0022] If triggered by the subgraph matching detection, then the evidence chain is a matched illegal subgraph structure;

[0023] If triggered by the structural anomaly detection, the evidence chain is a local graph structure within the time window that causes a dramatic change in graph structure entropy.

[0024] Preferably, in the step of structuring the real-time acquired multimodal live data stream into a sequence of atomic events, the video stream processing is implemented using a parallel pipeline architecture, specifically including:

[0025] The frame distribution and preprocessing step is used to capture image frames at a preset frame rate and perform standardized preprocessing; the parallel processing step synchronously sends the preprocessed image frames into the target detection pipeline, behavior recognition pipeline and optical character recognition pipeline for parallel processing.

[0026] The event generation and aggregation step is used to combine the output results of each pipeline with timestamps and uniformly format them into video atomic events.

[0027] Preferably, the process of structuring into a sequence of atomic events also includes entity linking and normalization steps:

[0028] The system maintains a dynamically updated entity knowledge base;

[0029] When a new atomic event is generated, the subject and / or object information therein is matched with the entity knowledge base and mapped to a unified entity identity identifier.

[0030] Preferably, the dynamically updated baseline threshold is calculated using an exponential moving average algorithm based on historical graph structure entropy change rate data.

[0031] A real-time detection system for live-streaming e-commerce based on multimodal fusion includes:

[0032] The event abstraction module is used to structure real-time acquired multimodal live data streams into atomic event sequences containing timestamps, modal sources, subjects, predicates, and objects;

[0033] The graph construction module is used to incrementally construct a dynamic semantic event graph based on the atomic event sequence, wherein each atomic event constituting the atomic event sequence serves as a node of the graph, and the preset relationships between events constitute the edges of the graph.

[0034] The detection engine is used to perform subgraph matching detection based on a violation pattern library and structural anomaly detection based on graph structure entropy change rate on the dynamic semantic event graph.

[0035] The decision alarm module is used to generate real-time detection alarms for live content based on the results of the subgraph matching detection or the structural anomaly detection.

[0036] This invention provides a real-time detection method and system for live-streaming e-commerce based on multimodal fusion. It has the following beneficial effects:

[0037] 1. This invention abstracts multimodal data streams such as video, audio, and text into structured atomic events and incrementally constructs a dynamic semantic event graph that can express the temporal, entity, and logical relationships between events. This enables deep fusion and contextual association of discrete and heterogeneous live information, overcoming the limitations of traditional methods that can only perform shallow feature splicing. It can gain insight into the intrinsic connections and combination patterns between different modal events from a global perspective, providing a unified and semantically rich data foundation for subsequent complex behavior recognition.

[0038] 2. This invention achieves the collaborative discovery of known violations and unknown anomalies by executing subgraph matching detection based on a violation pattern library and structural anomaly detection based on graph structure entropy change rate in parallel on a dynamic semantic event graph. The former uses expert knowledge to accurately identify violations with clear characteristics, while the latter does not rely on prior rules and perceives collective collaborative abnormal behavior by monitoring macroscopic mutations in the graph structure. The two complement each other, improving the coverage of the detection system and the flexibility to respond to sudden risks.

[0039] 3. This invention, while generating real-time detection alarms, solidifies an interpretable chain of evidence based on the detection source that triggered the alarm. Whether it is a matched violation subgraph structure or a local graph structure within a time window that causes a dramatic change in graph structure entropy, it can provide intuitive and traceable proof for the detection results, realize the transparency of the detection process and the credible verification of the results, facilitate manual review and subsequent processing, reduce the false judgment rate and improve the overall availability of the system. Attached Figure Description

[0040] Figure 1 This is a flowchart of the present invention;

[0041] Figure 2 This is a system architecture diagram of the present invention. Detailed Implementation

[0042] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0043] Example 1:

[0044] Please see the appendix Figure 1 This invention provides a real-time detection method for live-streaming e-commerce based on multimodal fusion. The core of this method is to perform deep analysis and structuring of unstructured live-streaming data streams, construct a dynamically evolving semantic graph, and achieve comprehensive monitoring of known violations and unknown anomalies through two complementary detection mechanisms.

[0045] In one embodiment, the main steps are as follows: First, step S1 is executed to structure the real-time acquired multimodal live data stream to form an atomic event sequence. Next, step S2 is executed to incrementally construct a dynamic semantic event graph based on the atomic event sequence. Then, step S3 is executed to perform subgraph matching detection and structural anomaly detection in parallel on the graph. Finally, step S4 is executed to generate an alarm based on the detection results.

[0046] Step S1: Structured abstraction of atomic events

[0047] This step aims to unify live streaming data from different sources and in various formats into discrete, structured atomic events, each atomic event... Each is defined as a standardized tuple:

[0048]

[0049] in:

[0050] For timestamps, As a modal source, As the main body of the event, For event predicates, As the object of the event, This is a collection of additional attributes.

[0051] In practice, the system synchronously acquires video, audio, and text streams through a data acquisition interface and uses a unified clock source to mark the timestamps. Subsequently, each data stream is sent to an independent parsing engine.

[0052] The video parsing engine runs object detection, behavior recognition, and optical character recognition models in parallel to extract and generate video atomic events from image frames, such as displaying objects, performing gestures, or displaying text on the screen.

[0053] To achieve efficient parallel processing of video streams, the video parsing engine in this embodiment of the invention employs a task-distribution-based pipeline architecture. Specifically, when the video stream enters the parsing engine:

[0054] Frame distribution and preprocessing: A frame distributor captures image frames from the video stream at a preset frame rate. Each frame is assigned a timestamp synchronized with the master clock and undergoes standardized preprocessing such as size normalization and color space conversion.

[0055] Parallel processing pipeline: The preprocessed image frames are simultaneously fed into three independent processing pipelines:

[0056] Object Detection Pipeline: This pipeline runs a lightweight and efficient object detection model specifically designed for real-time identification of key objects in live streams, such as products, brand logos, and QR codes. Its output is a series of bounding boxes with category labels and confidence scores.

[0057] Behavior recognition pipeline: Since behavior recognition requires timing information, this pipeline has a short-time frame buffer. When the buffer is full or a specific trigger condition is met, the pipeline runs a timing action recognition model to identify the anchor's specific behaviors, such as displaying products, pointing with a finger, clapping, etc.

[0058] Optical Character Recognition Pipeline: This pipeline focuses on recognizing static or dynamic text appearing in the image, such as price tags, slogans, and promotional information. It uses a two-stage OCR model to first locate the text area and then recognize the characters within it.

[0059] Event Generation and Aggregation: After the three pipelines have completed their processing, their output targets, behaviors, and texts are sent to an event generator. This generator, based on the output results of each pipeline and combined with the timestamp, formats them into video atomic events. For example, if the target detection pipeline detects product A, the event generator will generate a video atomic event with product A as the object and occurrence as the predicate. Finally, all generated video atomic events are sent to a unified queue, waiting for subsequent graph construction.

[0060] In this way, parallel operation is concretized into three independent but synchronously started processing pipelines, each responsible for extracting video information from different dimensions, thereby achieving comprehensive and efficient real-time analysis of video content.

[0061] The audio parsing engine uses automatic speech recognition technology to convert speech into text, and then uses a natural language processing unit to extract key semantic phrases to generate audio atomic events that speak specific content.

[0062] The process of extracting key semantic phrases by the natural language processing unit is not a simple keyword matching, but rather a deep semantic analysis method based on a combination of named entity recognition and relation extraction. The specific steps are as follows:

[0063] Step 1, Speech to Text: First, automatic speech recognition technology converts a continuous audio stream into a text string in real time.

[0064] Step 2, Named Entity Recognition: Next, the text string is fed into a pre-trained language model fine-tuned with data from the live-streaming e-commerce domain. This model can accurately identify various predefined entity types in the text, such as:

[0065] Product name, such as: This face mask;

[0066] Brands, such as: XX brand;

[0067] Efficacy / promises, such as: ten times compensation for fakes, immediate whitening effect;

[0068] Price / offer, such as: only 99, buy one get one free;

[0069] Interactive commands, such as: type 1 if you want a baby;

[0070] Step 3, Relation Extraction: After identifying entities, a relation extraction model analyzes the syntactic and semantic relationships between these entities to form structured subject-verb-object triples. For example, for the text "I recommend this XX face mask," the relation extraction model will extract relation triples such as "host," "recommend," and "XX face mask." This triple directly corresponds to the subject S, predicate P, and object O in the atomic event.

[0071] Step 4: Atomic Event Generation: Finally, the system generates one or more audio atomic events based on the extracted relation triples. For example, the above example will generate an audio atomic event with the subject being the broadcaster, the predicate being "recommendation", and the object being "XX face mask", and attach the corresponding timestamp and modality source.

[0072] The text parsing engine directly processes user comments, abstracting each comment into a text atomic event that publishes the comment, and can attach the results of sentiment analysis and sensitive word recognition as its attributes.

[0073] To ensure semantic consistency among nodes in the graph, this system includes an entity linking and normalization sub-step in the process of structured atomic events. This step aims to associate and unify the same entity with different modalities, time points, and representations into a single identity identifier. The specific implementation is as follows:

[0074] The system maintains a dynamically updated entity knowledge base, which contains unique IDs and attribute information of core entities such as products, brands, livestreamers, and active users. When a new atomic event is generated, its subject (S) and object (O) do not directly use the raw text or tags, but are processed by this module:

[0075] For product / brand entities: The system will comprehensively utilize the brand logo recognized by OCR, the product slang / abbreviations recognized by ASR, and a pre-built product information database. Through multi-source information comparison and fuzzy matching algorithms, the objects in the event will be mapped to unique product IDs in the knowledge base.

[0076] For user entities: the system uses the user's unique identifier as its unified identity in the graph, while variable information such as the user's nickname is stored as the attributes of the user node.

[0077] In this way, even if events of different modalities and at different times describe the same entity in different ways, they will eventually be associated with the same entity node in the graph, thus laying a solid foundation for establishing accurate cross-modal associations.

[0078] All generated atomic events are eventually sorted by timestamp to form a unified sequence of atomic events, which serves as input for subsequent steps.

[0079] Step S2: Incremental Construction of Dynamic Semantic Event Graph

[0080] This step transforms a one-dimensional sequence of atomic events into a graph structure rich in relational information. The system maintains a dynamic semantic event graph in memory, constrained by a sliding time window. Whenever a new atomic event E i Once generated, the system creates a corresponding node v in the graph. iAnd start the edge generation process, and determine v i With existing node v in the window j Does a pre-defined relationship exist between events?

[0081] The pre-defined relationships between events aim to reveal the connections between events in different dimensions.

[0082] Temporal continuity relationship: This relationship is used to capture consecutive actions of the same subject. It is established when the subjects of the two events are the same and the difference in their timestamps is less than a preset threshold. .

[0083] In a preferred embodiment of the invention, the time proximity threshold Set between 1 and 5 seconds, with a typical empirical value of 3 seconds, this range can effectively aggregate the host's continuous verbal announcements or users' continuous comments in a short period of time into a cluster of events with a temporal continuity, while avoiding the misconnection of different topics or behaviors with long time spans.

[0084] Subject-object relationship: This relationship is the core of achieving cross-modal fusion. The system uses a pre-trained cross-modal embedding model. This model can map object information from different sources, such as textual descriptions and image features, to the same high-dimensional semantic vector space. It calculates the cosine similarity between the embedding vectors of two event objects; if the similarity exceeds a preset threshold... Then, a subject-object relationship is established between these two event nodes.

[0085] In a preferred embodiment of the present invention, the cross-modal semantic similarity threshold Set within the range of 0.8 to 0.95, for example, 0.85. Setting a higher value can effectively filter out accidental and meaningless cross-modal co-occurrences, improving the accuracy of semantic associations in graph structures.

[0086] Logical relationships: Logical relationships are a higher-level generalization used to describe deeper logical dependencies between events, beyond temporal sequence and entity co-occurrence.

[0087] In one specific embodiment of the present invention, the logical association is implemented through a configurable rule engine, which maintains a set of logical rules, each rule defining a specific event combination pattern.

[0088] A logical rule for identifying question-and-answer interactions can be defined as follows: when an audio event of the nature of a question asked by a host appears in the graph, if one or more text events of the nature of a question asked by a user appear within a specified time window, then an edge is established to establish a logical association between the question event and these answer events. In this way, the system can understand and encode the business interaction logic in the live broadcast.

[0089] Step S3: Subgraph matching detection and structural anomaly detection

[0090] This step is the core of detection and decision-making, where two independent detection tasks are executed in parallel on a dynamic semantic event graph.

[0091] Subgraph matching detection based on a violation pattern library: The goal of this task is to accurately find known violations. The system pre-builds a violation pattern library, where each violation pattern is defined as a violation subgraph template. The template specifies in detail the set of event node types, relationship edge types between nodes, and attribute constraints that nodes must satisfy to constitute a violation. The violation pattern library is not a static library, but a knowledge base that can be dynamically maintained through the accompanying pattern definition and management tools. In one embodiment of the invention, the tool provides a visual graphical interface that allows security experts or operations personnel to draw and define a new violation subgraph template by dragging and dropping different types of event nodes and relationship edges on the canvas. Each template is compiled into a structured description file when saved, which defines in detail the node types, attribute constraints, and edge types and directions. This file can be dynamically loaded by the detection engine, thereby enabling the ability to hot-update violation detection rules without restarting the service, greatly improving the system's flexibility and timeliness in responding to new types of violations.

[0092] For example, a template can be defined as consisting of an audio event node containing an absolute commitment keyword and a video event node displaying a product, and the two must be connected by a subject-object relationship. During detection, a subgraph matching unit continuously executes an optimized subgraph isomorphic matching algorithm on the dynamic semantic event graph. Once a subgraph instance that matches any template in the library is found, it is determined to be a known violation.

[0093] Specifically, the optimization is manifested in the following ways:

[0094] Index-based candidate node filtering: Before matching, the system does not perform a brute-force search of the entire graph. Instead, it first uses a pre-built index to quickly filter out candidate nodes in the graph that may become matching instances based on the type and attributes of the nodes in the violation subgraph template. This greatly reduces the search space.

[0095] Quantitative matching: Since the graph is constructed incrementally, the matching algorithm is also executed incrementally. Whenever a new node or edge is added to the graph, the algorithm only checks whether a new illegal subgraph matching has been formed in the affected local area, instead of matching the entire graph from the beginning every time.

[0096] Employing an efficient matching algorithm kernel: Within the selected candidate regions, the system uses modern subgraph isomorphism algorithms optimized for large-scale and dynamic graphs, such as variants of VF2++, GraphQL, or CECI, to perform the final structure matching. These algorithms, through improved search strategies and pruning techniques, far outperform traditional Ullmann or VF2 algorithms in terms of performance.

[0097] Structural anomaly detection based on graph structure entropy change rate: This task does not rely on prior knowledge and aims to detect sudden, structural anomalies caused by group collaborative behavior.

[0098] The graph structure entropy used in this detection is a general technical term. It is an indicator used to quantify the overall structural complexity of a graph, and it can be calculated in various ways.

[0099] In a preferred embodiment of the present invention, the graph structure entropy is calculated by calculating the Shannon entropy based on the node degree distribution, and the calculation formula is as follows:

[0100]

[0101] in, To illustrate at time The set of node degrees, The degree is The percentage of nodes in the entire graph.

[0102] The detection process does not use the absolute value of entropy, but rather monitors its rate of change over time. A graph entropy analysis unit periodically calculates... The system calculates the rate of change relative to the previous period, and uses an exponential moving average algorithm to dynamically maintain a baseline threshold representing the normal fluctuation level.

[0103] Specifically, the calculation process for the rate of change of graph structure entropy is as follows:

[0104] The system calculates the structural entropy of the graph once at a fixed time period, assuming that at time t... The calculated graph structure entropy is At the time of the previous calculation cycle The entropy obtained is .

[0105] The rate of change relative to the previous period, i.e., the relative rate of change of entropy. It can be calculated using the following formula:

[0106]

[0107] in:

[0108] if If the value is 0, the rate of change can be recorded as infinity or a preset maximum value.

[0109] The rate of change It is a dimensionless relative value that represents the percentage increase or decrease in the complexity of the current periodic chart relative to the previous period.

[0110] When the calculated instantaneous rate of change exceeds this dynamic baseline, the system determines that a structural anomaly has occurred. This anomaly usually corresponds to collaborative behaviors such as malicious screen-scraping that cause a rapid increase in the local connectivity relationships in the graph within a short period of time.

[0111] Step S4: Alarm Generation and Evidence Consolidation

[0112] When any detection task in step S3 is triggered, the system generates a corresponding real-time detection alarm. At the same time, in order to ensure the traceability and interpretability of the detection results, the system will immediately solidify a chain of evidence.

[0113] The solidification operation specifically refers to the following: When an alarm is triggered, an evidence solidification unit serializes the relevant local graph structure and the detailed attributes of all its nodes and edges into an independent, self-contained graph data object. This object completely snapshots all relevant events and their relationships at the moment the alarm is triggered. Subsequently, this graph data object is stored in a dedicated evidence database and obtains a unique evidence ID.

[0114] The generated real-time detection alarm information will contain this evidence ID. When a human reviewer sees this alarm in the alarm center, they can use this ID to call an evidence visualization module. This module can parse the stored graph data object and re-render the local graph structure that triggered the alarm in a graphical way on the front-end interface. Reviewers can intuitively see which events and how they were combined to trigger the alarm, and can click on any node or edge in the graph to view its detailed original information, thereby achieving intuitive and reliable alarm review and handling, and effectively reducing the false judgment rate.

[0115] Example 2:

[0116] Please see the appendix Figure 2This invention provides a real-time detection system for live-streaming e-commerce based on multimodal fusion, which serves as a hardware or software functional entity to implement the detection method described in Embodiment 1. Structurally, the system includes an event abstraction module, a graph construction module, a detection engine, and a decision-making alarm module.

[0117] Event abstraction module:

[0118] The event abstraction module is used to structure real-time acquired multimodal live data streams into atomic event sequences containing timestamps, modality sources, subjects, predicates, and objects. To perform this function, the event abstraction module can include independent parsing units for different modalities.

[0119] A video parsing unit integrates object detection, behavior recognition, and optical character recognition models to extract and generate atomic events from the video stream.

[0120] An audio parsing unit, comprising an automatic speech recognition engine and a natural language processing unit, is used to extract and generate atomic events from the audio stream.

[0121] It also includes a text parsing unit containing models for sensitive word recognition and sentiment classification, used to extract and generate atomic events from the text stream.

[0122] The event abstraction module also includes an event aggregation unit, which sorts the atomic events generated by each parsing unit by timestamp to form a unified, time-sequential atomic event sequence and outputs it to the graph construction module.

[0123] Graph building module:

[0124] The graph building module receives the atomic event sequence output from the event abstraction module and incrementally builds a dynamic semantic event graph based on this sequence. The graph building module is configured to perform node addition and edge generation operations.

[0125] When a new atomic event is received, the module performs a node addition operation, creating a corresponding event node in the graph.

[0126] Then, the module performs an edge generation operation, establishing relationship edges between the newly created node and existing nodes in the graph based on a set of preset association rules.

[0127] To enable edge creation, the graph construction module is configured to create edges that include at least temporally continuous relationships, subject-object relationships, and logical relationships.

[0128] To establish the relationship between the subject and the object, the graph construction module integrates a cross-modal semantic matching unit. This unit carries a pre-trained cross-modal embedding model, which determines whether to establish a relationship by calculating the cosine similarity of the embedding vectors of different event objects.

[0129] To ensure the system's real-time performance and resource availability, the graph construction module is also configured to maintain a dynamic semantic event graph within a sliding time window, automatically removing outdated nodes and related edges that exceed the time window.

[0130] Detection engine:

[0131] The detection engine is used to perform subgraph matching detection based on a violation pattern library and structural anomaly detection based on graph structure entropy change rate on a dynamic semantic event graph. To achieve this dual-channel parallel detection function, the detection engine includes a subgraph matching unit and a graph entropy analysis unit in its structure.

[0132] Subgraph matching unit:

[0133] The subgraph matching unit is configured to compare the local structure of a dynamic semantic event graph with a predefined library of violation patterns.

[0134] This unit contains an optimized subgraph isomorphic matching algorithm and is connected to a violation pattern library that stores multiple violation subgraph templates.

[0135] When the graph structure is updated, the subgraph matching unit is activated, which efficiently searches the graph for subgraph instances that match any template. If a match is successful, it outputs a known violation detection signal and the corresponding violation subgraph to the decision alarm module.

[0136] Graph entropy analysis unit:

[0137] The graph entropy analysis unit is configured to calculate and monitor the rate of change of graph structure entropy over time for dynamic semantic event graphs.

[0138] The graph entropy analysis unit is specifically configured to periodically calculate the graph structure entropy of the entire graph or key regions, and further calculate its rate of change relative to the previous period.

[0139] This unit also contains a dynamic baseline calculation component, which uses an exponential moving average algorithm to dynamically update the baseline threshold.

[0140] When the graph entropy analysis unit detects that the entropy change rate exceeds this dynamic baseline threshold, it determines that a structural anomaly has occurred and outputs a structural anomaly detection signal to the decision alarm module.

[0141] Decision-making alarm module:

[0142] The decision alarm module is used to receive the detection results from the detection engine and generate the final real-time detection alarm.

[0143] In this embodiment, the decision alarm module also integrates an evidence solidification unit, which is configured to solidify a corresponding evidence chain based on the source of the detection signal while generating an alarm.

[0144] If a signal is received from the subgraph matching unit, the evidence solidification unit extracts and stores the complete illegal subgraph structure that was matched.

[0145] If a signal is received from the graph entropy analysis unit, the evidence solidification unit extracts and stores the local graph structure and related event information within the time window that caused the drastic change in graph structure entropy.

[0146] Finally, the decision alarm module outputs alarm information and evidence chain index to the user interface or downstream processing system.

[0147] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A real-time detection method for live-streaming e-commerce based on multimodal fusion, characterized in that, Includes the following steps: S1. The real-time collected multimodal live data stream is structured into a sequence of atomic events containing timestamps, modal sources, subjects, predicates, and objects; S2. Based on the atomic event sequence, an incremental dynamic semantic event graph is constructed, wherein each atomic event constituting the atomic event sequence serves as a node of the graph, and the preset relationships between events constitute the edges of the graph. S3. On the dynamic semantic event graph, perform subgraph matching detection based on the violation pattern library and structural anomaly detection based on the graph structure entropy change rate; S4. Based on the results of the subgraph matching detection or the structural anomaly detection, generate real-time detection alarms for the live content.

2. The real-time detection method for live-streaming e-commerce based on multimodal fusion according to claim 1, characterized in that, The steps of structuring real-time acquired multimodal live data streams into atomic event sequences specifically include: From the video stream in the multimodal live data stream, video atomic events are extracted and generated using object detection, behavior recognition, and optical character recognition technologies; Audio atomic events are extracted and generated from the audio stream in the multimodal live data stream using automatic speech recognition technology; From the text stream in the multimodal live data stream, natural language processing technology is used to extract and generate text atomic events; The video atomic events, audio atomic events, and text atomic events together constitute an atomic event sequence.

3. The real-time detection method for live-streaming e-commerce based on multimodal fusion according to claim 1, characterized in that, The preset inter-event relationships include at least: Temporal continuity is used to connect atomic events emitted consecutively by the same subject within a preset time threshold. Subject-object association is used to connect atomic events of different modalities that point to the same semantic object; And logical relationships, used to connect atomic event pairs that conform to a preset logical model.

4. The real-time detection method for live-streaming e-commerce based on multimodal fusion according to claim 1, characterized in that, The subgraph matching detection based on the violation pattern library includes: matching the local structure of the dynamic semantic event graph with a predefined violation subgraph template in the violation pattern library.

5. The real-time detection method for live-streaming e-commerce based on multimodal fusion according to claim 1, characterized in that, The structural anomaly detection step based on graph structure entropy change rate specifically includes: Calculate the rate of change of the graph structure entropy of the dynamic semantic event graph over time; A structural anomaly is determined to have occurred when the rate of change exceeds a dynamically updated baseline threshold.

6. The real-time detection method for live-streaming e-commerce based on multimodal fusion according to claim 1, characterized in that, Also includes: When generating the real-time detection alarm, a chain of evidence is established; If triggered by the subgraph matching detection, then the evidence chain is a matched illegal subgraph structure; If triggered by the structural anomaly detection, the evidence chain is a local graph structure within the time window that causes a dramatic change in graph structure entropy.

7. The real-time detection method for live-streaming e-commerce based on multimodal fusion according to claim 1, characterized in that, In the step of structuring the real-time acquired multimodal live data stream into a sequence of atomic events, the video stream processing is implemented using a parallel pipeline architecture, specifically including: The frame distribution and preprocessing steps are used to capture image frames at a preset frame rate and perform standardized preprocessing. The parallel processing step synchronously sends the preprocessed image frames into the target detection pipeline, behavior recognition pipeline, and optical character recognition pipeline for parallel processing; The event generation and aggregation step is used to combine the output results of each pipeline with timestamps and uniformly format them into video atomic events.

8. The real-time detection method for live-streaming e-commerce based on multimodal fusion according to claim 1, characterized in that, The process of structuring into a sequence of atomic events also includes entity linking and normalization steps: The system maintains a dynamically updated entity knowledge base; When a new atomic event is generated, the subject and / or object information therein is matched with the entity knowledge base and mapped to a unified entity identity identifier.

9. The real-time detection method for live-streaming e-commerce based on multimodal fusion according to claim 5, characterized in that, The dynamically updated baseline threshold is calculated using an exponential moving average algorithm based on historical graph structure entropy change rate data.

10. A real-time detection system for live-streaming e-commerce based on multimodal fusion, characterized in that, include: The event abstraction module is used to structure real-time acquired multimodal live data streams into atomic event sequences containing timestamps, modal sources, subjects, predicates, and objects; The graph construction module is used to incrementally construct a dynamic semantic event graph based on a sequence of atomic events, where each atomic event constituting the sequence of atomic events serves as a node of the graph, and the preset relationships between events constitute the edges of the graph. The detection engine is used to perform subgraph matching detection based on a violation pattern library and structural anomaly detection based on graph structure entropy change rate on dynamic semantic event graphs. The decision alarm module is used to generate real-time detection alarms for live content based on the results of subgraph matching detection or structural anomaly detection.

Citation Information

Patent Citations

  • Remote sensing image abnormal event detection method and device based on multi-modal representation learning

    CN115346132A

  • Multilingual event causality identification method and system based on meta-learning with knowledge

    US20250036969A1