Live broadcast anomaly detection method and electronic equipment

By using multimodal feature fusion and semantic similarity judgment, the accuracy problem of live broadcast anomaly detection has been solved, realizing automated and accurate anomaly detection and reducing the need for manual review.

CN120916008APending Publication Date: 2025-11-07BEIJING BLACK CAT GUARDIAN TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511012352.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

During live streaming, it is difficult for moderators to detect abnormal behavior in a timely manner, resulting in low accuracy in identifying anomalies.

Method used

By acquiring video, audio, and text streams from live broadcast data, multimodal features are extracted, cross-modal correction and intramodal aggregation are performed, and contribution fusion is combined to calculate the semantic similarity between the overall features and standard abnormal features. The similarity threshold is dynamically adjusted to determine live broadcast anomalies.

Benefits of technology

It enables timely and accurate detection of live stream anomalies, reduces manual review costs, and improves review efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120916008A_ABST
    Figure CN120916008A_ABST
Patent Text Reader

Abstract

The invention discloses a live broadcast anomaly detection method and electronic equipment, and belongs to the technical field of terminals. The method comprises the following steps: performing feature extraction on modal data in live broadcast data to obtain multi-modal features; wherein the multi-modal features comprise at least two modal features of a video feature, an audio feature and a text feature; performing cross-modal correction and intra-modal aggregation processing on the multi-modal features to obtain target multi-modal features; wherein the target multi-modal features comprise target modal features corresponding to the modal features; based on the contribution degree corresponding to each target modal feature, fusing the target multi-modal features to obtain an overall feature; calculating the semantic similarity between the overall feature and the standard live broadcast abnormal feature; and when the semantic similarity is greater than or equal to the similarity threshold, outputting a live broadcast anomaly result, thereby ensuring the accuracy of live broadcast anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of terminals, and in particular, relates to a live broadcast abnormality detection method and an electronic device. BACKGROUND

[0002] With the development of Internet technology, live broadcast is widely applied in different fields. During live broadcast, the live broadcast may have abnormalities, such as fraud and other irregularities of a host of the live broadcast. Therefore, it is necessary to determine whether the live broadcast has abnormalities.

[0003] In the prior art, an auditor can watch a live broadcast to audit whether the watched live broadcast has abnormalities. However, because the number of live broadcasts is large, the auditor may need a colleague to watch multiple live broadcasts, which may cause the auditor to not see abnormal live broadcast segments, thereby causing omission of live broadcast abnormalities and reducing the accuracy of abnormality determination. SUMMARY

[0004] The present application provides a live broadcast abnormality detection method and an electronic device to improve the accuracy of abnormality determination.

[0005] In a first aspect, the present application provides a live broadcast abnormality detection method, which can be applied to an electronic device or related components (such as a chip in the electronic device), and the method comprises:

[0006] Obtaining live broadcast data; wherein the live broadcast data comprises at least two modal data in video stream data, audio stream data, and text stream data of a live broadcast;

[0007] Performing feature extraction on the modal data in the live broadcast data to obtain multi-modal features; wherein the multi-modal features comprise at least two modal features in video features, audio features, and text features;

[0008] Performing cross-modal correction and intra-modal aggregation processing on the multi-modal features to obtain target multi-modal features; wherein the target multi-modal features comprise target modal features corresponding to each of the modal features;

[0009] Fusing the target multi-modal features based on a contribution degree corresponding to each of the target modal features to obtain an overall feature;

[0010] Calculating a semantic similarity between the overall feature and a standard live broadcast abnormality feature;

[0011] In a case where the semantic similarity is greater than or equal to a similarity threshold, outputting a live broadcast abnormality result.

[0012] Optionally, the cross-modal correction and intra-modal aggregation processing on the multi-modal features to obtain target multi-modal features comprises:

[0013] The cross-modal attention mechanism is adopted to correct each of the modal features in the multi-modal feature, to obtain a corrected modal feature corresponding to each of the modal features.

[0014] For each of the corrected modal features corresponding to the modal features, the corrected modal features are aggregated by attention pooling to obtain a target modal feature corresponding to the modal feature.

[0015] Optionally, the method further comprises:

[0016] Each of the target modal features is input into a quality evaluation network model to obtain a reliability score corresponding to each of the target modal features.

[0017] According to the reliability score corresponding to each of the target modal features, in combination with a context vector of the live data, a contribution degree corresponding to each of the target modal features is obtained.

[0018] Optionally, the similarity threshold is determined according to a first threshold; wherein the first threshold is obtained by adjusting a base threshold according to an influence factor; the influence factor includes one or more of a risk coefficient, a historical factor, a context factor and a reliability factor.

[0019] The risk coefficient is related to a live type corresponding to the live data, the historical factor is related to historical live data, the context factor is related to context data of the live data, and the reliability factor is related to quality, consistency and missing information of the modal data.

[0020] Optionally, the first threshold is determined according to formula one τ_dynamic=τ_base*(1-risk_factor)*(1-history_factor)*context_factor / reliability_factor.

[0021] Wherein, the τ_dynamic represents the first threshold, the τ_base represents the base threshold, the risk_factor represents the risk coefficient, the history_factor represents the historical factor, the context_factor represents the context factor, and the reliability_factor represents the reliability factor.

[0022] Optionally, the similarity threshold is determined according to a first threshold and a second threshold.

[0023] The first threshold value is obtained by adjusting a basic threshold value by an influencing factor; the influencing factor includes one or more of a risk coefficient, a historical factor, a context factor, and a reliability factor.

[0024] The second threshold value is obtained by prediction according to live streaming feature data; wherein the live streaming feature data includes one or more of live streaming scene features corresponding to the live streaming data, historical data, and live streaming abnormal type corresponding to the live streaming data.

[0025] Optionally, the similarity threshold value is obtained by weighted calculation on the first threshold value and the second threshold value.

[0026] The weight of the second threshold value is obtained by mapping the reliability factor.

[0027] In a second aspect, the present application further provides a live streaming abnormality detection device, which includes a data acquisition module, a feature extraction module, a feature processing module, a fusion module, a similarity calculation module, and an abnormality judgment module.

[0028] The data acquisition module is configured to acquire live streaming data; wherein the live streaming data includes at least two modal data of video stream data, audio stream data, and text stream data of live streaming.

[0029] The feature extraction module is configured to perform feature extraction on the modal data in the live streaming data to obtain multi-modal features; wherein the multi-modal features include at least two modal features of video features, audio features, and text features.

[0030] The feature processing module is configured to perform cross-modal correction and intra-modal aggregation processing on the multi-modal features to obtain target multi-modal features; wherein the target multi-modal features include target modal features corresponding to each of the modal features.

[0031] The fusion module is configured to fuse the target multi-modal features based on the contribution degrees of the target modal features corresponding to each of the target modal features to obtain overall features.

[0032] The similarity calculation module is configured to calculate semantic similarity between the overall features and standard live streaming abnormality features.

[0033] The abnormality judgment module is configured to output a live streaming abnormality result in a case where the semantic similarity is greater than or equal to a similarity threshold value.

[0034] In a third aspect, the present application further provides an electronic device including a memory and a processor, the memory storing a computer program or instructions, and the computer program or instructions being executed by the processor to cause the processor to perform the live streaming abnormality detection method according to the first aspect.

[0035] In a fourth aspect, the present application also provides a computer readable storage medium, which stores a computer program or instructions, and the computer program or instructions are executed by a processor to implement the live streaming exception detection method in the first aspect.

[0036] In a fifth aspect, the present application also provides a computer program product, which comprises a computer program or instructions, and the computer program or instructions are executed by a processor to implement the live streaming exception detection method in the first aspect.

[0037] In the present application, the modal data in the live streaming data is subjected to feature extraction to obtain multi-modal features. Then, the multi-modal features are subjected to cross-modal correction and intra-modal aggregation processing to obtain target modal features which can better reflect the live streaming content corresponding to the live streaming data. Then, the target multi-modal features are fused based on the contribution degrees corresponding to the target modal features to obtain an overall feature which can comprehensively reflect the vector of the overall semantics of the current live streaming segment of the live streaming content. Then, if the semantic similarity between the overall feature and the standard live streaming exception feature is greater than or equal to a similarity threshold, it indicates that the similarity between the overall feature and the standard live streaming exception feature is high, that is, the overall feature is an exception feature, and thus it can be determined that the live streaming content is abnormal, thereby realizing timely and accurate detection of live streaming exceptions and ensuring the accuracy of live streaming exception determination. Moreover, a large number of auditors are not required to perform auditing, and auditors are not required to simultaneously audit whether the live streaming content of multiple live streams is abnormal, thereby reducing labor costs and improving auditing efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained according to these drawings without creative labor.

[0039] In order to more completely understand the present application and its beneficial effects, the following will be described in conjunction with the drawings, wherein the same reference numerals in the following description represent the same parts.

[0040] Figure 1 Flowchart of the live streaming exception detection method provided by the embodiments of the present application Figure 1 .

[0041] Figure 2 Flowchart of the live streaming exception detection method provided by the embodiments of the present application Figure 2 .

[0042] Figure 3A flowchart of a live broadcast exception detection method provided by an embodiment of the present application Figure 3 .

[0043] Figure 4 A structural diagram of a live broadcast exception detection device provided by an embodiment of the present application

[0044] Figure 5 A structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort are within the protection scope of the present application.

[0046] In the embodiments of the present application, at least one refers to one or more; multiple refers to two or more than two. In the description of the present application, the terms "first", "second", "third" and the like are only used for distinguishing the purposes of description, and cannot be understood as indicating or implying relative importance, nor can be understood as indicating or implying order.

[0047] In the present specification, the reference "one embodiment" or "some embodiments" and the like means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Therefore, the terms "include", "contain", "have" and their variants in the present specification mean "include but not limited to", unless otherwise specifically emphasized.

[0048] It should be noted that in the embodiments of the present application, the association relationship of the associated objects described by "and / or" represents that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B together, and the existence of B alone. In addition, the character " / ", unless otherwise specified, generally represents a "or" relationship between the associated objects before and after it.

[0049] It should be noted that in the embodiments of the present application, "connection" can be understood as electrical connection, and the connection between two electrical elements can be direct or indirect connection between the two electrical elements. For example, A and B are connected, which can be direct connection between A and B, or indirect connection between A and B through one or more other electrical elements.

[0050] The embodiments of the present application provide a live broadcast exception detection method, which can be executed by an electronic device. As shown in Figure 1 the method can include S101-S108.

[0051] S101, the electronic device acquires live data. The live data includes at least two modalities of video stream data, audio stream data, and text stream data of live streaming.

[0052] For example, the live data can be current live data of online live streaming, or live data of historical live video.

[0053] In one case, the electronic device acts as a server. As shown in Figure 2 The user terminal collects live data and sends the collected live data to the server, so that the server can acquire live data of different user terminals to simultaneously detect whether the live data of different user terminals is abnormal. Moreover, since the server detects whether the live data is abnormal, the user terminal does not need to detect, which can reduce the pressure on the terminal side. It should be understood that the different user terminals may collect live data of different live rooms, but may also collect live data of the same live room. However, whether the live rooms are the same or not, the server detects respectively.

[0054] In another case, the electronic device acts as a user terminal. As shown in Figure 3 The user terminal collects live data to detect whether the live data is abnormal.

[0055] In some embodiments, the video stream data of the live streaming can be obtained from the video captured from the live data stream, such as determined by frame sampling. The frame sampling refers to intelligent key frame extraction or fixed interval sampling (for example, 1-5 frames per second, and the number of frames is dynamically adjusted according to the content) of the live video stream, so as to reduce redundancy and balance the calculation efficiency and information integrity.

[0056] In addition, the electronic device can also perform size normalization, enhancement, denoising, etc. on the sampled frames.

[0057] The audio stream data of the live streaming can be obtained from the audio captured from the live data stream. The audio stream data can be live speech text. For example, the electronic device divides the continuous audio stream into small segments (such as 30-60 seconds), performs noise suppression and voice enhancement, thereby realizing segmentation and noise reduction. Then, the electronic device can perform speech recognition (ASR) to convert the speech content in the audio into text in real time for subsequent text analysis.

[0058] The text stream data of the live broadcast can be obtained from text captured from the live data stream. The text can include user barrage, comments, interactive messages between the host and the platform, etc. The text can be determined by screen text recognition, such as optical character recognition (OCR). For example, static or dynamic text (such as titles, annotations, product information) in a video frame is recognized and extracted. Then, text cleaning is performed to remove irrelevant characters, emoticons (or convert them to specific labels), unify the code, etc.

[0059] In addition, optionally, the text stream data can further include one or more of host information, product links, historical violation records, etc.

[0060] In S102, the electronic device extracts features of the modality data in the live data to obtain multi-modal features. The multi-modal features include at least two modal features of video features, audio features, and text features.

[0061] In the embodiments of the present application, for each modality data in the live data, deep semantic features of the modality data are extracted to obtain modal features corresponding to the modality data, so as to convert original data of different modalities into a unified, high-dimensional, and semantic information-rich feature vector space.

[0062] Taking the live data including video stream data, audio stream data, and text stream data as an example, the electronic device extracts visual features of the video stream data to obtain video features corresponding to the video stream data, extracts audio features of the audio stream data to obtain audio features corresponding to the audio stream data, and extracts text features of the text stream data.

[0063] In some embodiments, the video features include static image features and dynamic video features. The static image features refer to deep semantic feature vectors of each frame of image extracted using a pre-trained convolutional neural network or vision.

[0064] The dynamic video features refer to spatiotemporal features learned from a frame sequence in combination with time sequence information.

[0065] In some embodiments, the audio features include acoustic features and text semantic features. The acoustic features refer to acoustic features such as mel-frequency cepstral coefficients (MFCC) and spectrograms of non-speech parts (such as background sound and specific sound effects) extracted.

[0066] The text semantic features refer to semantic vector representations of the converted text extracted using a pre-trained language model.

[0067] In some embodiments, the aforementioned text features include semantic vectors. Semantic vectors refer to the semantic vectors extracted from bullet comments, reviews, OCR text, etc., using a pre-trained language model.

[0068] It is understandable that the modal features mentioned above are actually feature sequences, including multiple features. For example, video features include two features: static image features and dynamic video features.

[0069] S103. The electronic device performs cross-modal correction and intra-modal aggregation processing on the multimodal features to obtain the target multimodal features. The target multimodal features include the target modal features corresponding to each modal feature.

[0070] Cross-modal correction involves learning the relationships between different modal features, associating and aligning them. For example, it associates the "red area" indicated by video features with the "apple" indicated by text features to achieve fine-grained alignment.

[0071] Intramodal aggregation refers to associating features within the same modality. As mentioned earlier, a modal feature is actually a feature sequence. Simply put, intramodal aggregation concatenates multiple features from the feature sequence into a single vector.

[0072] In some embodiments, the process of determining the above-mentioned target multimodal features may include:

[0073] A cross-modal attention mechanism is used to correct each modal feature in the multimodal features, so as to obtain the corrected modal features corresponding to each modal feature;

[0074] For each modal feature, the corrected modal features are aggregated using attention pooling to obtain the target modal features corresponding to the modal features. Accordingly, the target multimodal features include the target modal features corresponding to each modal feature. The target multimodal features can reflect the features of the live content corresponding to the live data from different dimensions.

[0075] For example, an electronic device can input various modal features into a cross-modal transformer to obtain corrected modal features corresponding to each modal feature after fine-grained attention correction, thus achieving cross-modal transformer interaction. Then, for each corrected modal feature, attention pooling is used to aggregate the corrected modal features into a single vector. Internally, the pooling uses the same-layer attention weights Aij to perform a weighted average of the tokens / patches of the corrected modal features.

[0076] S104, the electronic device fuses the target multi-modal features based on the contribution degrees corresponding to the respective target modal features, to obtain an overall feature.

[0077] In the embodiments of the present application, the contribution degrees of the respective target modal features are obtained by dynamically adjusting the contribution degrees of the respective modalities according to the reliability of the live data corresponding to the live content. Then, the electronic device adaptively fuses the target multi-modal features by using the contribution degrees corresponding to the respective target modal features, to obtain a vector (i.e., an overall feature) that can comprehensively reflect the overall semantics of the current segment of the live content, and to obtain a unified and more informative overall feature. The overall feature can be represented by a vector R.

[0078] In some embodiments, the process of determining the contribution degrees corresponding to the target modal features can include:

[0079] The respective target modal features are respectively input into a quality evaluation network model to obtain the reliability scores corresponding to the respective target modal features.

[0080] The contribution degrees corresponding to the respective target modal features are obtained according to the reliability scores corresponding to the respective target modal features and in combination with the context vector of the live data.

[0081] For example, first, the electronic device determines the reliability scores corresponding to the respective target modal features by using a quality evaluation network. That is, the reliability scores corresponding to the respective modal features are determined. Then, for each target modal feature, the reliability score corresponding to the target modal feature is concatenated with the context vector of the live data to obtain a concatenated feature corresponding to the target modal feature. Then, for each target modal feature, the electronic device inputs the concatenated feature corresponding to the target modal feature into a multilayer perceptron (MLP) to obtain an initial contribution degree corresponding to the target modal feature. Then, the electronic device can convert the initial contribution degree corresponding to the target modal feature into a contribution degree with a value between 0 and 1 by using a normalization function (Softmax). The sum of the contribution degrees corresponding to the respective target modal features is equal to 1.

[0082] The context vector can be obtained by performing feature extraction on the context information (or context data) of the live data. The process of feature extraction can be referred to in the foregoing, and will not be described here. For example, the context information includes historical modal information related to the current content and / or modal information of a future period of time, such as a live category, a historical portrait of a host, a time period, and the like.

[0083] In some embodiments, the electronic device can also input the modality features or modality data into the quality assessment network model to obtain a reliability score corresponding to each target modality feature.

[0084] The determination process of the contribution degree corresponding to the target modality feature is introduced above, and the process of fusing the target modality features based on the contribution degrees corresponding to the target modality features will be introduced below.

[0085] The electronic device can perform weighted calculation on the target modality features based on the contribution degrees corresponding to the target modality features to obtain an overall vector. Optionally, after the weighted calculation, a linear layer is used for normalization (LayerNorm) to obtain a final unified semantic vector.

[0086] It should be noted that the attention pooling provides a "micro" weight, and the contribution degree provides a "macro" weight, and the product of the two is the final effective weight of each token / patch. In addition, the role of the cross-modal Transforme is to allow tokens / patches of different modalities to interact through self-attention and cross-modal attention, adjust their representations, and make them more consistent with cross-modal semantic alignment.

[0087] The fusion based on the attention mechanism will be introduced below in combination with specific examples. The cross-modal attention mechanism is adopted, which allows features of different modalities to focus on each other, dynamically adjusts the weights of the modality features, and captures the complex correlations and complementary information between them.

[0088] The attention score as a dynamic weight is automatically calculated according to the similarity between the query vector and the key vector: Attention(Q, K, V) = softmax(QK^T / √d)V)

[0089] Each head learns different attention patterns, allowing the system to capture multiple interaction relationships.

[0090] Gating mechanism: control the flow of modality information through a learnable gating unit:

[0091] Each modality feature vector passes through a gating network layer: g_m = σ(W_g·f_m + b_g) where f_m is the feature vector of modality m, and g_m is the gating value (between 0 and 1).

[0092] When fusing modalities, the feature vector is multiplied by the corresponding gating value: f_fused = Σ(g_m·f_m). The gating value dynamically changes according to the current content context, such as when the video content is monotonous and the text information is rich, the weight of the text feature is increased.

[0093] In addition, the application is based on reinforcement learning weight optimization, and weight adjustment is regarded as a sequential decision problem. The state space is defined as the current multi-modal feature and context information. The action space is the adjustment of the weight coefficients of each modality. The reward function is based on the accuracy, confidence and artificial feedback of the violation identification. The optimal weight adjustment strategy is learned through deep Q learning (DQN) or policy gradient method. And with the accumulation of system experience, the weight adjustment strategy is gradually optimized, and the adaptability to complex scenes is improved.

[0094] Context-aware dynamic weighting dynamically adjusts the weight strategy according to the characteristics of the live scene and the historical live situation. For example, it is adjusted according to the historical portrait of the host. If the host corresponding to the live data has a specific type of violation history, the weight corresponding to the relevant modal feature will be increased. For example, if the host has used pictures to imply violations, the weight corresponding to the visual feature will be increased.

[0095] According to the live type (or live scene), the weight configuration is adjusted. The live type (such as game, shopping, entertainment) is identified, and targeted weight configuration is adopted. For example, in e-commerce live, the weight of the product display area and the price text is increased, that is, the weight corresponding to the video feature and the text feature is increased.

[0096] Optionally, the weight corresponding to the text feature can be smoothed by time sequence to make the weight not change abruptly but smoothly, avoiding judgment fluctuations. The smoothing can be realized by w_t = a w_t + (1-a) w_{t-1}. Wherein, a is the smoothing factor (which is between 0 and 1), w_t represents the current weight, and w_{t-1} represents the weight at the last time.

[0097] For example, the host promotes a product A, and the promotional phrase of product A is exaggerated. Correspondingly, the dynamic weight adjustment process includes an initial stage: the weights of each modality are similar by default (video feature: 0.35, audio feature: 0.35, text: 0.30).

[0098] When the host starts to show product A and the comparison chart, the electronic device detects the importance of the visual area through the attention mechanism, and the gate mechanism automatically adjusts the weight (video feature: 0.50, audio feature: 0.25, text feature: 0.25).

[0099] The electronic device analyzes the video content and finds that there are obvious image processing traces (such as sharpening, color enhancement) in the comparison chart. When the host speaks the promotional "XXX", the audio text and the interaction with the barrage appear a high correlation (audience inquiry and host response), and the electronic device captures this correlation through the cross-modal attention mechanism and adjusts the weight (video feature: 0.40, audio feature: 0.20, text feature: 0.40).

[0100] It should be understood that the historical learning-based reinforcement learning module finds that the text mode (anchor commentary + barrage interaction) is the most critical for identifying exaggerated propaganda in such live scenes, and further optimizes the weight distribution.

[0101] In the present application, by means of comparative learning, multi-task learning and the like, the features of different modalities are uniformly mapped to a shared semantic space, so that contents with similar semantics (regardless of their original modalities) are close in distance in the space. And according to the context of the current content and the reliability of the modal information, the contribution of each modality is dynamically adjusted during fusion.

[0102] In some embodiments, the modal contribution and the hierarchical architecture of adaptive weights are introduced as follows:

[0103] The adaptive weight hierarchical system is designed in the present application, which distinguishes between macroscopic modal contribution (or alternatively described as contribution) and microscopic feature weight.

[0104] Among them, the contribution is an adaptive evaluation of the importance of the whole modality from a macroscopic level, which is determined by modality reliability and context factors. The feature weight is an adaptive evaluation of the importance of each feature in the modality from a microscopic level, which is mainly dynamically calculated by attention mechanism and gate network.

[0105] The dynamic interaction between adaptive levels includes: from top to bottom: the contribution affects the feature extraction process, and the modality with high contribution can adaptively obtain more computing resources. From bottom to top: the feature weight distribution feedback affects the contribution, such as the modality whose weight is concentrated in the key area can adaptively improve its overall contribution.

[0106] In addition, adaptive weight joint optimization: joint optimization framework: Loss=L_task+λ_1·L_contribution+λ_2·L_attention. Among them, L_task is the main task loss, which is used to train the classifier whether it is illegal (i.e. abnormal).

[0107] L_contribution is a regular term for contribution, which aims to keep the contribution consistent with the modality reliability score, and at the same time encourages the sparsity of the contribution distribution. Normalize the reliability score to a target distribution, and then use KL divergence L2 distance to constrain the contribution to approach the target distribution. For example, L_contribution=Σ_m KL(G_m||ρ·Reli_m+(1-ρ) / M) where ρ∈[0,1] controls whether to completely trust the reliability evaluation.

[0108] L_attention is to constrain the cross-modal attention matrix A, so that the attention is not evenly distributed.

[0109] The sizes of λ 1 and λ 2 can be set according to actual conditions. For example, λ 1 ∈ [0.1, 0.5] and λ 2 ∈ [0.05, 0.3].

[0110] From the foregoing, according to the context of the current content and the reliability of each modality information, the contribution degree of each modality feature corresponding to the fusion is dynamically adjusted. Among them, the reliability is the basis of adaptive fusion weight, and the electronic device comprehensively evaluates the reliability of each modality feature (such as the above target modality feature), and provides an objective basis for weight allocation.

[0111] The above reliability evaluation can include video feature corresponding reliability evaluation. For example, the reliability evaluation can include one or more of picture quality evaluation, content stability, visual information entropy, and scene understanding confidence.

[0112] The picture quality evaluation indicates that the clarity, brightness, contrast and other indicators are calculated by using BRISQUE and other no-reference image quality evaluation algorithms.

[0113] The content stability indicates that the unstable video is adaptively reduced in weight by detecting picture jitter, scene switching frequency, compression artifacts, etc.

[0114] The visual information entropy indicates that the image information amount is evaluated, and the scene with rich information obtains a higher adaptive weight.

[0115] The scene understanding confidence indicates that the understanding degree of the visual recognition model to the current scene is measured, and the visual modality weight is adaptively adjusted.

[0116] The above reliability evaluation can include audio feature corresponding reliability evaluation. Optionally, the reliability evaluation can include one or more of signal-to-noise ratio (SNR), speech intelligibility, audio signal quality, and speech presence rate.

[0117] The signal-to-noise ratio indicates that the signal-to-noise ratio of the audio stream is calculated, and the audio weight is adaptively reduced in a noisy environment.

[0118] The speech intelligibility indicates that the high-definition speech obtains a higher adaptive weight by evaluating the volume stability, speech speed appropriateness and ASR confidence.

[0119] The audio signal quality indicates that the frequency spectrum characteristics, dynamic range and distortion degree are analyzed, and the contribution degree of the audio in the fusion is automatically adjusted.

[0120] The speech presence rate indicates that the proportion of speech is calculated, and the weight of the pure background sound segment in semantic analysis is adaptively reduced.

[0121] The reliability evaluation can include reliability evaluation of text features. Optionally, the reliability evaluation can include one or more of text source reliability, recognition confidence, text quality, and text relevance.

[0122] Text source reliability, which indicates distinguishing direct input text and automatic speech recognition (ASR) / OCR recognized text, and adjusting the weight adaptively according to the source reliability.

[0123] Recognition confidence, which indicates using the confidence score output by ASR / OCR, and reducing the weight of low-confidence text adaptively.

[0124] Text quality, which indicates evaluating the grammatical integrity, semantic coherence, and information density, and obtaining a higher adaptive weight for high-quality text.

[0125] Text relevance, which indicates calculating the relevance to the current scene theme, and obtaining a higher adaptive weight for content with high relevance.

[0126] Optionally, the reliability evaluation can also include multi-modal internal consistency evaluation. The multi-modal internal consistency evaluation can include one or more of cross-modal semantic consistency, time alignment, and context coherence.

[0127] Cross-modal semantic consistency, which indicates comparing the consistency of semantic information transmitted by different modalities, and obtaining a higher fusion weight for a modality combination with high consistency.

[0128] Time alignment, which indicates evaluating the time synchronization of different modal signals, and adaptively adjusting the weight of time-mismatched signals.

[0129] Context coherence, which indicates detecting content coherence, and identifying and reducing the adaptive weight of content mutation points.

[0130] S105, the electronic device calculates the semantic similarity between the overall features and the standard live broadcast abnormality features.

[0131] Illustratively, the standard live broadcast abnormality features (R_illegal_pattern) can be obtained from a multi-modal knowledge base. The standard live broadcast abnormality features can be all abnormality features in the multi-modal knowledge base, or part of them.

[0132] The multi-modal knowledge base contains various abnormality features. Optionally, the multi-modal knowledge base contains not only a list of sensitive words, samples of illegal pictures / videos, and illegal audio clips, but also semantic vector representations of these illegal content, i.e., abnormality features. It can be understood that the multi-modal knowledge base should be continuously updated to absorb new illegal patterns and countermeasures, that is, to update the abnormality features in the multi-modal knowledge base.

[0133] In S106, the electronic device determines whether the semantic similarity is greater than or equal to a similarity threshold.

[0134] In the present application, when the semantic similarity is greater than or equal to the similarity threshold (τ_category), it indicates that the semantic similarity between the overall feature and the standard live streaming abnormal feature is high, and the possibility of the overall feature being an abnormal feature is high, in other words, the live streaming content corresponding to the live streaming data has a high possibility of being abnormal, therefore, the electronic device can perform S107.

[0135] When the semantic similarity is less than the similarity threshold, it indicates that the semantic similarity between the overall feature and the standard live streaming abnormal feature is low, and the possibility of the overall feature being an abnormal feature is low, in other words, the live streaming content corresponding to the live streaming data has a low possibility of being abnormal, therefore, the electronic device can perform S108.

[0136] For each standard live streaming abnormal feature, the electronic device can calculate the semantic similarity between the overall feature of the live streaming content corresponding to the live streaming data and the standard live streaming abnormal feature. When there is a semantic similarity greater than or equal to the similarity threshold, S107 can be performed.

[0137] In some embodiments, the similarity threshold described above can be preset.

[0138] In another embodiment, since the illegal content often has ambiguity and variability, the present application adopts a fuzzy matching strategy. If the semantic similarity between R and a certain R_illegal_pattern is greater than or equal to τ_category, the possibility of live streaming content being abnormal is high, therefore, the live streaming abnormal result can be output. τ_category is not fixed, but can be dynamically adjusted according to the violation type, historical data, context information, etc. to balance the accuracy and recall rate.

[0139] In one possible design, the similarity threshold is determined according to a first threshold (τ_dynamic). Optionally, the electronic device can take the first threshold as the similarity threshold.

[0140] The first threshold is obtained by adjusting a base threshold according to an influencing factor; the influencing factor includes one or more of a risk coefficient, a historical factor, a context factor, and a reliability factor.

[0141] The risk factor is related to a live streaming type corresponding to the live streaming data, such as a live streaming abnormal type. For example, the live streaming abnormal type can include a high-risk abnormal type (such as a live streaming with high-risk content such as violence), a medium-risk abnormal type, and a low-risk abnormal type (such as a live streaming with low-risk content such as improper dressing). The risk factor is determined according to risk_factor = base_risk * severity_multiplier. risk_factor represents the risk factor, base_risk represents a basic risk value, which can be a preset pointer. severity_multiplier represents a live streaming coefficient corresponding to the live streaming abnormal type. The live streaming coefficient corresponding to the high-risk abnormal type is less than the live streaming coefficient corresponding to the medium-risk abnormal type, which is less than the live streaming coefficient corresponding to the low abnormal type.

[0142] It should be understood that a more sensitive threshold is adopted for a high-risk violation, and a higher threshold can be adopted for a low-risk violation to reduce false positives.

[0143] The history factor is related to historical live streaming data. The history factor can be determined according to history_factor = f(violation_history, feedback_consistency, model_confidence). history_factor represents the history factor. violation_history represents a host historical violation record, feedback_consistency represents artificial review feedback, and model_confidence represents historical model performance.

[0144] The host historical violation record indicates that the history factor is dynamically reduced and the monitoring sensitivity is improved for a host with a high historical violation rate.

[0145] The historical model performance indicates that the threshold is adaptively adjusted according to the accuracy and recall rate performance of the analysis model on a specific type of content.

[0146] The artificial review feedback indicates that the setting of the history factor is dynamically optimized according to the confirmation or denial of the determination result (i.e., the abnormal detection result) by the reviewer.

[0147] For example, the electronic device can input the host historical violation record, the artificial review feedback, and the historical model performance into the analysis model to obtain the history factor.

[0148] The context factor is related to context data of the live streaming data. The context data includes one or more of a live streaming room attribute, a real-time content theme, user real-time feedback, and a time factor.

[0149] Live type, representing different types of live streaming (such as games, education, talent, etc.) apply different thresholds.

[0150] Real-time content theme, representing automatically identifying the current live streaming theme, adjusting the threshold for sensitive topics (such as specific political topics).

[0151] User feedback data, representing analyzing the frequency of audience reports, the emotion of bullet screen, etc. as auxiliary signals for threshold adjustment.

[0152] Time factor, representing threshold dynamic adjustment strategy considering special time period.

[0153] Taking the above context data including live room attribute, real-time content theme, user real-time feedback, and time factor as examples, the context factor can be determined according to context_factor=w1*room_factor+

[0154] w2*topic_factor+w3*user_feedback+w4*time_factor. context_factor represents the context factor. w1, w2, w3, and w4 are weights, which can be preset. room_factor represents the value corresponding to the live type, and the values corresponding to different live types are different. topic_factor represents the value corresponding to the real-time content theme. user_feedback represents the value corresponding to the user feedback data.

[0155] time_factor represents the value corresponding to the time factor. For example, the value corresponding to a special time period is different from the value corresponding to a non-special time period.

[0156] The reliability factor is related to the quality, consistency, and missing information of the modal data.

[0157] The quality of modal data, i.e. signal quality, means that when the quality of multi-modal data is greater than a value 1 and the signal-to-noise ratio is greater than a value 2, a higher value can be used.

[0158] Consistency, i.e. modal consistency, means that when different modal data points to the same violation judgment, a lower value can be selected to reduce the similarity threshold and increase the sensitivity.

[0159] Missing information, i.e. missing of key modal, means that when some key modal data is missing or of poor quality, the corresponding value can be selected to increase the similarity threshold to avoid misjudgment.

[0160] The above reliability factor can be determined by reliability_factor=avg_modal_quality*

[0161] modal consistency * (1 - critical modal missing). The reliability factor represents a reliability factor, the avg modal quality represents a value corresponding to the signal quality, the modal consistency represents a value corresponding to the consistency, and the critical modal missing represents a value corresponding to the missing information.

[0162] In some embodiments, the above-mentioned influence factors include a risk factor, a history factor, a context factor, and a reliability factor. Accordingly, the above-mentioned first threshold value can be determined as τ dynamic = τ base * (1 - risk factor) * (1 - history factor) * context factor / reliability factor. Wherein, τ dynamic represents the first threshold value, and τ base represents the base threshold value.

[0163] Optionally, the base threshold value can be set according to requirements, for example, generally set between 0.75-0.85, representing the baseline similarity requirement. In addition, the threshold value based on the threshold value can also be determined according to the live broadcast exception type.

[0164] Optionally, the influence factor value range is generally between 0.7-1.3, which comprehensively affects the final similarity threshold value.

[0165] In some embodiments, the above-mentioned similarity threshold value is determined according to the first threshold value and the second threshold value (τ optimal). Wherein, the determination process of the first threshold value can refer to the above.

[0166] The second threshold value is obtained by predicting according to the live broadcast feature data; wherein, the live broadcast feature data includes one or more of the live broadcast scene features (such as games, sports, etc.) corresponding to the live broadcast data, the historical data (such as historical live broadcast exception types or historical modal data) and the live broadcast exception type corresponding to the live broadcast data.

[0167] Optionally, the electronic device can input the live broadcast feature data into the threshold value prediction model to obtain the second threshold value. For example, τ optimal = ThresholdPredictor (scene features, violation type, history data). scene features represents the live broadcast scene features, violation type represents the live broadcast exception type, and history data represents the historical data.

[0168] In this application, a special threshold prediction model is trained, and the training target is to maximize the F1 score or a custom scoring function based on business weights to obtain the second threshold described above.

[0169] Of course, the second threshold described above can also be determined by other adaptive threshold algorithms. Dynamic threshold based on score distribution: analyze the distribution characteristics of the similarity score to identify the best decision boundary. ODIN (Out-of-Distribution detection) algorithm variant: use temperature scaling and perturbation enhancement to improve the robustness of the threshold. Bayesian optimal threshold: τ_optimal=argmin_τ[α·FPR(τ)+(1-α)·FNR(τ)], where α is a configurable risk trade-off parameter.

[0170] Optionally, the similarity threshold is obtained by weighting the first threshold and the second threshold. For example, τ_final=β·τ_optimal+(1-β)·τ_dynamic. Wherein, τ_final represents the weight of the second threshold.

[0171] Optionally, the weight of the second threshold is obtained by mapping the reliability score corresponding to the target modality feature. The size of β is obtained by transforming the reliability score corresponding to the target modality feature by an activation function (Sigmoid). The higher the reliability score, the larger β. For example, the electronic device can calculate the average of the reliability scores corresponding to all target modality features, and then obtain β by Sigmoid transformation of the average. Of course, the electronic device can also process the reliability score corresponding to the target modality feature in other ways, and then map the processed reliability score to obtain β. For example, the electronic device can map the reliability score corresponding to a certain target modality feature to obtain β.

[0172] In the embodiment of the application, τ_dynamic is a threshold calculated by the rule engine in real time, which completely depends on interpretable indicators such as risk factors, context factors, and reliability factors. τ_optimal comes from a data-driven threshold prediction model ThresholdPredictor. The current scene features, violation categories, historical judgment statistics, etc. are taken as input, and an optimal threshold is output by regression, thereby obtaining the second threshold described above.

[0173] In the embodiment of the application, in a complex and variable live broadcast environment, a fixed threshold may not be able to adapt to subtle differences in different live broadcast scenes, and may easily lead to misjudgment or omission. The application realizes a comprehensive dynamic threshold adjustment mechanism, which can intelligently adjust the judgment standard according to multi-dimensional factors.

[0174] S107, the electronic device outputs the live broadcast exception result.

[0175] In one case, the electronic device is a user terminal. When the user terminal detects that there is an anomaly in the live content on the user terminal, the user terminal can send a live anomaly result to the server, so that the server performs corresponding operations in response to the live anomaly result. For example, an auditor closes the live broadcast corresponding to the live anomaly result through the server. For another example, the auditor further determines whether there is an anomaly by using the live anomaly result received by the server.

[0176] In addition, the user terminal can also display the live anomaly result to prompt the user that there is an anomaly in the current live broadcast.

[0177] In another case, the electronic device is a server. When the server determines that there is a live anomaly based on the live data uploaded by the user terminal, the server can display or send the live anomaly result to the device of an auditor, so that the relevant auditor performs corresponding operations according to the live anomaly result.

[0178] In some embodiments, the live anomaly result can include a live anomaly risk level. The electronic device performs multi-label classification (such as violence, fraud, etc.) on the identified live data with risks, and comprehensively evaluates the live anomaly risk level (such as high, medium, and low) according to the matching strength, the severity of the violation mode, the context, and other factors.

[0179] In some embodiments, the live anomaly result can include one or more of the output violation type, risk level, and confidence. In addition, the electronic device can also trigger corresponding handling measures (such as alarm and evidence preservation) in response to the live anomaly result. For example, evidence preservation can refer to automatically intercepting video clips, audio clips, text screenshots, etc. containing violation content as evidence, and recording timestamp, anchor ID, etc.

[0180] In some embodiments, the electronic device can receive feedback of manual review, which is used to update the multi-modal knowledge base, adjust the model parameters, and optimize the matching rules, forming a closed-loop iteration.

[0181] S108, the electronic device returns to S101.

[0182] When no live anomaly is detected, the electronic device can continue to obtain live data, so as to continue to detect whether there is a live anomaly by using the live data, ensure the continuity of anomaly detection, avoid the situation of missing detection, and ensure the accuracy of anomaly detection.

[0183] The application provides a live broadcast content violation information identification method based on multi-modal semantic fuzzy matching. By deeply fusing multi-modal information such as images, videos, audios, texts (such as bullet screens, comments, anchor oral broadcast transcripts, screen texts, etc.) in a live broadcast stream, complex violation intentions expressed by multi-dimensional information can be captured, and hidden violations that cannot be found by single modal technology, such as "normal picture but voice violation" and "normal text but video action violation", can be effectively identified. And the semantic fuzzy matching technology can better understand the context, identify code, metaphor, variant, etc., effectively reducing the missed and mistaken judgments caused by literal matching. That is, understanding and fuzzy matching are performed at the semantic level, so that various explicit and implicit abnormal content can be more accurately, robustly and intelligently identified.

[0184] It should be noted that the collection of data involved in the present application is authorized and agreed by the user.

[0185] The above mainly introduces the scheme provided by the embodiments of the application from the perspective of the method. It can be understood that the electronic device includes a hardware structure and / or a software module corresponding to the execution of each function in order to realize the above functions. The units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed by hardware or computer-driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of the embodiments of the present application.

[0186] The embodiments of the present application can divide the functional modules of the electronic device according to the above method examples, for example, each functional module can be divided according to each function, or two or more functions can be integrated in one processing unit. The integrated unit can be in the form of hardware or in the form of a software functional module. It should be noted that the division of units in the embodiments of the present application is illustrative, and is only a logical functional division. Actual implementation can have another division method.

[0187] Figure 4 A live broadcast anomaly detection device structure schematic diagram provided in the embodiments of the present application. Please refer to Figure 4 The live broadcast anomaly detection device has the functions of the method examples on the electronic device side described above, and the functions can be realized by hardware or by hardware executing corresponding software. The live broadcast anomaly detection device can be the electronic device introduced above, or can be arranged in the electronic device. For example, Figure 4As shown, the live broadcast anomaly detection apparatus 400 can include a data acquisition module 410, a feature extraction module 420, a feature processing module 430, a fusion module 440, a similarity calculation module 450, and an anomaly judgment module 460.

[0188] The data acquisition module 410 is configured to acquire live broadcast data, wherein the live broadcast data includes at least two modal data in video stream data, audio stream data, and text stream data of live broadcast;

[0189] The feature extraction module 420 is configured to perform feature extraction on the modal data in the live broadcast data to obtain multi-modal features, wherein the multi-modal features include at least two modal features in video features, audio features, and text features;

[0190] The feature processing module 430 is configured to perform cross-modal correction and intra-modal aggregation processing on the multi-modal features to obtain target multi-modal features, wherein the target multi-modal features include target modal features corresponding to each of the modal features;

[0191] The fusion module 440 is configured to fuse the target multi-modal features based on a contribution degree corresponding to each of the target modal features to obtain an overall feature;

[0192] The similarity calculation module 450 is configured to calculate semantic similarity between the overall feature and a standard live broadcast anomaly feature;

[0193] The anomaly judgment module 460 is configured to output a live broadcast anomaly result in a case where the semantic similarity is greater than or equal to a similarity threshold.

[0194] Figure 5 A structural schematic diagram of an electronic device provided in some embodiments of the present application. Figure 5 The dashed line in the above-mentioned unit or module indicates that the unit or module is optional. Figure 5 The electronic device 500 in the above-mentioned unit or module can be used to implement the method described in the above-mentioned method embodiment. The electronic device 500 can be a user terminal (such as a mobile phone, a tablet computer, a notebook computer, etc.), or the electronic device 500 is a server.

[0195] The electronic device 500 can include one or more processors 510. The processor 510 can support the electronic device 500 to implement the method described in the foregoing method embodiments. The processor 510 can be a general-purpose processor or a special-purpose processor. For example, the processor can be a central processing unit (CPU). Alternatively, the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0196] The electronic device 500 can further include one or more memories 520. The memory 520 stores a computer program. The memory 520 can be independent of the processor 510 or integrated in the processor 510.

[0197] The electronic device 500 can further include a transceiver 530. The processor 510 can communicate with other devices or chips through the transceiver 530. For example, the processor 510 can perform data transceiving with other devices or chips through the transceiver 530.

[0198] The computer program in the memory 520 can be executed by the processor 510, so that the processor 510 performs the method as described above.

[0199] It should be understood that each step in the above method embodiments can be completed by the integrated logic circuit of hardware in the processor or the instruction in the form of software. The method steps disclosed in combination with the embodiments of the present application can be directly embodied as hardware processor execution completion or hardware and software module combination execution completion in the processor.

[0200] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program, when the computer program is run on a computer, the computer executes the method in any of the above embodiments,

[0201] In the embodiments of the present application, the storage medium can be a magnetic disk, an optical disk, a read only memory (ROM), or a random access memory (RAM), etc.

[0202] It should be noted that, for the business processing method of the embodiments of the present application, a person of ordinary skill in the art can understand that all or part of the processes of implementing the business processing method of the embodiments of the present application can be completed by a computer program controlling related hardware. The computer program can be stored in a computer readable storage medium, such as a memory of an electronic device, and executed by at least one processor in the electronic device. In the execution process, the processes of the embodiments of the business processing method can be included.

[0203] The embodiments of the present application also provide a computer program product or computer program, which includes computer instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the electronic device to perform the method provided in various optional implementation manners in the above embodiments.

[0204] It should be noted that, for the live broadcast exception detection method of the embodiments of the present application, a person of ordinary skill in the art can understand that all or part of the processes of implementing the live broadcast exception detection method of the embodiments of the present application can be completed by a computer program controlling related hardware. The computer program can be stored in a computer readable storage medium, such as a memory of an electronic device, and executed by at least one processor in the electronic device. In the execution process, the processes of the embodiments of the live broadcast exception detection method can be included.

[0205] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0206] The above is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with the preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make slight changes or modifications to the above disclosed technical content to obtain equivalent embodiments without departing from the scope of the technical solution of the present application. Any simplification, modification, equivalent change and modification of the above embodiments made according to the technical essence of the present application are still within the scope of the technical solution of the present application.

Claims

1. A live abnormality detection method, characterized by, The method comprises: acquiring live data; wherein the live data comprises at least two modal data in video stream data, audio stream data and text stream data of live broadcast; feature extraction is performed on the modal data in the live data to obtain multi-modal features; wherein the multi-modal features comprise at least two modal features in video features, audio features and text features; cross-modal correction and intra-modal aggregation processing are performed on the multi-modal features to obtain target multi-modal features; wherein the target multi-modal features comprise target modal features corresponding to each of the modal features; based on the contribution degree corresponding to each of the target modal features, the target multi-modal features are fused to obtain an overall feature; the semantic similarity between the overall feature and a standard live broadcast abnormal feature is calculated; in the case where the semantic similarity is greater than or equal to a similarity threshold, a live broadcast abnormal result is output.

2. The method of claim 1, wherein, The cross-modal correction and intra-modal aggregation processing of the multi-modal features to obtain target multi-modal features comprises: using a cross-modal attention mechanism, each of the modal features in the multi-modal features is corrected to obtain a corrected modal feature corresponding to each of the modal features; for each of the corrected modal features corresponding to the modal features, the corrected modal features are aggregated by attention pooling to obtain a target modal feature corresponding to the modal feature.

3. The method according to claim 1 or 2, characterized in that, The method further comprises: each of the target modal features is input into a quality evaluation network model to obtain a reliability score corresponding to each of the target modal features; according to the reliability score corresponding to each of the target modal features, in combination with a context vector of the live data, a contribution degree corresponding to each of the target modal features is obtained.

4. The method according to claim 1 or 2, characterized in that, The similarity threshold is determined according to a first threshold; wherein the first threshold is obtained by adjusting a base threshold by an influencing factor; the influencing factor comprises one or more of a risk coefficient, a historical factor, a context factor and a reliability factor; the risk coefficient is related to a live broadcast type corresponding to the live data, the historical factor is related to historical live data, the context factor is related to context data of the live data, and the reliability factor is related to quality, consistency and missing information of the modal data.

5. The method of claim 4, wherein, The first threshold is determined according to formula one τ_dynamic=τ_base*(1-risk_factor)*(1-history_factor)*context_factor / reliability_factor; wherein τ_dynamic represents the first threshold, τ_base represents the base threshold, risk_factor represents the risk coefficient, history_factor represents the historical factor, context_factor represents the context factor, and reliability_factor represents the reliability factor.

6. The method of claim 1 or 2, wherein, The similarity threshold is determined according to a first threshold and a second threshold; The first threshold value is obtained by adjusting a basic threshold value according to an influencing factor; the influencing factor includes one or more of a risk coefficient, a historical factor, a context factor, and a reliability factor; The second threshold value is obtained by prediction according to live broadcast feature data; the live broadcast feature data includes one or more of live broadcast scene features corresponding to the live broadcast data, historical data, and live broadcast anomaly types corresponding to the live broadcast data.

7. The method of claim 6, wherein, The similarity threshold value is obtained by weighted calculation of the first threshold value and the second threshold value. The weight of the second threshold value is obtained by mapping a reliability score corresponding to the target modality feature.

8. An electronic device, comprising: The memory stores a computer program or instructions, and the processor executes the computer program or instructions to perform the steps of the live broadcast anomaly detection method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer program or instructions are stored on the memory and executed by the processor to implement the steps of the live broadcast anomaly detection method according to any one of claims 1 to 7.

10. A computer program product, characterised in that, The computer program or instructions are stored on the memory and executed by the processor to implement the steps of the live broadcast anomaly detection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • SNN multi-mode target identification method, system, device and medium

    CN117892175A

  • Abnormal behavior detection method and system based on cross-modal fusion

    CN119169524A

  • Multimedia content AI detection method and device, equipment and storage medium

    CN119484890A

  • AI-based factory abnormal behavior identification monitoring method and system

    CN120088737A

  • Method for triggering platform algorithm preset value point location

    CN120234631A

Cited By

  • Live broadcast stream hotlinking real-time tracing method and system based on multi-source heterogeneous signaling fusion

    CN121644838A

  • A Real-Time Source Tracing Method and System for Live Stream Hotlinking Based on Multi-Source Heterogeneous Signaling Fusion

    CN121644838B

  • Anchor intention recognition method and device

    CN121744017A