Method for triggering platform algorithm preset value point location

Through differentiated rule base and multimodal fusion feature vector calculation, combined with audience feedback dynamic adjustment of thresholds, the live broadcast misjudgment problem caused by singlemodal eigenvectors is solved, and more accurate judgment of violations is achieved.

CN120234631AActive Publication Date: 2025-07-01ZHEJIANG UNIVERSE SINGULARITY TECH CO LTD

Patent Information

Application Number
CN202510705704.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In the prior art, live broadcast environment specification methods based on single-modal feature vectors are prone to misjudgment, and fail to effectively utilize multimodal information, resulting in misjudgment and imbalance in the judgment of violations.

Method used

A differentiated rule library is used to store the weights of different live broadcast risk types, and the similarity is calculated through multimodal fusion feature vectors, and the threshold is dynamically adjusted in combination with audience feedback to achieve accurate judgment of violations in the live broadcast process.

Benefits of technology

It reduces the probability of misjudgment of single-modal eigenvectors, improves the accuracy and consistency of live broadcast environment specifications, and reduces the misjudgment rate of violations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234631A_ABST
    Figure CN120234631A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of live broadcast data analysis, and particularly relates to a method for triggering a platform algorithm preset value point location, comprising the following steps: setting a differentiation rule base according to live broadcast risk types, the differentiation rule base being used for storing weights corresponding to different live broadcast risk types; data acquisition and preprocessing: acquiring and preprocessing video data, audio data and text data; feature extraction: respectively extracting video features, audio features and text features from the preprocessed video data, audio data and text data; and weight allocation: based on the weights stored in the differentiation rule base, setting the differentiation rule base and according to the identified live broadcast risk type, setting corresponding weight allocation for different live broadcast risk types preferentially, thereby generating differentiation when the multi-modal fusion feature vector is calculated, avoiding the same weight allocation, and improving the accuracy of the multi-modal fusion feature vector. The penalty standards of different live broadcast types are consistent, and misjudgment and penalty imbalance are caused.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of live broadcast data analysis, and specifically relates to a method for triggering the preset value points of the platform algorithm. Background Art

[0002] Network live broadcast is a relatively mainstream form of media presentation nowadays. During the live broadcast process, there may be some illegal acts, which may cause discomfort to the audience; when the audience participates in the live broadcast, they will interact with the live broadcast room, including sending bullet screens and connecting lines, etc., and there may be prohibited words or acts of infringing on the intellectual property rights of others. Therefore, in the current live broadcast environment, it is necessary to strictly regulate the live broadcast environment; triggering the platform algorithm can be understood as an algorithm that automatically performs specific operations based on preset conditions.

[0003] In the prior art, for prohibited words or illegal acts, generally, single-modal features in the live broadcast video are collected and compared with known illegal acts to achieve the purpose of regulating the live broadcast environment. For example, in an e-commerce live broadcast, the host is showing a red product, and there is text description in the background; at this time, the host inadvertently says a sentence that is similar but not illegal to the prohibited word. According to the single-modal feature vector calculation in the prior art, potential misjudgment may occur as follows: For the video modality, the system extracts features such as colors and shapes in the video to generate a video feature vector, and uses this video feature vector to calculate the similarity with the preset illegal picture feature vector (such as the pattern of prohibited items). Since the color of the product is similar to the color feature of a certain prohibited item, although other features such as the shape do not match, the single-modal similarity calculation may misjudge as an illegal picture due to the higher weight of the color feature; for the audio modality, the system extracts features such as the frequency spectrum and pitch in the audio to generate an audio feature vector, and uses this audio feature vector to calculate the similarity with the preset illegal audio feature vector. The sentence spoken by the host is similar to a certain prohibited word in terms of pitch and duration. Although the specific words are different, the single-modal similarity calculation may misjudge as illegal audio due to these similar features; for the text modality, the system extracts features such as word frequency and sentiment tendency in the background text to generate a text feature vector, and uses this text feature vector to calculate the similarity with the preset illegal text feature vector. Some words in the background text are similar to the words in the illegal text. Although the overall context is different, the single-modal similarity calculation may misjudge as illegal text due to these similar words. Based on the above, the single-modal feature vector only considers the information of a certain modality and ignores the complementary information of other modalities, resulting in easy misjudgment when the features are similar but not completely matched; Therefore, the present invention provides a method for triggering the preset value points of the platform algorithm. Summary of the Invention

[0004] In order to make up for the deficiencies of the prior art and solve at least one technical problem proposed in the background art.

[0005] The technical solution adopted by the present invention to solve its technical problems is as follows: A method for triggering the preset value points of a platform algorithm according to the present invention includes the following steps: S1: Set a differential rule library according to the live broadcast risk type, and the differential rule library is used to store the weights corresponding to different live broadcast risk types; S2: Data collection and preprocessing, including video data, audio data, and text data collection and preprocessing; S3: Feature extraction, respectively extract video features, audio features, and text features from the preprocessed video data, audio data, and text data; S4: Weight assignment, based on the weights stored in the differential rule library, assign dynamic weights to video, audio, and text features; S5: Multimodal fusion, represent the extracted video, audio, and text features as feature vectors 、 、 , and perform weighted fusion based on the weights; S6: Algorithm preset point trigger judgment, including: Represent the preset point as a preset feature vector ; Calculate the similarity between the fused feature vector and the preset feature vector ; ; Set a first threshold , if , then trigger the preset point, otherwise it is not triggered.

[0006] Preferably, the method for video data collection and preprocessing is: Real-time collect the live video stream, extract video frames at a fixed frame rate to obtain a video frame sequence ; Adjust the resolution of the video frames uniformly, and then perform denoising processing on the images; The method for audio data collection and preprocessing is: Synchronously record the live audio stream to obtain a continuous audio signal; Denoise the audio signal, and then perform frame splitting and windowing processing on the audio signal to obtain an audio signal sequence ; The method for text data collection and preprocessing is: Grab the live subtitle and bullet screen text information, perform word segmentation on the text to obtain word units, and label the part of speech for each word to obtain a text information sequence .

[0007] Preferably, the method for weighted fusion based on the weights is as follows: Identify the live broadcast risk type, and obtain the differential weight distribution based on the differential rule base 、 、 ; According to the formula: where ; is any live broadcast type; According to the formula, calculate the fused feature vector ; Calculate the similarity between the fused feature vector and the preset feature vector The method is as follows: According to the formula, calculate the similarity between the fused feature vector and the preset feature vector .

[0008] Preferably, the method for calculating the similarity further includes: Set a time window ; Calculate the cumulative feature vector within the time window , expressed as: where represents the fused feature vector at time Update the similarity calculation formula of the fused feature vector and the preset feature vector , expressed as: Based on the calculation, obtain the updated .

[0009] Preferably, the generation method of the preset feature vector is as follows: Collect multi-modal data of violations and extract single-modal features; Based on the single-modal features, fuse and generate a multi-modal feature vector; Based on the multi-modal feature vector, construct a database of violation behaviors, which also includes: Classify the multi-modal feature vector and Identify a unique sample ID for each multimodal feature vector.

[0010] Preferably, the differentiation rule library is also used to store a grading threshold, and match the first threshold based on the live broadcast risk type , and dynamically adjust according to real-time feedback.

[0011] Preferably, the differentiation rule library further includes a feedback mechanism, and the feedback mechanism is used to dynamically adjust the grading threshold according to real-time feedback; The dynamic adjustment method of the grading threshold is: Collect the audience feedback data during the live broadcast in real time; Based on natural language technology, distinguish the audience feedback data and count the number of negative emotions in the bullet comments , the number of reports and the interaction rate ; Based on the audience feedback data, calculate the adjustment amplitude of the grading threshold , according to the formula: where, is the bullet comment emotion influence coefficient, is the report influence coefficient, is the total number of bullet comments, is the interaction influence coefficient; According to the calculated adjustment amplitude update the first threshold , expressed as: where, is the second threshold; Apply the calculated second threshold to the updated , and judge whether a preset point is triggered.

[0012] Preferably, the calculation method of the interaction rate is: Count the number of user likes , shares and comments within a unit time; According to the formula: where, is the number of real-time online people in the live broadcast room.

[0013] Preferably, the time window The length is adjusted dynamically according to the live broadcast type, and the method is as follows: Identify the live broadcast risk type and define the time window according to the live broadcast risk type The length of, including: The first time window, s , satisfying that the live broadcast type is high risk; The second time window, s , satisfying that the live broadcast type is medium risk; The third time window, s , satisfying that the live broadcast type is low risk; Apply the matched time window To The update calculation of.

[0014] Preferably, the differential rule library further stores a risk level - first threshold mapping table, and the classification threshold is bound to the live broadcast type risk type. The specific rule is to automatically adapt the live broadcast risk type and the first threshold through the risk level - first threshold mapping table stored in the differential rule library. Automatically adapt.

[0015] The beneficial effects of the present invention are as follows: 1. For a method of triggering the preset value point of the platform algorithm according to the present invention, by setting a differential rule library and according to the identified live broadcast risk type, corresponding weight distributions can be preferentially set for different live broadcast risk types, so that when calculating the multi - modal fusion feature vector, differentiation can be generated, avoiding the same weight distribution, resulting in the same penalty criteria between different live broadcast types, causing misjudgment and penalty imbalance problems.

[0016] 2. For a method of triggering the preset value point of the platform algorithm according to the present invention, by based on the multi - modal fusion feature, and then performing a similarity calculation with known or preset violation behaviors in sequence to obtain the similarity , according to the size of the similarity , that is, comparing the similarity with the first threshold to judge whether there is a violation during the live broadcast process, and based on the similarity calculation with the preset violation behaviors in sequence, selecting the violation behavior corresponding to the maximum similarity value, which represents the possible violation behavior in the live broadcast room. Based on the multi - modal fusion feature vector on the same time axis, the probability of misjudgment is reduced. Brief Description of the Drawings

[0017] The present invention will be further described below with reference to the drawings.

[0018] Figure 1 Is the flowchart of the present invention. Detailed implementation mode

[0019] In order to make the technical means, creative features, achieved purposes and effects realized by the present invention easy to understand, the present invention will be further described below in conjunction with specific implementation modes.

[0020] As Figure 1 shown, a method for triggering the preset value points of the platform algorithm according to an embodiment of the present invention includes the following steps: S1: Set a differential rule library according to the live broadcast risk type, and the differential rule library is used to store the weights corresponding to different live broadcast risk types; S2: Data collection and preprocessing, including video data, audio data and text data collection and preprocessing; S3: Feature extraction, respectively extract video features, audio features and text features from the preprocessed video data, audio data and text data; S4: Weight assignment, based on the weights stored in the differential rule library, assign dynamic weights to video, audio and text features; S5: Multimodal fusion, represent the extracted video, audio and text features as feature vectors , , , and perform weighted fusion based on the weights; S6: Algorithm preset point trigger judgment, including: Represent the preset point as a preset feature vector ; Calculate the similarity of the fused feature vector and the preset feature vector ; ; Set a first threshold , if , then trigger the preset point, otherwise it is not triggered.

[0021] Calculating according to the single-modal feature vector in the prior art will cause potential misjudgment, as follows: For the video modality, the system extracts features such as color and shape in the video, generates a video feature vector, and calculates the similarity between the video feature vector and a preset feature vector of a violation scene (such as a contraband pattern). Since the color of the commodity is similar to the color feature of a certain contraband, although other features such as shape do not match, the single-modal similarity calculation may misjudge it as a violation scene due to the relatively high weight of the color feature; for the audio modality, the system extracts features such as spectrum and pitch in the audio, generates an audio feature vector, and calculates the similarity between the audio feature vector and a preset feature vector of a violation audio. The sentence spoken by the anchor is similar to a certain violation word in terms of pitch and duration. Although the specific words are different, the single-modal similarity calculation may misjudge it as a violation audio due to these similar features; for the text modality, the system extracts features such as word frequency and sentiment tendency in the background text, generates a text feature vector, and calculates the similarity between the text feature vector and a preset feature vector of a violation text. Some words in the background text are similar to the words in the violation text. Although the overall context is different, the single-modal similarity calculation may misjudge it as a violation text due to these similar words.

[0022] Based on the above, the single-modal feature vector only considers the information of a certain modality and ignores the complementary information of other modalities, resulting in easy misjudgment when the features are similar but not completely matched; based on this, in an embodiment of the present invention, when analyzing a live video, multi-modal features of the live video are collected respectively, preprocessed respectively, and then fused to obtain the multi-modal fusion feature of the live video. Based on the multi-modal fusion feature, the similarity is calculated successively with known or preset violation behaviors to obtain the similarity. , according to the similarity size, that is, the similarity and the first threshold Compare to determine whether there are violations during the live broadcast, and perform sequential similarity calculations based on preset violation behaviors, and select the violation behavior corresponding to the maximum similarity value, which represents the possible violation behavior in this live broadcast room. In addition, due to different live broadcast risk types, the corresponding violation behaviors may exist in videos, audio, and bullet screens respectively. Therefore, when calculating the multi-modal fusion features, it is also necessary to distinguish the weight distributions corresponding to different live broadcast risk types. It can be understood that during high-risk live broadcasts, such as game live broadcasts, there may be a high probability of prohibited words or uncivilized language. Then, for game live broadcasts, the weight ratio of the audio is relatively high. On the contrary, for low-risk live broadcasts, such as education live broadcasts, the probability of violations in the audio or video is relatively low, and the probability of violations in the text is relatively high. Therefore, the weight ratio of the text in education live broadcasts is relatively high. Based on the above differential rule library, according to the identified live broadcast risk type, the corresponding weight distribution can be preferentially set for different live broadcast risk types, so that when calculating the multi-modal fusion feature vector, differentiation will occur, avoiding the same weight distribution, resulting in the same penalty criteria between different live broadcast types, causing misjudgment and penalty imbalance problems.

[0023] In one embodiment, the method for video data acquisition and preprocessing is: real-time collect the live video stream, extract video frames at a fixed frame rate to obtain a video frame sequence ; adjust the resolution of the video frames uniformly, and then perform denoising processing on the images; The method for audio data acquisition and preprocessing is: synchronously record the live audio stream to obtain a continuous audio signal; perform denoising on the audio signal, and then perform frame segmentation and windowing processing on the audio signal to obtain an audio signal sequence ; The method for text data acquisition and preprocessing is: capture the live caption and bullet screen text information, perform word segmentation on the text to obtain word units, and label the part of speech for each word to obtain a text information sequence 。

[0024] When collecting live broadcast data, video data, audio data, and text data are mainly collected. Among them, for the preprocessing of video data, audio data, and text data, based on the sequence established with the time axis, it can ensure the time axis synchronization between multi-modal features during the later calculation of the multi-modal fusion feature vector, reducing the probability of misjudgment.

[0025] In one embodiment, the method for weighted fusion based on the weights is: Identify the live broadcast risk type, and obtain the differential weight distribution based on the differential rule library 、 、 ; According to the formula: in, ; Any live broadcast type; According to the formula, the fusion feature vector is calculated ; Calculate the fused feature vector With the preset feature vector Similarity The method is: According to the formula, the fusion feature vector is calculated With the preset feature vector Similarity .

[0026] Since the traditional method of comparing single-modal feature vectors with illegal behaviors has a high misjudgment rate, in one embodiment of the present invention, a multi-modal fusion feature vector synchronized with the time axis is used. Preset feature vectors corresponding to violations By making a comparison, the misjudgment or penalty imbalance caused by the single-mode feature vector can be reduced. Specifically, the weight distribution corresponding to the live broadcast risk type is obtained according to the differentiated rule library. , , , and then calculate the multimodal fusion feature vector based on the formula ,in , and At the same time The collected features are used to calculate the fused feature vector , applied to similarity calculation, and then obtain the fusion feature vector With the preset feature vector Similarity , based on the above, the similarity is calculated After that, with the preset first threshold A comparison is made to determine whether a preset value point is triggered, and the preset value point can be understood as whether any violation is triggered.

[0027] In one embodiment, the similarity is calculated The method also includes: Set time window ; Calculation time window The cumulative eigenvector in , expressed as: Among them, represents the fused feature vector at a moment; Update the fused feature vector and the preset feature vector The similarity calculation formula is expressed as: Based on the calculation, the updated .

[0028] Due to the fused feature vector calculated above and the preset feature vector The similarity , where , and are all features collected at the same moment , so misjudgment may be caused by instantaneous fluctuations. In an embodiment of the present invention, a time window is also introduced to smooth the short-term abnormal misjudgment caused by instantaneous fluctuations. Taking the detection of violent scenes as an example: Set the time window ; Real-time collect video, audio and text data within 5s; Extract the average color histogram and motion vector for each frame of video, extract the average audio energy and spectral features for each segment of audio, and extract the frequency of occurrence of violence-related keywords for text; Average the feature vectors within 5s to form a cumulative feature vector ; Calculate and the similarity with the preset feature vector of violent scenes ; If the similarity exceeds the first threshold , an alarm is triggered to indicate that a violent scene is detected; in the above embodiment, by introducing a time window and corresponding formula transformation, it is possible to more accurately determine whether a preset point is triggered during the live broadcast process, thereby reducing the misjudgment rate.

[0029] In one embodiment, the generation method of the preset feature vector is as follows: Collect illegal multimodal data and extract unimodal features; Based on the unimodal features, fuse and generate a multimodal feature vector; Based on the multimodal feature vector, construct an illegal behavior database, which also includes: Classify the multimodal feature vector and Identify a unique sample ID for each multimodal feature vector.

[0030] The generation of the preset vector P is the basis for the representation of multimodal fusion feature vectors. For example, in e-commerce live streaming, if the preset vector P contains features such as "red products + promotional terms", the system can quickly identify illegal advertisements; the specific steps include classifying and storing the fused multimodal feature vectors according to the types of violations and assigning unique IDs, and the real-time live data can be quickly matched with the known violation patterns in the database through similarity calculation; the database supports continuous addition of new samples for model retraining to improve detection capabilities; then, the illegal behaviors are classified in a fine-grained manner (such as violence being divided into physical conflicts, weapon displays, etc.), and a unique identifier is provided for each sample, and different types of violations correspond to different disposal strategies (such as interrupting the live broadcast or popping up a warning window), and the sample source can be traced through the ID to assist the review personnel in reviewing misjudged cases.

[0031] In one embodiment, the differential rule base is further used to store grading thresholds and match a first threshold based on the live broadcast risk type and dynamically adjust according to real-time feedback.

[0032] In the above description, since the occurrence locations of the illegal behaviors corresponding to different live broadcast risk types are different, the occurrence location described here mainly refers to the feature type, such as video features, audio features, or text features. Therefore, different weight assignments need to be matched. In addition, due to different live broadcast risk types, their inherent risks are also different. For example, for low-risk live broadcasts and high-risk live broadcasts, the calculated similarity distributions are also different. The similarity distribution corresponding to a low-risk live broadcast should correspond to a smaller numerical distribution, while the similarity distribution corresponding to a high-risk live broadcast should correspond to a larger numerical distribution. If the same first threshold is used, high-risk live broadcasts are more likely to be penalized, while low-risk live broadcasts are more difficult to be penalized. Based on the above, in one implementation manner of the present invention, grading thresholds are also set based on the live broadcast risk type. The grading thresholds are used to match different live broadcast risk types, that is, it can be understood that multiple first thresholds are set to match different live broadcast risk types. When analyzing whether the live broadcast data is illegal, by identifying the live broadcast risk type, a first threshold is matched for it ; Exemplarily: For high-risk live broadcast types, after identification, the matched first threshold ; For medium-risk live broadcast types, after identification, the matched first threshold ; For low-risk live broadcast types, after identification, the matched first threshold ; Based on the above, for high-risk live broadcast types, their live broadcast content may lead to a high similarity with the violation behaviors recorded in the violation behavior library. Therefore, the first threshold for matching is relatively high. For the corresponding low-risk live broadcast types, since there are few misjudgments in their historical live broadcasts, or rather, in their live broadcast history, the live broadcast content is quite different from the violation behaviors, so the similarity is normally relatively low. Therefore, a lower similarity threshold is matched for them.

[0033] In one embodiment, the differential rule library further includes a feedback mechanism, and the feedback mechanism is used to dynamically adjust the classification threshold according to real-time feedback; The method for dynamically adjusting the classification threshold is as follows: Collect the audience feedback data during the live broadcast in real time; Based on natural language technology, distinguish the audience feedback data and count the number of negative emotions in the bullet comments , the number of reports and the interaction rate ; Based on the audience feedback data, calculate the adjustment range of the classification threshold , according to the formula: where is the bullet comment emotion influence coefficient, is the report influence coefficient, is the total number of bullet comments, is the interaction influence coefficient; According to the calculated adjustment range update the first threshold , expressed as: where is the second threshold; Apply the calculated second threshold to the updated to determine whether a preset point is triggered.

[0034] Based on the above, after identifying the live broadcast risk type, the differential rule library matches the first threshold for live broadcast data analysis . Based on this, it is also necessary to consider the real-time feedback of the audience during the live broadcast, so as to perform a floating adjustment on the first threshold . The purpose is to improve the accuracy of the penalty through the floating-adjusted first threshold . In one embodiment of the present invention, by collecting the audience feedback data in real time and counting the number of negative emotions in the bullet comments based on natural language technology , the number of reports and the interaction rate ; the threshold adjustment range of the live broadcast data is calculated based on the above data according to the formula Based on the calculated threshold adjustment range It is also necessary to determine whether the threshold adjustment range meets the adjustment requirements. Specifically, set and When , it means that the threshold adjustment range calculated this time is less than the minimum adjustment value, so it is not adjusted, and the threshold is restored to the first threshold While when , adjustment is required, and the second threshold is calculated according to the formula Use the second threshold as the updated first threshold Based on the above, through the audience feedback data, the first threshold of the analyzed live broadcast data is adjusted floatingly , which is beneficial to quickly respond to live broadcast anomalies and reduce the imbalance probability of live broadcast penalties. Exemplarily: Assume the first threshold , negative bullet screen rate: , reporting rate: , interaction rate: , bullet screen emotion influence coefficient: , reporting influence coefficient: , interaction influence coefficient: ; Calculation process: Based on the above calculation, the threshold adjustment range , so the first threshold is not updated.

[0035] In one embodiment, the calculation method of the interaction rate is as follows: Statistical number of user likes , sharing number and comment number within a unit time; According to the formula: where is the number of real-time online users in the live broadcast room.

[0036] The above interaction rate is used to correct the adjustment range . When the interaction rate When it is higher, it can suppress the excessive lowering of the threshold due to the instantaneous amount of reports.

[0037] In one embodiment, the time window The length is dynamically adjusted according to the live broadcast type, the method is: Identify the risk type of live broadcast and define the time window according to the risk type of live broadcast Length, including: The first time window, s , the live broadcast type meets the high risk; The second time window, s , the live broadcast type meets the medium risk; The third time window, s , the live broadcast type is low risk; The time window that will match Apply to The update calculation is in progress.

[0038] After identifying the live broadcast type, you can also match time windows of different lengths based on the live broadcast risk type. , thereby reducing the imbalance in penalties for live broadcast violations. In addition, linking the time window with the risk level of the live broadcast type can improve the robustness of multimodal fusion.

[0039] The differentiation rule base also stores a risk level-first threshold mapping table, and the classification threshold is bound to the live broadcast type risk type. The specific rule is to map the live broadcast risk type with the first threshold through the risk level-first threshold mapping table stored in the differentiation rule base. Automatic adaptation.

[0040] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A method for triggering the preset point positions of a platform algorithm, characterized in that, It includes the following steps: S1: Set a differential rule base according to the live broadcast risk types, where the differential rule base is used to store the weights corresponding to different live broadcast risk types; S2: Data collection and preprocessing, including video data, audio data, and text data collection and preprocessing; S3: Feature extraction, extracting video features, audio features, and text features from the preprocessed video data, audio data, and text data respectively; S4: Weight assignment, assigning dynamic weights to video, audio, and text features based on the weights stored in the differential rule base; S5: Multimodal fusion, representing the extracted video, audio, and text features as feature vectors , , , and performing weighted fusion based on weights; S6: Algorithm preset point trigger judgment, including: Represent the preset point positions as preset feature vectors ; Calculate the fused feature vector with the preset feature vector similarity ; Set the first threshold , if , the preset point is triggered; otherwise, it is not triggered.

2. A method for triggering the preset value point of the platform algorithm according to claim 1, characterized in that: The method for video data acquisition and preprocessing is as follows: Real-time collect the live video stream, extract video frames at a fixed frame rate to obtain a video frame sequence ; Unify the resolution adjustment of video frames, and then perform denoising processing on the images; The method for audio data acquisition and preprocessing is as follows: synchronously record the live audio stream to obtain continuous audio signals; denoise the audio signals, and then frame and window the audio signals to obtain an audio signal sequence ; The method for text data collection and preprocessing is as follows: capture live caption and bullet screen text information, perform word segmentation on the text to obtain word units, and label the part of speech for each word to obtain a text information sequence 。 3. A method for triggering a preset value point of a platform algorithm according to claim 2, characterized in that: The method of weighted fusion based on the weights is: Identify live broadcast risk types and obtain differentiated weight allocation based on a differentiated rule base 、 、 ; According to the formula: Among them, ; is any live broadcast type; According to the formula, the fused feature vector is calculated ; Calculating the fused feature vector and the preset feature vector similarity The method is as follows: According to the formula, the fused feature vector is calculated and the preset feature vector for similarity .

4. A method for triggering a preset value point of a platform algorithm according to claim 3, characterized in that: Calculating the similarity The method further includes: Set time window ; Calculation time window The cumulative eigenvector within , expressed as: Among them, denotes the fused feature vector at the moment; Updated fused feature vector and a preset feature vector The similarity calculation formula is expressed as: Based on the calculation, the updated is obtained.

5. A method for triggering the preset value points of the platform algorithm according to claim 4, characterized in that: The aforementioned preset feature vector is generated in the following manner: Collect multi-modal data of violations and extract single-modal features; Based on the single-modal features, fuse and generate a multi-modal feature vector; Based on the multi-modal feature vector, construct a database of violation behaviors, and also include: Classify the multi-modal feature vectors and Assign a unique sample ID to each multi-modal feature vector.

6. A method for triggering the preset value points of the platform algorithm according to claim 5, characterized in that: The differential rule library is also used to store the classification thresholds, match the first threshold based on the live broadcast risk type , and dynamically adjust according to the real-time feedback.

7. A method for triggering a preset value point of a platform algorithm according to claim 6, characterized in that: The differential rule base also includes a feedback mechanism, and the feedback mechanism is used to dynamically adjust the grading threshold according to real-time feedback; The method for dynamically adjusting the grading threshold is: Real-time collect the audience feedback data during the live broadcast process; Distinguish the audience feedback data based on natural language technology, and count the number of negative emotions in the bullet comments , the number of reports and the interaction rate ; Calculate the adjustment range of the grading threshold based on the audience feedback data , according to the formula: Among them, is the bullet screen emotion influence coefficient, is the reporting influence coefficient, is the total number of bullet screens, is the interaction influence coefficient; According to the calculated adjustment range Update the first threshold , expressed as: Among them, is the second threshold value; Apply the calculated second threshold to the updated to determine whether a preset point is triggered.

8. A method for triggering the preset value point of the platform algorithm according to claim 7, characterized in that: The interaction rate is calculated as follows: Count the number of user likes within a unit of time , the number of shares and the number of comments ; According to the formula: Among them, is the real-time online number of people in the live broadcast room.

9. A method for triggering the preset value point of the platform algorithm according to claim 8, characterized in that: The time window has a length that is dynamically adjusted according to the live broadcast type, and the method is as follows: Identify the live broadcast risk types and define the time window according to the live broadcast risk types The length of which includes: The first time window, s , meeting the live broadcast type being high-risk; The second time window, s , meets the requirement that the live broadcast type is medium risk; The third time window, s , where the live broadcast type meets the low-risk requirement; Apply the matching time window to in the updated calculation.

10. A method for triggering the preset value points of the platform algorithm according to claim 9, characterized in that: The differential rule library also stores a risk level - first threshold mapping table, and the classification threshold is bound to the live broadcast type risk type. The specific rule is that through the risk level - first threshold mapping table stored in the differential rule library, the live broadcast risk type is automatically adapted to the first threshold automatically.

Citation Information

Patent Citations

  • Video auditing method and device, equipment and storage medium

    CN109495766A

  • Live video identification method and device, computer equipment and storage medium

    CN114760484A

  • Multi-mode network content security intelligent auditing system and method thereof

    CN118312922A

  • Dangerous behavior identification and early warning method based on multi-modal analysis

    CN119360278A

  • Multimedia content AI detection method and device, equipment and storage medium

    CN119484890A

Cited By

  • Live broadcast anomaly detection method and electronic equipment

    CN120916008A