Processing method for realizing multi-modal data auditing based on machine learning

Through the multimodal data audit method based on machine learning, text, audio and video data are identified and processed, and the accuracy and stability of multimodal data audit in the existing technology are solved, achieving efficient violation identification and accurate risk assessment.

CN119989242AActive Publication Date: 2025-05-13BEIJING LIUJINSUIYUE TECH CO LTD

Patent Information

Application Number
CN202510465294.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively identify and review violation information in multimodal data, resulting in poor stability of audit results and low analysis accuracy.

Method used

Using machine learning-based processing methods, by identifying the content and format of the samples to be reviewed, feature data of text, audio and video are extracted, a unified feature data group is constructed, and risk assessment is conducted based on the proportion of violations and matching degree, risk coefficients are dynamically adjusted, and processing level is finally determined.

Benefits of technology

It realizes unified auditing and processing of multimodal data, improves the accuracy of cross-modal violation identification, reduces the misjudgment rate, and improves the robustness and adaptability of audits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989242A_ABST
    Figure CN119989242A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a processing method for realizing multi-modal data auditing based on machine learning, which comprises the following steps: identifying the content of a to-be-audited sample, determining a sample format and a data processing mode, and extracting different format features in the to-be-audited sample to construct a feature data group; determining a risk coefficient of the to-be-audited sample according to the feature data set; obtaining the position of the audio or video in the to-be-audited sample, determining a matching range according to the position, obtaining the content matching degree of the text and the audio or video in the matching range, and judging whether to adjust the risk coefficient or not according to the content matching degree; and comparing the final risk coefficient with a coefficient threshold, and determining the processing level of the to-be-audited sample according to a comparison result. According to the method and the device, unified auditing processing of various formats of data such as texts, audios and videos is realized, and the accuracy of cross-modal violation recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a processing method for realizing multimodal data audit based on machine learning. Background Art

[0002] With the widespread use of social media, short video platforms and online communication tools, the integration of audio, video and text has exploded. The diversification of user-generated content makes it difficult for traditional single-modal audit methods to meet the needs of complex audit scenarios. Especially in content security management, illegal information may appear in the form of a combination of different modes, such as text content combined with audio / video information to form implicit illegal expressions. Therefore, how to efficiently and accurately audit multimodal data has become a key direction for the development of content audit technology.

[0003] However, the current single-modality audit method is difficult to fully cover the risk assessment of cross-modal data. For example, relying solely on text analysis may not be able to identify potential illegal content in audio and video, while relying solely on audio and video analysis may ignore the importance of text information. In content that combines audio, video and text, illegal information is often distributed in different modes, making it difficult for the audit system to accurately extract and match relevant information, affecting the accuracy of the audit results. Existing risk assessment methods fail to fully consider the correlation between audio, video and text, and it is difficult to make a reasonable comprehensive judgment on multimodal data, resulting in poor stability of the audit results and affecting the effectiveness of content security management.

[0004] Therefore, it is necessary to design a processing method for multimodal data audit based on machine learning to solve the problems existing in current technology. Summary of the invention

[0005] In view of this, the present invention proposes a processing method for multimodal data review based on machine learning, aiming to solve the problems of poor stability and low analysis accuracy in current multimodal data review.

[0006] The present invention proposes a processing method for realizing multimodal data audit based on machine learning, comprising:

[0007] Identify the content of the sample to be reviewed, determine the sample format and determine the data processing method based on the sample format, extract different format features in the sample to be reviewed based on the data processing method, and construct a feature data group; the sample format includes text audio and text video; the data processing method includes natural language processing, image processing and audio processing;

[0008] Determine the risk factor of the sample to be reviewed based on the feature data group; when the sample format of the sample to be reviewed is text audio, determine the risk factor based on the audio violation ratio and the text violation ratio; when the sample format of the sample to be reviewed is text video, determine the risk factor based on the video violation ratio and the text violation ratio;

[0009] Obtaining the location of the audio or video in the sample to be reviewed, determining a matching range according to the location, obtaining a content matching degree between the text in the matching range and the audio or video, and determining whether to adjust the risk factor according to the content matching degree;

[0010] When it is determined that the risk coefficient is to be adjusted, the characteristic data group is compared with the historical audit data, a similarity set is determined according to the comparison result, and an adjustment coefficient is determined according to the similarity set to adjust the risk coefficient to obtain a final risk coefficient;

[0011] The final risk coefficient is compared with the coefficient threshold, and the processing level of the sample to be reviewed is determined according to the comparison result.

[0012] Furthermore, based on the data processing method, different format features of the sample to be reviewed are extracted to construct a feature data group, including:

[0013] The feature data set includes text features, audio features and video features;

[0014] The natural language processing includes: removing special characters, segmenting the text in the sample to be reviewed based on BERT, extracting keywords and calculating the word frequency and inverse document frequency of the keywords, obtaining the TF-IDF value of each keyword according to the word frequency and inverse document frequency, sorting the keywords according to the TF-IDF values, and obtaining the text features according to the sorting;

[0015] The audio processing includes: extracting audio text based on speech recognition, obtaining keywords in the audio text, extracting high-dimensional audio features using Wav2Vec and extracting timbre features using Mel frequency cepstral coefficients, and obtaining the audio features according to the keywords, high-dimensional audio features and timbre features;

[0016] The image processing includes: calculating the motion changes between video frames based on optical flow to obtain a key frame sequence, performing target detection and scene recognition on the key frames respectively, and obtaining target detection features and scene recognition features; extracting video and audio, and obtaining video and audio features based on the video and audio, and obtaining the video features based on the video and audio features, target detection features and scene recognition features.

[0017] Furthermore, when determining the risk factor of the sample to be reviewed according to the characteristic data group, it includes:

[0018] When determining the risk factor based on the audio violation ratio and the text violation ratio,

[0019] Compare the text features with the illegal text database to determine the percentage of illegal texts;

[0020] According to the audio features, the proportion of illegal timbre, the proportion of illegal high-dimensional audio, and the proportion of illegal audio words are obtained, and the illegal audio proportion is determined;

[0021] Determine the text-audio weight according to the ratio of text to audio, and obtain the risk coefficient according to the text-audio weight;

[0022] When determining the risk factor based on the video violation ratio and the text violation ratio,

[0023] Obtaining the target detection violation ratio, the scene recognition violation ratio, and the violation video word ratio according to the video features, and determining the video violation ratio;

[0024] The text-video weight is determined according to the ratio of text to video, and the risk coefficient is obtained according to the text-video weight.

[0025] Further, when determining the matching range according to the location, it includes:

[0026] When the audio or video is in the first 30% of the sample to be reviewed, the first 50% of the text in the sample to be reviewed is used as the matching range;

[0027] When the audio or video is in the first 30-50% of the sample to be reviewed, the first 70% of the text in the sample to be reviewed is used as the matching range;

[0028] When the audio or video is in the last 50% of the sample to be reviewed, all the text in the sample to be reviewed is used as the matching range.

[0029] Furthermore, obtaining the content matching degree between the text in the matching range and the audio or video includes:

[0030] When the sample format of the sample to be reviewed is text audio, the content matching degree is the similarity between the audio text and the text;

[0031] When the sample format of the sample to be reviewed is a text video, extract the video audio text according to the video audio, obtain the content similarity between the video audio text and the text, and when the content similarity is not zero, use the content similarity as the content matching degree; when the content similarity is zero, obtain the feature similarity between the video feature and the text feature, and use the feature similarity as the content matching degree, and the value range of the content matching degree is [0, 1];

[0032] Further, judging whether to adjust the risk factor according to the content matching degree includes:

[0033] When the content matching degree is less than or equal to the matching degree threshold, it is determined that the risk coefficient is adjusted; when the content matching degree is greater than the matching degree threshold, it is determined that the risk coefficient is not adjusted and the risk coefficient is used as the final risk coefficient.

[0034] Furthermore, the characteristic data group is compared with the historical audit data, and a similar set is determined according to the comparison result, including:

[0035] The historical audit data includes a plurality of historical feature data groups, a plurality of historical content matching degrees, and a plurality of historical final risk coefficients, and each historical feature data group corresponds to a historical content matching degree and a historical final risk coefficient;

[0036] When there is data in the historical audit data whose historical content matching degree is equal to the content matching degree, the historical feature data group corresponding to the historical content matching degree and the historical final risk coefficient are constructed as a similarity set;

[0037] When there is no data in the historical review data whose historical content matching degree is equal to the content matching degree, select the data in the historical review data whose absolute value of the difference between the historical content matching degree and the content matching degree is less than or equal to 0.1, and construct the corresponding historical feature data group and the historical final risk coefficient into a similarity set.

[0038] Furthermore, the adjustment coefficient is determined according to the similarity set to adjust the risk coefficient to obtain the final risk coefficient, including:

[0039] The correlation between the feature data group and each of the historical feature data groups is calculated based on cosine similarity, and the average correlation is obtained. The data in the similarity set whose correlation is greater than the average correlation is classified into a first set, the data in the similarity set whose correlation is less than the average correlation is classified into a second set, and the data in the similarity set whose correlation is equal to the average correlation is classified into a third set.

[0040] The adjustment coefficient is determined according to the first set, the second set and the third set to adjust the risk coefficient to obtain a final risk coefficient.

[0041] Further, when the adjustment coefficient is determined according to the first set, the second set, and the third set to adjust the risk coefficient, it includes:

[0042] Obtaining a ratio of each historical final risk coefficient in the set to the risk coefficient, and obtaining an average ratio according to the ratio of each historical final risk coefficient in the third set to the risk coefficient;

[0043] Obtain a first variance of all ratios in the first set and the average ratio, and obtain a second variance of all ratios in the second set and the average ratio, sum the first variance and the second variance and take the square root to obtain a sum of variances, sum half of the sum of variances and the average ratio to obtain the adjustment coefficient, and take the product of the adjustment coefficient and the risk coefficient as the final risk coefficient.

[0044] Furthermore, the final risk coefficient is compared with the coefficient threshold, and the processing level of the sample to be reviewed is determined according to the comparison result, including:

[0045] The final risk coefficient is compared with the first coefficient threshold and the second coefficient threshold respectively, and the processing level of the sample to be reviewed is determined according to the comparison results; the first coefficient threshold is less than the second coefficient threshold;

[0046] When the final risk coefficient is less than or equal to the first coefficient threshold, the processing level of the sample to be reviewed is determined to be the first level; when the final risk coefficient is greater than the first coefficient threshold and less than or equal to the second coefficient threshold, the processing level of the sample to be reviewed is determined to be the second level; when the final risk coefficient is greater than the second coefficient threshold, the processing level of the sample to be reviewed is determined to be the third level; and the first level indicates that the risk level is less than the second level, and the second level indicates that the risk level is less than the third level.

[0047] Compared with the prior art, the beneficial effect of the present invention is that it realizes the unified review and processing of data in various formats such as text, audio and video, and improves the accuracy of cross-modal violation identification. By identifying the content of the sample to be reviewed and determining its format, natural language processing, image processing and audio processing and other technical means are used for text audio and text video samples to ensure that data of different modalities can be effectively extracted, and a unified feature data group is constructed to improve data fusion capabilities. Introducing risk coefficients In the review of text audio and text video, a comprehensive evaluation is conducted based on the violation ratio of audio, video and text, respectively, so as to ensure that the violation content of different modal data is reasonably quantified and avoid the influence of a single modality on the overall review results. Introducing cross-modal matching degree analysis, by obtaining the location of audio or video content, determining its matching range with text, and calculating the matching degree, to determine whether the risk coefficient needs to be adjusted, thereby further reducing the misjudgment rate and improving the robustness of the review. By comparing the feature data group with the historical review data, constructing a similar set and determining the adjustment coefficient based on the similarity, the risk assessment is made more accurate, and the adaptive ability and intelligent level of the review are enhanced. By combining the risk coefficient with the threshold comparison, the processing level of the sample to be reviewed is determined, which improves the efficiency and reliability of content review and effectively solves the problems in existing technologies such as the difficulty in comprehensively identifying multimodal violation information, insufficient review accuracy, and unstable risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0049] Figure 1 A flowchart of a processing method for implementing multimodal data review based on machine learning provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0050] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features described in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0051] In some embodiments of the present application, see Figure 1As shown, a processing method for implementing multimodal data review based on machine learning includes:

[0052] S100: Identify the content of the sample to be reviewed, determine the sample format and determine the data processing method based on the sample format, extract different format features in the sample to be reviewed based on the data processing method, and construct a feature data group. The sample format includes text audio and text video. The data processing method includes natural language processing, image processing and audio processing.

[0053] S200: Determine the risk factor of the sample to be reviewed based on the feature data group. When the sample format of the sample to be reviewed is text and audio, the risk factor is determined based on the audio violation ratio and the text violation ratio. When the sample format of the sample to be reviewed is text and video, the risk factor is determined based on the video violation ratio and the text violation ratio.

[0054] S300: Obtain the location of the audio or video in the sample to be reviewed, determine the matching range based on the location, obtain the content matching degree between the text and the audio or video within the matching range, and determine whether to adjust the risk factor based on the content matching degree.

[0055] S400: When it is determined that the risk coefficient is to be adjusted, the characteristic data group is compared with the historical audit data, a similarity set is determined according to the comparison result, and an adjustment coefficient is determined according to the similarity set to adjust the risk coefficient to obtain a final risk coefficient.

[0056] S500: Compare the final risk factor with the factor threshold, and determine the processing level of the sample to be reviewed based on the comparison result.

[0057] Specifically, in step S100, the content of the sample to be audited is identified, its data format is determined, and an appropriate data processing method is selected based on the format to extract feature information of different modalities, and finally form a feature data group. It supports multiple data formats such as text, audio and video, and uses natural language processing, image processing and audio processing respectively to achieve effective parsing and information extraction of cross-modal data. In step S200, the risk coefficient is calculated based on the constructed feature data group, and different evaluation strategies are adopted in combination with the characteristics of the data format. When the sample is text audio, the audio violation ratio and the text violation ratio are calculated, and the overall risk coefficient is determined based on the ratio of the two. When the sample is text video, risk assessment is performed based on the video violation ratio and the text violation ratio. It ensures that different modal data can be reasonably quantified to avoid a single modality affecting the overall audit result. In step S300, in order to further improve the accuracy of the audit, the specific location of the audio or video content in the sample is obtained, and the matching range is determined based on the location. Within the matching range, the matching degree between the text and the audio or video content is calculated to determine the consistency of their semantics and content, so as to analyze whether there are implicit illegal expressions between the text and the audio and video. For example, if the text content and the audio content do not match, but each has a risk of violation separately, the matching degree calculation can be used to determine whether to adjust the risk assessment to reduce misjudgment. If it is determined that the risk factor needs to be adjusted, then enter step S400. If it is determined that the risk factor does not need to be adjusted, then enter step S500. In step S400, the extracted feature data group is compared with the historical audit data to find similar historical cases, and a similar set is constructed based on the comparison results. The adjustment coefficient is calculated based on the similar set, and the risk coefficient of the current sample is dynamically adjusted to make the risk assessment more consistent with the historical audit experience and improve the accuracy and adaptability of the risk assessment. In step S500, the risk coefficient finally calculated is compared with the preset risk coefficient threshold, and the processing level of the sample to be reviewed is determined based on the comparison result. If the risk coefficient is higher than the threshold, the sample is directly marked as high risk, triggering further manual review or automatic interception. If the risk factor is lower than the threshold, the sample risk can be judged to be low, reducing unnecessary interventions and thus improving audit efficiency.

[0058] It is understandable that through multimodal data fusion technology, in-depth analysis of text, audio and video data is achieved, and matching calculation and historical data comparison are introduced in the risk assessment process to improve the accuracy and stability of the audit. Compared with the traditional single-modality audit method, it can effectively identify cross-modal violation information and reduce the misjudgment rate. At the same time, through the construction of similar sets and the application of adjustment coefficients, it can dynamically optimize the audit strategy and enhance the intelligence level of audit decision-making. The step-by-step risk assessment method ensures the comprehensive processing capabilities of different modal data while ensuring efficient audits, making the audit more flexible and adaptable, and able to meet complex and changing content audit needs.

[0059] In some embodiments of the present application, different format features in the sample to be reviewed are extracted based on the data processing method, and when a feature data group is constructed, it includes: the feature data group includes text features, audio features and video features.

[0060] Natural language processing includes: removing special characters, segmenting the text in the audit sample based on BERT, extracting keywords and calculating the word frequency and inverse document frequency of the keywords, obtaining the TF-IDF value of each keyword based on the word frequency and inverse document frequency, sorting the keywords according to the TF-IDF value, and obtaining text features based on the sorting.

[0061] Audio processing includes: extracting audio text based on speech recognition, obtaining keywords in the audio text, extracting high-dimensional audio features using Wav2Vec and extracting timbre features using Mel-frequency cepstral coefficients, and obtaining audio features based on keywords, high-dimensional audio features and timbre features.

[0062] Image processing includes: calculating the motion changes between video frames based on optical flow, obtaining key frame sequences, performing target detection and scene recognition on key frames, obtaining target detection features and scene recognition features. Extracting video and audio, and obtaining video and audio features based on the video and audio, and obtaining video features based on the video and audio features, target detection features, and scene recognition features.

[0063] Specifically, in the process of processing text data, natural language processing methods are used to first remove special characters to ensure the purity and readability of text data. Subsequently, the BERT (Bidirectional Encoder Representations from Transformers) model is used to segment the text to obtain contextual information and improve text comprehension. In order to extract key semantic information, the keywords in the text are calculated, and the importance of each keyword is calculated based on the TF-IDF (Term Frequency-Inverse Document Frequency) method. Keywords with higher TF-IDF values ​​represent the core content of the text. By sorting the keywords, the core semantic information of the text is effectively extracted and used as text features.

[0064] Specifically, in terms of audio data processing, speech recognition (ASR, Automatic Speech Recognition) technology is used to convert audio content into text, and keywords are extracted from the audio text to ensure that the semantic information of the audio content is fully utilized. At the same time, in order to obtain richer audio features, Wav2Vec is used for high-dimensional audio feature extraction. This model is based on self-supervised learning and can effectively extract deep-level features in audio signals and improve the accuracy of audio recognition. In addition, Mel-Frequency Cepstral Coefficients (MFCC, Mel-Frequency Cepstral Coefficients) are used to extract timbre features. MFCC can effectively capture the spectral characteristics of audio and can identify different voice or sound effect characteristics. Combining keywords, high-dimensional audio features and timbre features to form complete audio feature data provides richer information for multimodal risk assessment.

[0065] Specifically, for video data, optical flow is used to calculate the motion changes between video frames to obtain the dynamic features of the video content and extract key frame sequences, thereby reducing the amount of data processing calculations while retaining the core information of the video. Based on the key frames, target detection and scene recognition are performed. Target detection is used to identify specific objects or people in the video, while scene recognition is used to understand the environment and context of the video. In addition, audio is extracted from the video, and the audio content is analyzed through video and audio feature extraction methods to ensure the collaborative processing capabilities of audio and video data. Finally, the video and audio features, target detection features, and scene recognition features are integrated to generate complete video feature data.

[0066] It is understandable that through the combination of natural language processing, audio processing and video processing technologies, efficient feature extraction of multimodal data is achieved, providing accurate and reliable feature data for multimodal risk assessment. Compared with traditional audit methods, BERT semantic analysis, TF-ID keyword screening, Wav2Vec deep audio feature extraction, MFCC timbre analysis, optical flow calculation, target detection and scene recognition technologies are used to ensure that key information of different modal data can be accurately extracted, and the intelligent level of audit is improved through cross-modal feature fusion. Especially in complex scenarios where text, audio and video are combined, it can effectively reduce misjudgments and missed judgments in the audit process, improve the ability to identify hidden illegal content, and at the same time reduce unnecessary computing resource consumption and improve audit efficiency.

[0067] In some embodiments of the present application, when determining the risk factor of the sample to be reviewed according to the feature data group, it includes:

[0068] When determining the risk factor based on the audio violation ratio and text violation ratio,

[0069] Compare the text features with the illegal text database to determine the percentage of text violations.

[0070] According to the audio features, the proportion of illegal timbre, the proportion of illegal high-dimensional audio and the proportion of illegal audio words are obtained, and the proportion of audio violations is determined.

[0071] The text-audio weight is determined according to the ratio of text to audio, and the risk coefficient is obtained according to the text-audio weight.

[0072] When determining the risk factor based on the video violation ratio and text violation ratio,

[0073] According to the video features, the target detection violation ratio, scene recognition violation ratio and violation video word ratio are obtained, and the video violation ratio is determined.

[0074] The text-video weight is determined according to the ratio of text to video, and the risk coefficient is obtained according to the text-video weight.

[0075] Specifically, the text features of the sample to be reviewed are extracted and compared with the illegal text database. The proportion of illegal information in the text content is determined by matching, and then the proportion of text violations is determined. Proportion of illegal timbre: Mel frequency cepstral coefficient (MFCC) and audio fingerprint technology are used to detect whether there are illegal timbre, such as counterfeit voice, abnormal tone, etc. Proportion of illegal high-dimensional audio: Deep learning methods such as Wav2Vec are used to analyze the high-dimensional features of audio, and potential illegal content is identified based on known illegal audio patterns. Proportion of illegal audio words: The audio text is transcribed based on speech recognition (ASR) and matched with the illegal text database to determine the degree of violation of the audio content. Combining the above three proportions, the proportion of audio violations is calculated comprehensively. Since the importance of text and audio may be different in different content, for example, audio is more important in commentary audio content, while text is dominant in subtitle interpretation content. Therefore, the text audio weight is calculated by the ratio of text to audio. According to the text audio weight, the text violation proportion and audio violation proportion are weighted and calculated to obtain the final risk coefficient.

[0076] Specifically, the same method is used to extract the text violation ratio for text and audio data, and then the video violation ratio is calculated. Based on object detection, the objects in the video are analyzed, such as whether they contain sensitive people, prohibited items, etc., and the ratio of illegal targets is calculated. Scene recognition is used to analyze the overall environment of the video, such as whether it involves violent scenes, illegal places, etc., and the ratio of illegal scenes is calculated. Video subtitles, embedded text and other content are extracted and compared with the illegal text database to calculate the ratio of illegal text in the video. Since the main information sources of different video content may be different (for example, in news reports, text subtitles may dominate, while short videos without subtitles rely on visual content), the text video weight is calculated by the ratio of text to video. According to the text video weight, the text violation ratio and the video violation ratio are weighted to obtain the final risk coefficient.

[0077] It is understandable that accurate and efficient risk assessment is achieved through the calculation of violation ratios and adaptive weight calculations of the three major modalities of text, audio, and video. Compared with the traditional independent modal analysis method, it can comprehensively consider multiple data characteristics to ensure the comprehensiveness and accuracy of the audit results. In particular, the combination of timbre, speech text, high-dimensional audio feature analysis, and target detection, scene recognition, and video text recognition can effectively capture hidden illegal content and reduce the rate of misjudgment and missed judgment. Through text audio weight calculation and text video weight calculation, the adaptability to different content types is improved.

[0078] In some embodiments of the present application, when determining the matching range according to the location, it includes:

[0079] When the audio or video is in the first 30% of the samples to be reviewed, the first 50% of the text in the samples to be reviewed is used as the matching range.

[0080] When the audio or video is in the first 30-50% of the sample to be reviewed, the first 70% of the text in the sample to be reviewed is used as the matching range.

[0081] When the audio or video is in the last 50% of the sample to be reviewed, the entire text in the sample to be reviewed is used as the matching range.

[0082] It is understandable that audio and video appear at the beginning of the content and are highly related to the first half of the text, while the subsequent text may involve different topics. Too large a matching range may introduce noise, thereby reducing the accuracy of the match. When audio and video content appears in the first 30%-50% range, the matching range is expanded and the first 70% of the text content is selected for matching. At this time, the audio and video may be in the transition part, involving the previous and next content, so a larger range of text needs to be covered to ensure the accuracy of the content matching. When the audio and video content is in the last 50%, the entire text content is selected as the matching range. Audio and video appear in the second half and may be related to the full text information. Especially in summary content, audio and video may be a supplement to the full text, so all text information needs to be included in the matching range to prevent missing key information. Through the dynamic adjustment of the position-aware matching range, the problem of unreasonable matching range between text and audio and video in traditional multimodal auditing methods is effectively solved. Compared with the method of fixed matching window, it can adaptively adjust the matching range, reduce noise interference, and improve the accuracy and stability of auditing. Especially in the combination of long text and audio and video content, it can reduce the interference of irrelevant content and improve the accuracy of cross-modal auditing.

[0083] In some embodiments of the present application, when obtaining the content matching degree of text and audio or video within the matching range, it includes: when the sample format of the sample to be reviewed is text audio, the content matching degree is the content similarity between the audio text and the text. When the sample format of the sample to be reviewed is text video, the video audio text is extracted according to the video audio, and the content similarity between the video audio text and the text is obtained. When the content similarity is not zero, the content similarity is used as the content matching degree. When the content similarity is zero, the feature similarity between the video feature and the text feature is obtained, and the feature similarity is used as the content matching degree. The value range of the content matching degree is [0, 1].

[0084] It is understandable that compared with single text matching, the addition of video feature matching can still provide matching basis through image target detection and scene recognition when the audio text matching degree is low, thus improving the accuracy of the audit. In addition, the matching degree value range is set in [0,1], making risk assessment more quantifiable and facilitating subsequent decision adjustments for intelligent audit.

[0085] In some embodiments of the present application, when determining whether to adjust the risk factor according to the content matching degree, it includes: when the content matching degree is less than or equal to the matching degree threshold, determining to adjust the risk factor. When the content matching degree is greater than the matching degree threshold, determining not to adjust the risk factor, and taking the risk factor as the final risk factor.

[0086] In some embodiments of the present application, the feature data group is compared with the historical audit data, and when the similarity set is determined based on the comparison result, it includes: the historical audit data includes several historical feature data groups, several historical content matching degrees and several historical final risk coefficients, and each historical feature data group corresponds to a historical content matching degree and a historical final risk coefficient.

[0087] Specifically, when there is data in the historical audit data where the historical content matching degree is equal to the content matching degree, the historical feature data group corresponding to the historical content matching degree and the historical final risk coefficient are constructed as a similar set. When there is no data in the historical audit data where the historical content matching degree is equal to the content matching degree, data in the historical audit data where the absolute value of the difference between the historical content matching degree and the content matching degree is less than or equal to 0.1 are selected, and the corresponding historical feature data group and the historical final risk coefficient are constructed as a similar set.

[0088] Specifically, historical audit data comes from the storage and recording of a large number of audited samples, including the corresponding feature data groups, content matching and final risk coefficients. These data accumulate the analysis results of multimodal content such as text, audio, and video in the past audit process, and provide a historical basis for the review of new samples. Specifically, if the content matching degree is low, it means that the correlation between the text content and the audio / video is weak, and there may be hidden cross-modal violations, and the risk coefficient needs to be adjusted to improve the accuracy of risk assessment.

[0089] It is understandable that risk adjustment is controlled by content matching threshold, which avoids unnecessary risk factor adjustment and improves audit efficiency. At the same time, by using the similarity comparison of historical audit data, it can still provide a reasonable reference basis when the data matching degree is low, reducing misjudgment and missed judgment. The intelligent optimization method based on historical data effectively improves the self-learning ability and decision-making stability of the audit, and enhances the accuracy and adaptability of cross-modal data audit.

[0090] In some embodiments of the present application, the risk coefficient is adjusted according to the similarity set to determine the adjustment coefficient, and when the final risk coefficient is obtained, it includes: calculating the correlation between the feature data group and each historical feature data group based on cosine similarity, and obtaining the average correlation, and classifying the data in the similarity set with a correlation greater than the average correlation into the first set, the data in the similarity set with a correlation less than the average correlation into the second set, and the data in the similarity set with a correlation equal to the average correlation into the third set. The risk coefficient is adjusted according to the first set, the second set, and the third set to obtain the final risk coefficient.

[0091] It is understandable that in order to ensure that the most similar set to the current sample to be audited can be screened out from the historical audit data, the cosine similarity calculation method is used to measure the correlation between the feature data group of the current sample and the feature data group of the historical sample, and the average correlation of all historical samples is calculated. By setting the threshold and dividing the set, the historical data with higher than average correlation is classified into the first set (highly similar data), the data with lower than average correlation is classified into the second set (low correlation data), and the data with equal average correlation is classified into the third set (benchmark data). The accuracy and effectiveness of the screening of historical audit data are ensured, so that risk assessment can draw on the historical audit experience closest to the current sample, and further improve the stability and consistency of the audit.

[0092] In some embodiments of the present application, when the risk factor is adjusted by determining the adjustment coefficient according to the first set, the second set, and the third set, the method includes: obtaining the ratio of each historical final risk coefficient to the risk coefficient in the set, and obtaining the average ratio according to the ratio of each historical final risk coefficient to the risk coefficient in the third set. Obtain the first variance of all ratios in the first set and the average ratio, and obtain the second variance of all ratios in the second set and the average ratio, sum the first variance and the second variance and then take the square root to obtain the sum of the variances, sum one-half of the sum of the variances with the average ratio to obtain the adjustment coefficient, and use the product of the adjustment coefficient and the risk coefficient as the final risk coefficient.

[0093] It is understandable that by combining similarity analysis with historical data and statistical calculations, similar historical data can be reasonably screened through cosine similarity, and the adjustment coefficient can be calculated based on the variance and ratio of historical data to make risk assessment more accurate and stable. By dynamically adjusting the risk coefficient, audit errors can be effectively reduced.

[0094] In some embodiments of the present application, the final risk coefficient is compared with the coefficient threshold, and the processing level of the sample to be reviewed is determined according to the comparison result, including:

[0095] The final risk coefficient is compared with the first coefficient threshold and the second coefficient threshold respectively, and the processing level of the sample to be reviewed is determined according to the comparison results. The first coefficient threshold is less than the second coefficient threshold.

[0096] When the final risk coefficient is less than or equal to the first coefficient threshold, the processing level of the sample to be reviewed is determined to be the first level. When the final risk coefficient is greater than the first coefficient threshold and less than or equal to the second coefficient threshold, the processing level of the sample to be reviewed is determined to be the second level. When the final risk coefficient is greater than the second coefficient threshold, the processing level of the sample to be reviewed is determined to be the third level. The first level indicates that the risk level is less than the second level, and the second level indicates that the risk level is less than the third level.

[0097] Specifically, when the final risk coefficient is less than or equal to the first coefficient threshold, it indicates that the sample has a low risk of violation, so it is classified as the first level, that is, the low risk level. When the final risk coefficient is between the first coefficient threshold and the second coefficient threshold, it indicates that the sample has a certain possibility of violation, but has not reached the highest risk level, so it is classified as the second level, that is, the medium risk level. When the final risk coefficient is greater than the second coefficient threshold, it indicates that the sample has a high risk of violation, so it is classified as the third level, that is, the high risk level.

[0098] It is understandable that through the multi-level threshold comparison mechanism based on risk factors, the refined classification of samples to be reviewed is achieved, avoiding the problem of misjudgment of review that may be caused by simple "compliance / violation" binary classification. By setting three risk levels of low, medium and high, different levels of review strategies can be flexibly matched. For example, low-risk samples can be automatically passed, medium-risk samples can be manually reviewed, and high-risk samples can be directly intercepted or strictly reviewed, thereby improving the intelligence, accuracy and operability of the review, and ultimately improving the stability and reliability of content review.

[0099] In the above embodiments, unified audit processing of data in multiple formats such as text, audio and video is realized, and the accuracy of cross-modal violation identification is improved. By identifying the content of the sample to be audited and determining its format, natural language processing, image processing and audio processing and other technical means are used for text audio and text video samples to ensure that data of different modalities can be effectively extracted, and a unified feature data group is constructed to improve data fusion capabilities. Introducing risk coefficients In the audit of text audio and text video, a comprehensive assessment is conducted based on the proportion of violations of audio, video and text, respectively, so as to ensure that the illegal content of different modal data is reasonably quantified and avoid a single modality affecting the overall audit results. Introducing cross-modal matching analysis, by obtaining the location of audio or video content, determining its matching range with the text, and calculating the matching degree, to determine whether the risk coefficient needs to be adjusted, thereby further reducing the misjudgment rate and improving the robustness of the audit. By comparing the feature data group with the historical audit data, constructing a similar set and determining the adjustment coefficient based on the similarity, the risk assessment is made more accurate, and the adaptive ability and intelligent level of the audit are enhanced. By combining the risk coefficient with the threshold comparison, the processing level of the sample to be reviewed is determined, which improves the efficiency and reliability of content review and effectively solves the problems in existing technologies such as the difficulty in comprehensively identifying multimodal violation information, insufficient review accuracy, and unstable risk assessment.

[0100] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0101] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0102] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for implementing multimodal data audit based on machine learning, characterized in that: include: Identify the content of the sample to be reviewed, determine the sample format and determine the data processing method based on the sample format, extract different format features in the sample to be reviewed based on the data processing method, and construct a feature data group; The sample formats include text audio and text video; the data processing methods include natural language processing, image processing and audio processing; Determine the risk factor of the sample to be reviewed based on the feature data group; when the sample format of the sample to be reviewed is text and audio, determine the risk factor based on the audio violation ratio and the text violation ratio; When the sample format of the sample to be reviewed is text video, the risk coefficient is determined according to the video violation ratio and the text violation ratio; Obtaining the location of the audio or video in the sample to be reviewed, determining a matching range according to the location, obtaining a content matching degree between the text in the matching range and the audio or video, and determining whether to adjust the risk factor according to the content matching degree; When it is determined that the risk coefficient is to be adjusted, the characteristic data group is compared with the historical audit data, a similarity set is determined according to the comparison result, and an adjustment coefficient is determined according to the similarity set to adjust the risk coefficient to obtain a final risk coefficient; The final risk coefficient is compared with the coefficient threshold, and the processing level of the sample to be reviewed is determined according to the comparison result.

2. The method for implementing multimodal data audit based on machine learning according to claim 1, characterized in that: Extracting different format features from the sample to be reviewed based on the data processing method and constructing a feature data group includes: The feature data set includes text features, audio features and video features; The natural language processing includes: removing special characters, segmenting the text in the sample to be reviewed based on BERT, extracting keywords and calculating the word frequency and inverse document frequency of the keywords, obtaining the TF-IDF value of each keyword according to the word frequency and inverse document frequency, sorting the keywords according to the TF-IDF values, and obtaining the text features according to the sorting; The audio processing includes: extracting audio text based on speech recognition, obtaining keywords in the audio text, extracting high-dimensional audio features using Wav2Vec and extracting timbre features using Mel frequency cepstral coefficients, and obtaining the audio features according to the keywords, high-dimensional audio features and timbre features; The image processing includes: calculating the motion changes between video frames based on optical flow to obtain a key frame sequence, performing target detection and scene recognition on the key frames respectively, and obtaining target detection features and scene recognition features; extracting video and audio, and obtaining video and audio features based on the video and audio, and obtaining the video features based on the video and audio features, target detection features and scene recognition features.

3. The method for implementing multimodal data audit based on machine learning according to claim 2, characterized in that: When determining the risk factor of the sample to be reviewed according to the characteristic data group, it includes: When determining the risk factor based on the audio violation ratio and the text violation ratio, Compare the text features with the illegal text database to determine the percentage of illegal texts; According to the audio features, the proportion of illegal timbre, the proportion of illegal high-dimensional audio, and the proportion of illegal audio words are obtained, and the illegal audio proportion is determined; Determine the text-audio weight according to the ratio of text to audio, and obtain the risk coefficient according to the text-audio weight; When determining the risk factor based on the video violation ratio and the text violation ratio, Obtaining the target detection violation ratio, the scene recognition violation ratio, and the violation video word ratio according to the video features, and determining the video violation ratio; The text-video weight is determined according to the ratio of text to video, and the risk coefficient is obtained according to the text-video weight.

4. The method for implementing multimodal data audit based on machine learning according to claim 2, characterized in that: When determining the matching range according to the location, it includes: When the audio or video is in the first 30% of the sample to be reviewed, the first 50% of the text in the sample to be reviewed is used as the matching range; When the audio or video is in the first 30-50% of the sample to be reviewed, the first 70% of the text in the sample to be reviewed is used as the matching range; When the audio or video is in the last 50% of the sample to be reviewed, all the text in the sample to be reviewed is used as the matching range.

5. The method for implementing multimodal data audit based on machine learning according to claim 4 is characterized in that: When obtaining the content matching degree between the text in the matching range and the audio or video, it includes: When the sample format of the sample to be reviewed is text audio, the content matching degree is the similarity between the audio text and the text; When the sample format of the sample to be reviewed is text video, video and audio text is extracted according to the video and audio, and the content similarity between the video and audio text and the text is obtained. When the content similarity is not zero, the content similarity is used as the content matching degree; when the content similarity is zero, the feature similarity between the video feature and the text feature is obtained, and the feature similarity is used as the content matching degree. The value range of the content matching degree is [0, 1].

6. The method for implementing multimodal data audit based on machine learning according to claim 5, characterized in that: When judging whether to adjust the risk factor according to the content matching degree, it includes: When the content matching degree is less than or equal to the matching degree threshold, it is determined that the risk coefficient is adjusted; when the content matching degree is greater than the matching degree threshold, it is determined that the risk coefficient is not adjusted and the risk coefficient is used as the final risk coefficient.

7. The method for implementing multimodal data audit based on machine learning according to claim 6, characterized in that: When comparing the characteristic data group with the historical audit data and determining a similar set based on the comparison result, it includes: The historical audit data includes a plurality of historical feature data groups, a plurality of historical content matching degrees, and a plurality of historical final risk coefficients, and each historical feature data group corresponds to a historical content matching degree and a historical final risk coefficient; When there is data in the historical audit data whose historical content matching degree is equal to the content matching degree, the historical feature data group corresponding to the historical content matching degree and the historical final risk coefficient are constructed as a similarity set; When there is no data in the historical review data whose historical content matching degree is equal to the content matching degree, select the data in the historical review data whose absolute value of the difference between the historical content matching degree and the content matching degree is less than or equal to 0.1, and construct the corresponding historical feature data group and the historical final risk coefficient into a similarity set.

8. The method for implementing multimodal data audit based on machine learning according to claim 7, characterized in that: Determining an adjustment coefficient according to the similarity set to adjust the risk coefficient to obtain a final risk coefficient includes: Calculate the correlation between the feature data group and each of the historical feature data groups based on cosine similarity, and obtain the average correlation, classify the data in the similarity set with a correlation greater than the average correlation into a first set, classify the data in the similarity set with a correlation less than the average correlation into a second set, and classify the data in the similarity set with a correlation equal to the average correlation into a third set; The adjustment coefficient is determined according to the first set, the second set and the third set to adjust the risk coefficient to obtain a final risk coefficient.

9. The method for implementing multimodal data audit based on machine learning according to claim 8, characterized in that: When the adjustment coefficient is determined according to the first set, the second set, and the third set to adjust the risk coefficient, it includes: Obtaining a ratio of each historical final risk coefficient in the set to the risk coefficient, and obtaining an average ratio according to the ratio of each historical final risk coefficient in the third set to the risk coefficient; Obtain a first variance of all ratios in the first set and the average ratio, and obtain a second variance of all ratios in the second set and the average ratio, sum the first variance and the second variance and take the square root to obtain a sum of variances, sum half of the sum of variances and the average ratio to obtain the adjustment coefficient, and take the product of the adjustment coefficient and the risk coefficient as the final risk coefficient.

10. The method for implementing multimodal data audit based on machine learning according to claim 9, characterized in that: The final risk coefficient is compared with the coefficient threshold, and the processing level of the sample to be reviewed is determined according to the comparison result, including: The final risk coefficient is compared with the first coefficient threshold and the second coefficient threshold respectively, and the processing level of the sample to be reviewed is determined according to the comparison results; the first coefficient threshold is less than the second coefficient threshold; When the final risk coefficient is less than or equal to the first coefficient threshold, the processing level of the sample to be reviewed is determined to be the first level; when the final risk coefficient is greater than the first coefficient threshold and less than or equal to the second coefficient threshold, the processing level of the sample to be reviewed is determined to be the second level; when the final risk coefficient is greater than the second coefficient threshold, the processing level of the sample to be reviewed is determined to be the third level; and the first level indicates that the risk level is less than the second level, and the second level indicates that the risk level is less than the third level.

Citation Information

Patent Citations

  • Short video auditing method based on multiple modes

    CN115512259A

  • Industrial internet data abnormity detection method, system and equipment

    CN116628554A

  • Sampling content using machine learning to identify low-quality content

    US20170262635A1

Cited By

  • Video content automatic auditing method and system based on AI drive

    CN120726547A