A processing method for multi-modal data review based on machine learning
Through the multimodal data audit method based on machine learning, the problem of poor stability of multimodal data audit in the existing technology is solved, and unified processing and risk assessment of text, audio and video data is realized, improving the accuracy and robustness of the audit.
Patent Information
- Application Number
- CN202510465294.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-15
AI Technical Summary
It is difficult for the existing technology to effectively review multimodal data, especially in content security management. The single-modal audit method is difficult to fully cover the risk assessment of cross-modal data, resulting in poor stability of audit results.
Using machine learning-based processing methods, we can identify the content and format of the samples to be reviewed, extract different modal features, build feature data groups, calculate risk coefficients, and dynamically adjust risk assessment through cross-modal matching analysis and historical audit data comparison.
It realizes unified auditing and processing of multimodal data, improves the accuracy of cross-modal violation identification, reduces the misjudgment rate, and improves the robustness and adaptability of audits.
Smart Images

Figure CN119989242B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular, to a processing method for realizing multi-modal data audit based on machine learning. Background Art
[0002] With the wide application of social media, short video platforms and online communication tools, the integrated content of audio / video and text has shown an explosive growth. The diversification of user-generated content makes it difficult for traditional single-modal audit methods to meet the requirements of complex audit scenarios. Especially in content security management, illegal information may appear in the form of combinations of different modalities, such as the combination of text content and audio / video information to form implicit illegal expressions. Therefore, how to efficiently and accurately audit multi-modal data has become a key direction for the development of content audit technology.
[0003] However, the current single-modal audit methods are difficult to comprehensively cover the risk assessment of cross-modal data. For example, relying solely on text analysis may not be able to identify potential illegal content in audio / video, while relying solely on audio / video analysis may ignore the importance of text information. In the content of the combination of audio / video and text, illegal information is often distributed in different modalities, resulting in difficulty for the audit system to accurately extract and match relevant information, affecting the accuracy of the audit results. Existing risk assessment methods fail to fully consider the relevance among audio, video and text, and are difficult to make a reasonable comprehensive judgment on multi-modal data, resulting in poor stability of the audit results and affecting the effectiveness of content security management.
[0004] Therefore, it is necessary to design a processing method for realizing multi-modal data audit based on machine learning to solve the problems existing in the current technology. Summary of the Invention
[0005] In view of this, the present invention proposes a processing method for realizing multi-modal data audit based on machine learning, aiming to solve the problems of poor stability and low analysis accuracy during the current multi-modal data audit.
[0006] The present invention proposes a processing method for realizing multi-modal data audit based on machine learning, including:
[0007] Identifying the content of the sample to be audited, determining the sample format and determining the data processing method based on the sample format, extracting different format features in the sample to be audited based on the data processing method, and constructing a feature data group; the sample format includes text-audio and text-video; the data processing method includes natural language processing, image processing and audio processing;
[0008] Determine the risk coefficient of the sample to be reviewed according to the characteristic data group; when the sample format of the sample to be reviewed is text-audio, determine the risk coefficient according to the audio violation ratio and the text violation ratio; when the sample format of the sample to be reviewed is text-video, determine the risk coefficient according to the video violation ratio and the text violation ratio;
[0009] Obtain the location of the audio or video in the sample to be reviewed, determine the matching range according to the location, obtain the content matching degree between the text in the matching range and the audio or video, and judge whether to adjust the risk coefficient according to the content matching degree;
[0010] When it is determined to adjust the risk coefficient, compare the characteristic data group with the historical review data, determine the similarity set according to the comparison result, determine the adjustment coefficient according to the similarity set to adjust the risk coefficient, and obtain the final risk coefficient;
[0011] Compare the final risk coefficient with the coefficient threshold, and determine the processing level of the sample to be reviewed according to the comparison result.
[0012] Further, when extracting different format features in the sample to be reviewed based on the data processing method and constructing the characteristic data group, it includes:
[0013] The characteristic data group includes text features, audio features and video features;
[0014] The natural language processing includes: removing special characters, segmenting the text in the sample to be reviewed based on BERT, extracting keywords and calculating the word frequency and inverse document frequency of the keywords, obtaining the TF-IDF value of each keyword according to the word frequency and inverse document frequency, sorting the keywords according to the TF-IDF value, and obtaining the text features according to the sorting;
[0015] The audio processing includes: extracting audio text based on speech recognition, obtaining keywords in the audio text, extracting high-dimensional audio features using Wav2Vec and timbre features using Mel-frequency cepstral coefficients, and obtaining the audio features according to the keywords, high-dimensional audio features and timbre features;
[0016] The image processing includes: calculating the motion change between video frames based on optical flow to obtain a key frame sequence, performing object detection and scene recognition on the key frames respectively to obtain object detection features and scene recognition features; extracting video audio, and obtaining video audio features according to the video audio, and obtaining the video features according to the video audio features, object detection features and scene recognition features.
[0017] Further, when determining the risk coefficient of the sample to be reviewed according to the characteristic data group, it includes:
[0018] When determining the risk coefficient according to the proportion of audio violations and the proportion of text violations,
[0019] Compare the text features with the illegal text database to determine the proportion of text violations;
[0020] Obtain the proportion of illegal voices, the proportion of illegal high-dimensional audio, and the proportion of illegal audio words according to the audio features, and determine the proportion of audio violations;
[0021] Determine the text-audio weight according to the text-audio ratio relationship, and obtain the risk coefficient according to the text-audio weight;
[0022] When determining the risk coefficient according to the proportion of video violations and the proportion of text violations,
[0023] Obtain the proportion of target detection violations, the proportion of scene recognition violations, and the proportion of illegal video words according to the video features, and determine the proportion of video violations;
[0024] Determine the text-video weight according to the text-video ratio relationship, and obtain the risk coefficient according to the text-video weight.
[0025] Further, when determining the matching range according to the location, it includes:
[0026] When the audio or video is in the first 30% of the sample to be reviewed, take the first 50% of the text in the sample to be reviewed as the matching range;
[0027] When the audio or video is in the 30%-50% of the sample to be reviewed, take the first 70% of the text in the sample to be reviewed as the matching range;
[0028] When the audio or video is in the last 50% of the sample to be reviewed, take all the text in the sample to be reviewed as the matching range.
[0029] Further, when obtaining the content matching degree between the text in the matching range and the audio or video, it includes:
[0030] When the sample format of the sample to be reviewed is text-audio, the content matching degree is the similarity degree of the content between the audio text and the text;
[0031] When the sample format of the sample to be audited is a text video, extract the video audio text from the video audio, obtain the content similarity between the video audio text and the text, and when the content similarity is not zero, use the content similarity as the content matching degree; when the content similarity is zero, obtain the feature similarity between the video features and the text features, and use the feature similarity as the content matching degree, and the value range of the content matching degree is [0, 1].
[0032] Further, when judging whether to adjust the risk coefficient according to the content matching degree, it includes:
[0033] When the content matching degree is less than or equal to the matching degree threshold, it is determined that the risk coefficient is adjusted; when the content matching degree is greater than the matching degree threshold, it is determined that the risk coefficient is not adjusted, and the risk coefficient is used as the final risk coefficient.
[0034] Further, when comparing the feature data group with the historical audit data and determining the similarity set according to the comparison result, it includes:
[0035] The historical audit data includes several historical feature data groups, several historical content matching degrees, and several historical final risk coefficients, and each historical feature data group corresponds to a historical content matching degree and a historical final risk coefficient;
[0036] When there is data in the historical audit data with a historical content matching degree equal to the content matching degree, construct a similarity set with the historical feature data group and the historical final risk coefficient corresponding to the historical content matching degree;
[0037] When there is no data in the historical audit data with a historical content matching degree equal to the content matching degree, select the data in the historical audit data with the absolute value of the difference between the historical content matching degree and the content matching degree less than or equal to 0.1, and construct a similarity set with the corresponding historical feature data group and the historical final risk coefficient.
[0038] Further, when determining the adjustment coefficient according to the similarity set to adjust the risk coefficient and obtaining the final risk coefficient, it includes:
[0039] Calculate the correlation between the feature data group and each historical feature data group based on the cosine similarity, and obtain the average correlation. Classify the data in the similarity set with a correlation greater than the average correlation into the first set, classify the data in the similarity set with a correlation less than the average correlation into the second set, and classify the data in the similarity set with a correlation equal to the average correlation into the third set.
[0040] Determine the adjustment coefficient according to the first set, the second set and the third set, and adjust the risk coefficient to obtain the final risk coefficient.
[0041] Further, when determining the adjustment coefficient according to the first set, the second set and the third set to adjust the risk coefficient, it includes:
[0042] Obtain the ratio of each historical final risk coefficient in the set to the risk coefficient, and obtain the average ratio according to the ratio of each historical final risk coefficient in the third set to the risk coefficient;
[0043] Obtain the first variance of all ratios in the first set and the average ratio, and obtain the second variance of all ratios in the second set and the average ratio. Square root the sum of the first variance and the second variance to obtain the variance sum, and sum the half of the variance sum and the average ratio to obtain the adjustment coefficient. Multiply the adjustment coefficient by the risk coefficient as the final risk coefficient.
[0044] Further, when comparing the final risk coefficient with the coefficient threshold and determining the processing level of the sample to be audited according to the comparison result, it includes:
[0045] Compare the final risk coefficient with the first coefficient threshold and the second coefficient threshold respectively, and determine the processing level of the sample to be audited according to the comparison result; the first coefficient threshold is less than the second coefficient threshold;
[0046] When the final risk coefficient is less than or equal to the first coefficient threshold, determine that the processing level of the sample to be audited is the first level; when the final risk coefficient is greater than the first coefficient threshold and less than or equal to the second coefficient threshold, determine that the processing level of the sample to be audited is the second level; when the final risk coefficient is greater than the second coefficient threshold, determine that the processing level of the sample to be audited is the third level; and the first level indicates that the risk degree is less than the second level, and the second level indicates that the risk degree is less than the third level.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows: It realizes the unified review and processing of various formats of data such as text, audio, and video, and improves the accuracy of cross-modal violation recognition. By identifying the content of the sample to be reviewed and determining its format, natural language processing, image processing, audio processing and other technical means are adopted for text-audio and text-video samples to ensure that effective feature extraction can be carried out for data in different modalities, and a unified feature data group is constructed to improve the data fusion ability. The risk coefficient is introduced. In the review of text-audio and text-video, comprehensive evaluation is carried out respectively based on the violation ratios of audio, video and text, so as to ensure that the violation content of data in different modalities is reasonably quantified, and avoid the influence of a single modality on the overall review result. The cross-modal matching degree analysis is introduced. By obtaining the location of the audio or video content, the matching range with the text is determined, and the matching degree is calculated to judge whether the risk coefficient needs to be adjusted, so as to further reduce the misjudgment rate and improve the robustness of the review. By comparing the feature data group with the historical review data, a similarity set is constructed and the adjustment coefficient is determined based on the similarity, making the risk assessment more accurate, enhancing the adaptability and intelligent level of the review. Combining the risk coefficient with the threshold comparison to determine the processing level of the sample to be reviewed, improving the efficiency and reliability of content review, and effectively solving the problems in the prior art such as the difficulty in comprehensively identifying multi-modal violation information, insufficient review accuracy, and unstable risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0049] Figure 1 It is a flowchart of a processing method for realizing multi-modal data review based on machine learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] Hereinafter, the exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. Hereinafter, the present invention will be described in detail with reference to the drawings and in combination with the embodiments.
[0051] In some embodiments of the present application, refer to Figure 1As shown, a processing method for multi-modal data review based on machine learning includes:
[0052] S100: Identify the content of the sample to be reviewed, determine the sample format, and based on the sample format, determine the data processing method. Extract different format features from the sample to be reviewed based on the data processing method, and construct a feature data group. The sample format includes text-audio and text-video. The data processing methods include natural language processing, image processing, and audio processing.
[0053] S200: Determine the risk coefficient of the sample to be reviewed according to the feature data group. When the sample format of the sample to be reviewed is text-audio, determine the risk coefficient according to the audio violation ratio and the text violation ratio. When the sample format of the sample to be reviewed is text-video, determine the risk coefficient according to the video violation ratio and the text violation ratio.
[0054] S300: Obtain the location of the audio or video in the sample to be reviewed, determine the matching range according to the location, obtain the content matching degree between the text and the audio or video within the matching range, and determine whether to adjust the risk coefficient according to the content matching degree.
[0055] S400: When it is determined to adjust the risk coefficient, compare the feature data group with the historical review data, determine the similarity set according to the comparison result, and determine the adjustment coefficient according to the similarity set to adjust the risk coefficient to obtain the final risk coefficient.
[0056] S500: Compare the final risk coefficient with the coefficient threshold, and determine the processing level of the sample to be reviewed according to the comparison result.
[0057] Specifically, in step S100, the content of the sample to be reviewed is identified to determine its data format, and an appropriate data processing method is selected according to this format to extract feature information of different modalities, and finally a feature data group is formed. It supports various data formats such as text, audio, and video, and natural language processing, image processing, and audio processing are respectively used to achieve effective parsing and information extraction of cross-modal data. In step S200, the risk coefficient is calculated based on the constructed feature data group, and different evaluation strategies are adopted in combination with the characteristics of the data format. When the sample is text-audio, calculate the proportion of audio violations and the proportion of text violations, and determine the overall risk coefficient according to the ratio of the two. When the sample is text-video, risk assessment is carried out based on the proportion of video violations and the proportion of text violations. It ensures that different modality data can be reasonably quantified and avoids the influence of a single modality on the overall review result. In step S300, in order to further improve the accuracy of the review, the specific location of the audio or video content in the sample is obtained, and the matching range is determined based on this location. Within the matching range, calculate the matching degree between the text and the audio or video content, and judge the consistency of its semantics and content to analyze whether there are implicit violation expressions between the text and the audio / video. For example, if the text content and the audio content do not match, but each has a separate risk of violation, the risk assessment can be judged whether to be adjusted through the matching degree calculation to reduce misjudgment. If it is determined that the risk coefficient needs to be adjusted, then enter step S400. If it is determined that the risk coefficient does not need to be adjusted, then enter step S500. In step S400, the extracted feature data group is compared with the historical review data to find similar historical cases, and a similar set is constructed based on the comparison result. Calculate the adjustment coefficient based on the similar set and dynamically adjust the risk coefficient of the current sample to make the risk assessment more in line with historical review experience and improve the accuracy and self-adaptability of the risk assessment. In step S500, the finally calculated risk coefficient is compared with the preset risk coefficient threshold, and the processing level of the sample to be reviewed is determined according to the comparison result. If the risk coefficient is higher than the threshold, directly mark the sample as high risk and trigger further manual review or automatic interception. If the risk coefficient is lower than the threshold, it can be determined that the sample risk is relatively low, reducing unnecessary intervention, thereby improving the review efficiency.
[0058] It can be understood that through the multi-modal data fusion technology, in-depth analysis of text, audio, and video data is achieved, and the calculation of matching degree and comparison with historical data are introduced in the risk assessment process to improve the accuracy and stability of the review. Compared with the traditional single-modal review method, it can effectively identify cross-modal violation information, reduce the misjudgment rate, and at the same time, through the construction of similar sets and the application of adjustment coefficients, the review strategy can be dynamically optimized to enhance the intelligence level of the review decision-making. The distributed risk assessment method ensures the comprehensive processing ability of different modal data while guaranteeing efficient review, making the review more flexible and adaptable, and capable of meeting the complex and changeable content review requirements.
[0059] In some embodiments of the present application, when constructing a feature data group by extracting different format features from a sample to be reviewed based on the data processing method, it includes: the feature data group includes text features, audio features, and video features.
[0060] Natural language processing includes: removing special characters, segmenting the text in the sample to be reviewed based on BERT, extracting keywords and calculating the word frequency and inverse document frequency of the keywords, obtaining the TF-IDF value of each keyword according to the word frequency and inverse document frequency, sorting the keywords according to the TF-IDF value, and obtaining text features according to the sorting.
[0061] Audio processing includes: extracting audio text based on speech recognition, obtaining keywords in the audio text, extracting high-dimensional audio features using Wav2Vec and timbre features using Mel-frequency cepstral coefficients, and obtaining audio features according to the keywords, high-dimensional audio features, and timbre features.
[0062] Image processing includes: calculating the motion change between video frames based on optical flow to obtain a sequence of key frames, performing object detection and scene recognition on the key frames respectively to obtain object detection features and scene recognition features. Extracting video audio, and obtaining video audio features according to the video audio, and obtaining video features according to the video audio features, object detection features, and scene recognition features.
[0063] Specifically, in the process of processing text data, natural language processing methods are adopted. First, special characters are removed to ensure the purity and readability of the text data. Subsequently, the BERT (Bidirectional Encoder Representations from Transformers) model is used to tokenize the text to obtain context-related information and improve text understanding ability. To extract key semantic information, the keywords in the text are calculated, and the importance of each keyword is calculated based on the TF-IDF (Term Frequency-Inverse Document Frequency) method. Keywords with higher TF-IDF values represent the core content of the text. By sorting the keywords, the core semantic information of the text is effectively extracted and used as text features.
[0064] Specifically, in terms of audio data processing, the Automatic Speech Recognition (ASR) technology is used to convert audio content into text, and keywords are extracted from the audio text to ensure the full utilization of the semantic information of the audio content. At the same time, to obtain richer audio features, Wav2Vec is used for high-dimensional audio feature extraction. This model is based on a self-supervised learning method and can effectively extract deep features in audio signals to improve the accuracy of audio recognition. In addition, the Mel-Frequency Cepstral Coefficients (MFCC) are used to extract timbre features. MFCC can effectively capture the spectral characteristics of audio and can identify different speech or sound effects. Combining keywords, high-dimensional audio features, and timbre features constitutes complete audio feature data, providing richer information for multi-modal risk assessment.
[0065] Specifically, for video data, optical flow is used to calculate the motion changes between video frames to obtain the dynamic features of video content and extract a sequence of key frames, thereby reducing the computational complexity of data processing while retaining the core information of the video. Based on the key frames, object detection and scene recognition are performed. Object detection is used to identify specific objects or people in the video, while scene recognition is used to understand the environment and context of the video. In addition, the audio in the video is extracted, and the audio content is analyzed through video audio feature extraction methods to ensure the co-processing ability of audio-visual data. Finally, video audio features, object detection features, and scene recognition features are fused to generate complete video feature data.
[0066] It is understandable that through the combination of natural language processing, audio processing and video processing technologies, efficient feature extraction of multimodal data is achieved, providing accurate and reliable feature data for multimodal risk assessment. Compared with traditional audit methods, BERT semantic analysis, TF-ID keyword screening, Wav2Vec deep audio feature extraction, MFCC timbre analysis, optical flow calculation, target detection and scene recognition technologies are used to ensure that key information of different modal data can be accurately extracted, and the intelligent level of audit is improved through cross-modal feature fusion. Especially in complex scenarios where text, audio and video are combined, it can effectively reduce misjudgments and missed judgments in the audit process, improve the ability to identify hidden illegal content, and at the same time reduce unnecessary computing resource consumption and improve audit efficiency.
[0067] In some embodiments of the present application, when determining the risk factor of the sample to be reviewed according to the feature data group, it includes:
[0068] When determining the risk factor based on the audio violation ratio and text violation ratio,
[0069] Compare the text features with the illegal text database to determine the percentage of text violations.
[0070] According to the audio features, the proportion of illegal timbre, the proportion of illegal high-dimensional audio and the proportion of illegal audio words are obtained, and the proportion of audio violations is determined.
[0071] The text-audio weight is determined according to the ratio of text to audio, and the risk coefficient is obtained according to the text-audio weight.
[0072] When determining the risk factor based on the video violation ratio and text violation ratio,
[0073] According to the video features, the target detection violation ratio, scene recognition violation ratio and violation video word ratio are obtained, and the video violation ratio is determined.
[0074] The text-video weight is determined according to the ratio of text to video, and the risk coefficient is obtained according to the text-video weight.
[0075] Specifically, extract the text features of the samples to be reviewed and compare them with the illegal text database. Determine the proportion of illegal information involved in the text content through matching, and then determine the proportion of text violations. Proportion of illegal tones: Use Mel Frequency Cepstral Coefficients (MFCC) and audio fingerprint technology to detect whether there are illegal tones, such as counterfeit voices, abnormal tones, etc. Proportion of illegal high-dimensional audio: Adopt deep learning methods such as Wav2Vec to analyze the high-dimensional features of the audio and identify potential illegal content based on known illegal audio patterns. Proportion of illegal audio words: Transcribe the audio text based on Automatic Speech Recognition (ASR) and match it with the illegal text database to judge the degree of illegality of the audio content. Combine the above three proportions to comprehensively calculate the proportion of audio violations. Since the importance of text and audio may be different in different contents, for example, audio is more important in commentary audio content, while text dominates in subtitle interpretation content, the text-audio weight is calculated through the proportional relationship between text and audio. According to the text-audio weight, the proportion of text violations and the proportion of audio violations are weighted and calculated to obtain the final risk coefficient.
[0076] Specifically, for text-audio data, use the same method to extract the proportion of text violations, and then calculate the proportion of video violations. Based on Object Detection, analyze the objects in the video, such as whether it contains sensitive people, prohibited items, etc., and calculate the proportion of illegal objects. Use Scene Recognition to analyze the overall environment of the video, such as whether it involves violent scenes, illegal places, etc., and calculate the proportion of illegal scenes. Extract the content of video subtitles, embedded text, etc. and compare them with the illegal text database to calculate the proportion of illegal text in the video. Since the main information sources of different video contents may be different (for example, in news reports, text subtitles may dominate, while unsubtitled short videos rely on visual content), the text-video weight is calculated through the proportional relationship between text and video. According to the text-video weight, the proportion of text violations and the proportion of video violations are weighted and calculated to obtain the final risk coefficient.
[0077] It can be understood that through the calculation of the proportion of violations and the calculation of adaptive weights in the three major modalities of text, audio, and video, accurate and efficient risk assessment is achieved. Compared with traditional independent modality analysis methods, it can comprehensively consider various data characteristics to ensure the comprehensiveness and accuracy of the review results. In particular, the combination of tone, speech text, high-dimensional audio feature analysis, and object detection, scene recognition, and video text recognition can effectively capture hidden illegal content and reduce the false judgment and missed judgment rates. Through the calculation of text-audio weight and text-video weight, the adaptability to different content types is improved.
[0078] In some embodiments of the present application, when determining the matching range according to the location, it includes:
[0079] When the audio or video is in the first 30% of the sample to be audited, the first 50% of the text in the sample to be audited is taken as the matching range.
[0080] When the audio or video is in the 30%-50% range of the sample to be audited, the first 70% of the text in the sample to be audited is taken as the matching range.
[0081] When the audio or video is in the last 50% of the sample to be audited, all the text in the sample to be audited is taken as the matching range.
[0082] It can be understood that when the audio and video appear at the beginning of the content, they are highly correlated with the first half of the text. However, the subsequent text may involve different topics, and an overly large matching range may introduce noise, thereby reducing the accuracy of the matching. When the audio and video content appears in the 30%-50% range, the matching range is expanded, and the first 70% of the text content is selected for matching. At this time, the audio and video may be in the transition part and involve the content before and after. Therefore, a larger range of text needs to be covered to ensure the matching accuracy of the content. When the audio and video content is in the last 50%, the entire text content is selected as the matching range. The audio and video appear in the second half, which may be related to the full text information. Especially in summary content, the audio and video may be a supplement to the full text. Therefore, all text information needs to be included in the matching range to prevent missing key information. Through the dynamic adjustment of the matching range by position perception, the problem of unreasonable matching range between text and audio / video in traditional multimodal audit methods is effectively solved. Compared with the method of fixed matching window, it can adaptively adjust the matching range, reduce noise interference, and improve the accuracy and stability of the audit. Especially in the combination of long text and audio / video content, it can reduce the interference of irrelevant content and improve the accuracy of cross-modal audit.
[0083] In some embodiments of the present application, when obtaining the content matching degree between the text within the matching range and the audio or video, it includes: when the sample format of the sample to be audited is text-audio, the content matching degree is the similarity degree between the audio text and the text. When the sample format of the sample to be audited is text-video, the video audio text is extracted from the video audio, and the similarity degree between the video audio text and the text is obtained. When the similarity degree is not zero, the similarity degree is taken as the content matching degree. When the similarity degree is zero, the similarity degree between the video feature and the text feature is obtained, and the similarity degree is taken as the content matching degree. The value range of the content matching degree is [0, 1].
[0084] It can be understood that compared with single text comparison, video feature matching is added. When the matching degree of the audio text is low, it can still provide a matching basis through image object detection and scene recognition, improving the accuracy of the audit. In addition, the value range of the matching degree is set in [0, 1], making the risk assessment more quantifiable and facilitating subsequent decision-making adjustment for intelligent audit.
[0085] In some embodiments of the present application, when determining whether to adjust the risk factor according to the content matching degree, it includes: when the content matching degree is less than or equal to the matching degree threshold, determining to adjust the risk factor. When the content matching degree is greater than the matching degree threshold, determining not to adjust the risk factor, and taking the risk factor as the final risk factor.
[0086] In some embodiments of the present application, the feature data group is compared with the historical audit data, and when the similarity set is determined based on the comparison result, it includes: the historical audit data includes several historical feature data groups, several historical content matching degrees and several historical final risk coefficients, and each historical feature data group corresponds to a historical content matching degree and a historical final risk coefficient.
[0087] Specifically, when there is data in the historical audit data where the historical content matching degree is equal to the content matching degree, the historical feature data group corresponding to the historical content matching degree and the historical final risk coefficient are constructed as a similar set. When there is no data in the historical audit data where the historical content matching degree is equal to the content matching degree, data in the historical audit data where the absolute value of the difference between the historical content matching degree and the content matching degree is less than or equal to 0.1 are selected, and the corresponding historical feature data group and the historical final risk coefficient are constructed as a similar set.
[0088] Specifically, historical audit data comes from the storage and recording of a large number of audited samples, including the corresponding feature data groups, content matching and final risk coefficients. These data accumulate the analysis results of multimodal content such as text, audio, and video in the past audit process, and provide a historical basis for the review of new samples. Specifically, if the content matching degree is low, it means that the correlation between the text content and the audio / video is weak, and there may be hidden cross-modal violations, and the risk coefficient needs to be adjusted to improve the accuracy of risk assessment.
[0089] It is understandable that risk adjustment is controlled by content matching threshold, which avoids unnecessary risk factor adjustment and improves audit efficiency. At the same time, by using the similarity comparison of historical audit data, it can still provide a reasonable reference basis when the data matching degree is low, reducing misjudgment and missed judgment. The intelligent optimization method based on historical data effectively improves the self-learning ability and decision-making stability of the audit, and enhances the accuracy and adaptability of cross-modal data audit.
[0090] In some embodiments of the present application, when determining an adjustment coefficient based on a similarity set to adjust a risk coefficient and obtaining a final risk coefficient, it includes: calculating the correlation between a feature data group and each historical feature data group based on cosine similarity, and obtaining the average correlation. Data in the similarity set with a correlation greater than the average correlation is classified into a first set, data with a correlation less than the average correlation is classified into a second set, and data with a correlation equal to the average correlation is classified into a third set. Determining an adjustment coefficient based on the first set, the second set, and the third set to adjust the risk coefficient and obtaining the final risk coefficient.
[0091] It can be understood that, in order to ensure that the most similar set to the current sample to be audited can be selected from the historical audit data, the cosine similarity calculation method is used to measure the correlation between the feature data group of the current sample and the feature data groups of historical samples, and the average correlation of all historical samples is calculated. By setting a threshold and dividing the sets, historical data with a correlation higher than the average correlation is classified into the first set (highly similar data), data with a correlation lower than the average correlation is classified into the second set (low-correlation data), and data with a correlation equal to the average correlation is classified into the third set (benchmark data). This ensures the accuracy and effectiveness of the screening of historical audit data, enabling risk assessment to draw on the historical audit experience closest to the current sample and further improving the stability and consistency of the audit.
[0092] In some embodiments of the present application, when determining an adjustment coefficient based on the first set, the second set, and the third set to adjust the risk coefficient, it includes: obtaining the ratio of each historical final risk coefficient in the set to the risk coefficient, and obtaining the average ratio based on the ratio of each historical final risk coefficient to the risk coefficient in the third set. Obtaining the first variance between all ratios in the first set and the average ratio, and obtaining the second variance between all ratios in the second set and the average ratio. Taking the square root after summing the first variance and the second variance to obtain the variance sum, and summing half of the variance sum and the average ratio to obtain the adjustment coefficient, and taking the product of the adjustment coefficient and the risk coefficient as the final risk coefficient.
[0093] It can be understood that by combining similarity analysis with historical data and statistical calculations, similar historical data is reasonably screened through cosine similarity, and an adjustment coefficient is calculated based on the variance and ratio of historical data, making the risk assessment more accurate and stable. By dynamically adjusting the risk coefficient, the audit error can be effectively reduced.
[0094] In some embodiments of the present application, when comparing the final risk coefficient with a coefficient threshold and determining the processing level of the sample to be audited according to the comparison result, it includes:
[0095] Compare the final risk coefficient with the first coefficient threshold and the second coefficient threshold respectively, and determine the processing level of the sample to be reviewed according to the comparison results. The first coefficient threshold is less than the second coefficient threshold.
[0096] When the final risk coefficient is less than or equal to the first coefficient threshold, determine that the processing level of the sample to be reviewed is the first level. When the final risk coefficient is greater than the first coefficient threshold and less than or equal to the second coefficient threshold, determine that the processing level of the sample to be reviewed is the second level. When the final risk coefficient is greater than the second coefficient threshold, determine that the processing level of the sample to be reviewed is the third level. And the first level indicates that the risk degree is less than the second level, and the second level indicates that the risk degree is less than the third level.
[0097] Specifically, when the final risk coefficient is less than or equal to the first coefficient threshold, it means that the violation risk of the sample is relatively low, so it is classified as the first level, that is, the low-risk level. When the final risk coefficient is between the first coefficient threshold and the second coefficient threshold, it indicates that there is a certain possibility of violation of the sample, but it has not reached the highest risk level, so it is classified as the second level, that is, the medium-risk level. When the final risk coefficient is greater than the second coefficient threshold, it indicates that the sample has a high violation risk, so it is classified as the third level, that is, the high-risk level.
[0098] It can be understood that through the multi-level threshold comparison mechanism based on the risk coefficient, the refined classification of the samples to be reviewed is realized, avoiding the misjudgment problem in the review that may be caused by the simple binary classification of "compliance / violation". By setting three risk levels of low, medium, and high, different levels of review strategies can be flexibly matched. For example, low-risk samples can pass automatically, medium-risk samples can be manually reviewed, and high-risk samples can be directly intercepted or strictly reviewed, so as to improve the intelligence, accuracy, and operability of the review, and ultimately enhance the stability and reliability of content review.
[0099] In the above embodiments, unified review and processing of data in various formats such as text, audio, and video are achieved, improving the accuracy of cross-modal violation recognition. By identifying the content of the sample to be reviewed and determining its format, natural language processing, image processing, audio processing and other technical means are adopted for text-audio and text-video samples to ensure that effective feature extraction can be performed on data in different modalities, and a unified feature data group is constructed to improve the data fusion ability. The risk coefficient is introduced. In the review of text-audio and text-video, comprehensive evaluation is respectively carried out based on the violation ratios of audio, video and text, so as to ensure that the violation content of data in different modalities is reasonably quantified and avoid the influence of a single modality on the overall review result. The cross-modal matching degree analysis is introduced. By obtaining the location of the audio or video content, its matching range with the text is determined, and the matching degree is calculated to judge whether the risk coefficient needs to be adjusted, so as to further reduce the false positive rate and improve the robustness of the review. By comparing the feature data group with the historical review data, a similarity set is constructed and the adjustment coefficient is determined based on the similarity, making the risk assessment more accurate and enhancing the adaptive ability and intelligent level of the review. Combining the risk coefficient with the threshold comparison to determine the processing level of the sample to be reviewed improves the efficiency and reliability of content review, and effectively solves the problems in the prior art such as difficult to comprehensively identify multi-modal violation information, insufficient review accuracy, and unstable risk assessment.
[0100] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0101] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be realized by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0102] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the function.
[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the function.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for implementing multimodal data audit based on machine learning, characterized in that: include: Identify the content of the sample to be reviewed, determine the sample format and determine the data processing method based on the sample format, extract different format features in the sample to be reviewed based on the data processing method, and construct a feature data group; The sample formats include text audio and text video; the data processing methods include natural language processing, image processing and audio processing; Determine the risk factor of the sample to be reviewed based on the feature data group; when the sample format of the sample to be reviewed is text and audio, determine the risk factor based on the audio violation ratio and the text violation ratio; When the sample format of the sample to be reviewed is text video, the risk coefficient is determined according to the video violation ratio and the text violation ratio; Obtaining the location of the audio or video in the sample to be reviewed, determining a matching range according to the location, obtaining a content matching degree between the text in the matching range and the audio or video, and determining whether to adjust the risk factor according to the content matching degree; When it is determined that the risk coefficient is to be adjusted, the characteristic data group is compared with the historical audit data, a similarity set is determined according to the comparison result, and an adjustment coefficient is determined according to the similarity set to adjust the risk coefficient to obtain a final risk coefficient; The final risk coefficient is compared with the coefficient threshold, and the processing level of the sample to be reviewed is determined according to the comparison result.
2. The method for implementing multimodal data audit based on machine learning according to claim 1, characterized in that: Extracting different format features from the sample to be reviewed based on the data processing method and constructing a feature data group includes: The feature data set includes text features, audio features and video features; The natural language processing includes: removing special characters, segmenting the text in the sample to be reviewed based on BERT, extracting keywords and calculating the word frequency and inverse document frequency of the keywords, obtaining the TF-IDF value of each keyword according to the word frequency and inverse document frequency, sorting the keywords according to the TF-IDF values, and obtaining the text features according to the sorting; The audio processing includes: extracting audio text based on speech recognition, obtaining keywords in the audio text, extracting high-dimensional audio features using Wav2Vec and extracting timbre features using Mel-frequency cepstrum coefficients, and obtaining the audio features according to the keywords, high-dimensional audio features and timbre features; The image processing includes: calculating the motion changes between video frames based on optical flow to obtain a key frame sequence, performing target detection and scene recognition on the key frames respectively, and obtaining target detection features and scene recognition features; extracting video and audio, and obtaining video and audio features based on the video and audio, and obtaining the video features based on the video and audio features, target detection features and scene recognition features.
3. The method for implementing multimodal data audit based on machine learning according to claim 2, characterized in that: When determining the risk factor of the sample to be reviewed according to the characteristic data group, it includes: When determining the risk factor based on the audio violation ratio and the text violation ratio, Compare the text features with the illegal text database to determine the percentage of illegal texts; According to the audio features, the proportion of illegal timbre, the proportion of illegal high-dimensional audio, and the proportion of illegal audio words are obtained, and the illegal audio proportion is determined; Determine the text-audio weight according to the ratio of text to audio, and obtain the risk coefficient according to the text-audio weight; When determining the risk factor based on the video violation ratio and the text violation ratio, Obtaining the target detection violation ratio, the scene recognition violation ratio, and the violation video word ratio according to the video features, and determining the video violation ratio; The text-video weight is determined according to the ratio of text to video, and the risk coefficient is obtained according to the text-video weight.
4. The method for implementing multimodal data audit based on machine learning according to claim 2, characterized in that: When determining the matching range according to the location, it includes: When the audio or video is in the first 30% of the sample to be reviewed, the first 50% of the text in the sample to be reviewed is used as the matching range; When the audio or video is in the first 30-50% of the sample to be reviewed, the first 70% of the text in the sample to be reviewed is used as the matching range; When the audio or video is in the last 50% of the sample to be reviewed, all the text in the sample to be reviewed is used as the matching range.
5. The method for implementing multimodal data audit based on machine learning according to claim 4 is characterized in that: When obtaining the content matching degree between the text in the matching range and the audio or video, it includes: When the sample format of the sample to be reviewed is text audio, the content matching degree is the similarity between the audio text and the text; When the sample format of the sample to be reviewed is text video, video and audio text is extracted according to the video and audio, and the content similarity between the video and audio text and the text is obtained. When the content similarity is not zero, the content similarity is used as the content matching degree; when the content similarity is zero, the feature similarity between the video feature and the text feature is obtained, and the feature similarity is used as the content matching degree. The value range of the content matching degree is [0, 1].
6. The method for implementing multimodal data audit based on machine learning according to claim 5, characterized in that: When judging whether to adjust the risk factor according to the content matching degree, it includes: When the content matching degree is less than or equal to the matching degree threshold, it is determined that the risk coefficient is adjusted; when the content matching degree is greater than the matching degree threshold, it is determined that the risk coefficient is not adjusted and the risk coefficient is used as the final risk coefficient.
7. The method for implementing multimodal data audit based on machine learning according to claim 6, characterized in that: When comparing the characteristic data group with the historical audit data and determining a similar set based on the comparison result, it includes: The historical audit data includes a plurality of historical feature data groups, a plurality of historical content matching degrees, and a plurality of historical final risk coefficients, and each historical feature data group corresponds to a historical content matching degree and a historical final risk coefficient; When there is data in the historical audit data whose historical content matching degree is equal to the content matching degree, the historical feature data group corresponding to the historical content matching degree and the historical final risk coefficient are constructed as a similarity set; When there is no data in the historical review data whose historical content matching degree is equal to the content matching degree, select the data in the historical review data whose absolute value of the difference between the historical content matching degree and the content matching degree is less than or equal to 0.1, and construct the corresponding historical feature data group and the historical final risk coefficient into a similarity set.
8. The method for implementing multimodal data audit based on machine learning according to claim 7, characterized in that: Determining an adjustment coefficient according to the similarity set to adjust the risk coefficient to obtain a final risk coefficient includes: Calculate the correlation between the feature data group and each of the historical feature data groups based on cosine similarity, and obtain the average correlation, classify the data in the similarity set with a correlation greater than the average correlation into a first set, classify the data in the similarity set with a correlation less than the average correlation into a second set, and classify the data in the similarity set with a correlation equal to the average correlation into a third set; The adjustment coefficient is determined according to the first set, the second set and the third set to adjust the risk coefficient to obtain a final risk coefficient.
9. The method for implementing multimodal data audit based on machine learning according to claim 8, characterized in that: When the adjustment coefficient is determined according to the first set, the second set, and the third set to adjust the risk coefficient, it includes: Obtaining a ratio of each historical final risk coefficient in the set to the risk coefficient, and obtaining an average ratio according to the ratio of each historical final risk coefficient in the third set to the risk coefficient; Obtain a first variance of all ratios in the first set and the average ratio, and obtain a second variance of all ratios in the second set and the average ratio, sum the first variance and the second variance and take the square root to obtain a sum of variances, sum half of the sum of variances and the average ratio to obtain the adjustment coefficient, and take the product of the adjustment coefficient and the risk coefficient as the final risk coefficient.
10. The method for implementing multimodal data audit based on machine learning according to claim 9, characterized in that: The final risk coefficient is compared with the coefficient threshold, and the processing level of the sample to be reviewed is determined according to the comparison result, including: The final risk coefficient is compared with the first coefficient threshold and the second coefficient threshold respectively, and the processing level of the sample to be reviewed is determined according to the comparison results; the first coefficient threshold is less than the second coefficient threshold; When the final risk coefficient is less than or equal to the first coefficient threshold, the processing level of the sample to be reviewed is determined to be the first level; when the final risk coefficient is greater than the first coefficient threshold and less than or equal to the second coefficient threshold, the processing level of the sample to be reviewed is determined to be the second level; when the final risk coefficient is greater than the second coefficient threshold, the processing level of the sample to be reviewed is determined to be the third level; and the first level indicates that the risk level is less than the second level, and the second level indicates that the risk level is less than the third level.
Citation Information
Patent Citations
Short video auditing method based on multiple modes
CN115512259A
Industrial internet data abnormity detection method, system and equipment
CN116628554A