Cloud platform auditing method based on multi-modal data processing
Through the combination of multimodal data processing and deep learning models, the problems of weak cross-modal correlation and high computing resource consumption in cloud platform audits are solved, and efficient and accurate violation detection and interpretable audit results are achieved.
Patent Information
- Application Number
- CN202510288508.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-12
AI Technical Summary
The existing cloud platform audit methods have problems such as weak cross-modal correlation, high computing resource consumption, serious false alarms and missed reports, and lack of interpretability of audit results, which are difficult to meet the needs of efficient and accurate audits.
A cloud platform audit method based on multimodal data processing is adopted to build a deep learning audit model through steps such as data preprocessing, feature extraction, cross-modal feature fusion, violation judgment and interpretability analysis to realize unified representation and violation detection of text, image, audio, video and log data.
It improves the accuracy and efficiency of the audit system, reduces the consumption of computing resources, reduces the false positives and missed reports, and enhances the transparency of audit results through interpretability analysis.
Smart Images

Figure CN120223932A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud platform auditing, and particularly to a cloud platform auditing method based on multi-modal data processing. Background Art
[0002] There are a large number of data storage and computing tasks in the cloud platform, involving various types of data, such as text, images, videos, audio, etc. With the increasing demand for content supervision, the platform needs to audit the stored files, streaming media, and the content uploaded by users for illegal content.
[0003] Most of the current cloud platform auditing methods adopt deep learning combined with a rule engine, such as keyword matching or natural language processing (NLP) detection, and computer vision technology is used to identify copyright infringement and other content; however, the auditing processes of various types of data are independent, lacking correlation analysis, and are prone to misjudgment or missed judgment. For example, the text log shows normal, but the accompanying image or video content is illegal, or the video picture is compliant, but the background audio and voice content is illegal; in addition, the current auditing also has the problem of high consumption of computing resources. In a large-scale cloud platform environment, the computing cost of parallel processing of massive data is very high, resulting in limited auditing efficiency.
[0004] In view of the above problems, some traditional auditing solutions adopt a hierarchical optimization strategy, increase manual auditing, or use a lightweight pre-screening model for rapid filtering, but these methods are difficult to fundamentally improve the auditing efficiency and accuracy. Therefore, there is an urgent need for a cloud platform auditing method based on multi-modal data processing to solve such problems. Summary of the Invention
[0005] In view of the existing problems above, the present invention is proposed.
[0006] The present invention provides a cloud platform auditing method based on multi-modal data processing to solve the problems of weak cross-modal correlation, high consumption of computing resources, serious false positives and false negatives, and lack of interpretability of auditing results in traditional auditing methods, which are difficult to meet the requirements of efficient and accurate auditing.
[0007] To solve the above technical problems, the present invention provides the following technical solutions:
[0008] An embodiment of the present invention provides a cloud platform auditing method based on multi-modal data processing, which includes:
[0009] Step S1, collecting the storage and transmission data of the cloud platform and performing data preprocessing, where the data preprocessing includes format standardization, noise removal, integrity verification, and metadata extraction;
[0010] Step S2, based on the data preprocessed in step S1, respectively extracting features from text, image, audio, video, and log data to obtain features of each modality;
[0011] Step S3: Based on the modal features extracted in Step S2, perform cross-modal feature fusion to obtain unified multi-modal features;
[0012] Step S4: Based on the multi-modal features generated in Step S3, construct an audit model to determine data violations;
[0013] Step S5: Based on the audit results in Step S4, perform interpretability analysis and generate an audit report.
[0014] As a preferred solution of the cloud platform audit method based on multi-modal data processing described in the present invention, wherein: the stored and transmitted data includes text, images, audio, video, and log data.
[0015] As a preferred solution of the cloud platform audit method based on multi-modal data processing described in the present invention, wherein: the feature extraction includes:
[0016] Use a natural language processing model to perform text analysis to identify keywords, named entities, and sentiment tendencies;
[0017] Use a computer vision model to perform image and video analysis to detect target objects, optical characters, and violation elements;
[0018] Use automatic speech recognition technology to transcribe audio and perform semantic, emotional, and timbre analysis;
[0019] Parse the log data to extract cloud platform call, access records, and abnormal traffic characteristics.
[0020] As a preferred solution of the cloud platform audit method based on multi-modal data processing described in the present invention, wherein: the steps of performing feature extraction on text, images, audio, video, and log data are,
[0021] Use natural language processing methods to extract key features from text data and convert the text data into word vectors
[0022]
[0023] Among them, represents the global vector representation of the text at time step t, is the number of text words at time step t, is the weight of the i-th word, is the embedding vector corresponding to this word,
[0024] Calculate the keyword weight based on TFIDF and named entity recognition, and the calculation formula is:
[0025]
[0026] Among them, is the score of the j-th keyword, is the frequency of occurrence of this keyword in the text, and M T is the total number of all texts, is the number of texts containing this keyword,
[0027] The sentiment tendency is calculated using a bidirectional long short-term memory network, expressed as:
[0028]
[0029] Among them, represents the hidden state at time step t, is the hidden state of the previous time step, is the hidden state weight matrix, is the word vector at the current time step t, is the bias term, and σ is the activation function;
[0030] The computer vision model is used to extract image and video features,
[0031] The convolutional neural network is used for object detection, and the detection process is expressed as:
[0032]
[0033] Among them, represents the feature map of the image, is the convolutional kernel, is the input image data, is the bias term, and * represents the convolution operation,
[0034] Based on connectionist temporal classification, calculate:
[0035]
[0036] Among them, is the character probability predicted by OCR, is the set of possible character paths, is the character path at time step t, is the total number of time steps, is the predicted probability of the character corresponding to time step t;
[0037] Automatic speech recognition is used to extract audio features,
[0038] Based on the mel spectrogram coefficients, calculate the speech features, and the calculation formula is:
[0039]
[0040] Among them, is the nth A dimensional MFCC coefficient, is the kth frequency component, representing the signal energy after Mel filtering, is the Fourier transform window size,
[0041] The sentiment tendency is calculated using a deep neural network. The calculation process is as follows:
[0042]
[0043] Among them, represents the hidden state at time step t, is the weight matrix, is the input feature, is the bias term;
[0044] Parse the cloud platform call, access record, and abnormal traffic characteristics,
[0045] Perform access modeling and calculate based on the time series prediction LSTM:
[0046]
[0047] Among them, is the hidden state at the current time step, is the hidden state at the previous time step, is the log event feature, and are the weight matrices, is the bias,
[0048] Calculate the abnormal distribution using the probability density estimation KDE:
[0049]
[0050] Among them, p(x L ) is the probability density of the data point x L , is the number of log data samples, h L is the bandwidth, is the value of sample i, and K(·) is the kernel function.
[0051] As a preferred solution of the cloud platform audit method based on multi-modal data processing described in the present invention, wherein: the cross-modal feature fusion method includes:
[0052] Calculate the correlation between different modal features using a cross-modal consistency contrast network;
[0053] Adopt a feature alignment method based on the attention mechanism to extract cross-modal important features;
[0054] A shared feature encoder is used to uniformly represent different modality features.
[0055] As a preferred solution of the cloud platform audit method based on multi-modal data processing described in the present invention, wherein: the step of performing cross-modal feature fusion is,
[0056] A cross-modal consistency contrast network is used to calculate the similarity of different modality features in the common feature space, and the calculation formula is:
[0057]
[0058] Wherein, represents the Euclidean distance between modality p and modality q, p, q ∈ {T, I, A, V, L} corresponding to the five modalities of text, image, audio, video, and log, is the feature vector extracted for modality p, is the feature vector extracted for modality q, ||·||^2 represents the square of the second norm,
[0059] The cross-modal consistency is defined by a similarity function, and the function formula is:
[0060]
[0061] Wherein, is the similarity between modality p and modality q, is the temperature parameter of modality p,
[0062] An attention mechanism is used to extract cross-modal important features, and the extraction formula is:
[0063]
[0064] Wherein, is the attention weight of modality r, R M is the total number of all modalities, is the attention weight matrix of modality r, is the feature vector extracted for modality r,
[0065] Calculate the cross-modal feature fusion representation, and the calculation formula is:
[0066]
[0067] Wherein, is the cross-modal fusion feature representation;
[0068] The different modality features are unified by the shared feature encoder to obtain the final multi-modal features, which are expressed as:
[0069]
[0070] Among them, is the final multi-modal feature representation, is the shared feature encoding matrix, is the bias term.
[0071] As a preferred solution of the cloud platform audit method based on multi-modal data processing described in the present invention, wherein: the audit methods of the audit model include:
[0072] Combined with a rule engine for rapid content screening;
[0073] Adopt a deep learning model for end-to-end classification and output a violation score;
[0074] Analyze the log behavior data through an anomaly detection model to identify potential anomalies;
[0075] Set a dynamic risk threshold for secondary review of high-risk content.
[0076] As a preferred solution of the cloud platform audit method based on multi-modal data processing described in the present invention, wherein: the steps of constructing an audit model based on the multi-modal features generated in step S3 and performing a violation determination on the data are as follows:
[0077] The audit model calculates the violation score based on deep learning, and the scoring formula is:
[0078]
[0079] Among them, is the violation score, is the output layer weight matrix, is the bias term, and σ is the activation function.
[0080] Define the basis for violation determination:
[0081]
[0082] Among them, is the violation determination result, is the violation determination threshold, I(·) is the indicator function. If then otherwise
[0083] Based on the dynamic threshold adjustment, calculate the anomaly behavior determination threshold, and the calculation formula is:
[0084]
[0085] Among them, is the dynamic risk threshold, is the mean of historical data, is the adjustment coefficient, is the standard deviation of data features,
[0086] Calculate the data abnormality based on the anomaly detection model, and the calculation formula is:
[0087]
[0088] Among them, is the anomaly detection flag, I(·) is the indicator function, when at this time otherwise
[0089] As a preferred solution of the cloud platform audit method based on multi-modal data processing described in the present invention, wherein: the audit report includes:
[0090] Adopt visualization technology to mark suspicious areas in images and videos;
[0091] Combine the language model to generate audit conclusions and violation evidence;
[0092] Add time indexes to the audit results of audio and video.
[0093] As a preferred solution of the cloud platform audit method based on multi-modal data processing described in the present invention, wherein: the steps of performing interpretability analysis based on the audit results of step S4 and generating an audit report are,
[0094] For images and videos, mark suspicious areas, and the marking method is:
[0095]
[0096] Among them, is the set of suspicious area coordinates, is the pixel coordinate in the image or video frame, is the pixel violation probability, is the risk threshold of the suspicious area,
[0097] Use the natural language generation model to generate audit conclusions
[0098]
[0099] Among them, is the audit report text, is based on multi-modal features generated language description function,
[0100] Add an audit time index for audio and video:
[0101]
[0102] Among them, is the set of time indexes in the audit report, is the audio violation timestamp, is the video violation timestamp, and are time steps and are the violation determination results.
[0103] The beneficial effects of the present invention are as follows: In the present invention, independent feature extraction of text, images, audio, video, and log data enables each modality information to be refined and analyzed in the corresponding feature space. In the cross-modal feature fusion stage, the cross-modal consistency contrast network CMCN is used to calculate the correlation between different modality data, and the feature alignment is optimized through the attention mechanism, so that data such as text, images, audio, and video can be represented in the same feature space, and multi-modal information can be mutually verified, reducing misjudgments and missed judgments caused by independent audits.
[0104] In the present invention, an audit model based on deep learning is constructed in the audit decision-making link, and combined with anomaly detection and adaptive risk control thresholds, end-to-end determination of illegal content is realized, improving the intelligence level of the audit system; in addition, by visualizing and annotating suspicious areas of images and videos, using the natural language generation NLG model to generate audit conclusions, and adding a time index to the audio and video audit results, the audit results are made more transparent, facilitating manual review and adjustment of audit strategies.
[0105] In summary, compared with traditional methods, the present invention reduces the consumption of computing resources, improves the audit efficiency, and reduces the false alarm rate and missed report rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0106] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0107] Figure 1 is a schematic flow chart of the cloud platform audit method based on multi-modal data processing of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0108] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention in conjunction with the drawings of the specification.
[0109] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways than those specifically described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0110] Secondly, as used herein, "one embodiment" or "an embodiment" refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other with other embodiments.
[0111] Example 1, referring to Figure 1 , this embodiment provides a cloud platform auditing method based on multi-modal data processing, including the following steps:
[0112] Step S1, collect the storage and transmission data of the cloud platform and perform data preprocessing. The data preprocessing includes format standardization, noise removal, integrity verification, and metadata extraction;
[0113] The storage and transmission data includes text, images, audio, video, and log data;
[0114] Step S2, based on the data preprocessed in Step S1, perform feature extraction on the text, image, audio, video, and log data respectively to obtain features of each modality;
[0115] The feature extraction includes:
[0116] Use a natural language processing model to perform text analysis to identify keywords, named entities, and sentiment tendencies;
[0117] Use a computer vision model to perform image and video analysis to detect target objects, optical characters, and violation elements;
[0118] Use automatic speech recognition technology to transcribe audio and perform semantic, emotion, and tone analysis;
[0119] Parse the log data to extract cloud platform call, access records, and abnormal traffic features;
[0120] The steps of performing feature extraction on the text, image, audio, video, and log data are,
[0121] Use natural language processing methods to extract key features from text data and convert the text data into word vectors
[0122]
[0123] Among them, represents the global vector representation of the text at time step t, is the number of text words at time step t, is the weight of the i-th word, is the embedding vector corresponding to this word,
[0124] Calculate the keyword weights based on TFIDF and named entity recognition, and the calculation formula is:
[0125]
[0126] Among them, is the score of the j-th keyword, is the frequency of occurrence of this keyword in the text, M T is the total number of all texts, is the number of texts containing this keyword,
[0127] Use a bidirectional long short-term memory network to calculate the sentiment tendency, expressed as:
[0128]
[0129] Among them, represents the hidden state at time step t, is the hidden state of the previous time step, is the hidden state weight matrix, is the word vector at the current time step t, is the bias term, and σ is the activation function;
[0130] Use a computer vision model to extract image and video features,
[0131] Use a convolutional neural network for object detection, and the detection process is expressed as:
[0132]
[0133] Among them, represents the feature map of the image, is the convolutional kernel, is the input image data, is the bias term, and * represents the convolution operation,
[0134] Calculate based on connectionist temporal classification:
[0135]
[0136] Among them, is the character probability predicted by OCR, is the set of possible character paths, is the character path at time step t, is the total number of time steps, is the predicted probability of the character corresponding to time step t;
[0137] Extract audio features using automatic speech recognition,
[0138] Calculate speech features based on Mel spectrum coefficients, and the calculation formula is:
[0139]
[0140] where, is the n A -dimensional MFCC coefficient, is the kth frequency component, representing the signal energy after Mel filtering, is the Fourier transform window size,
[0141] Calculate the emotional tendency using a deep neural network, and the calculation process is:
[0142]
[0143] where, represents the hidden state at time step t, is the weight matrix, is the input feature, is the bias term;
[0144] Parse the cloud platform calls, access records, and abnormal traffic features,
[0145] Perform access modeling and calculate based on the time series prediction LSTM:
[0146]
[0147] where, is the hidden state at the current time step, is the hidden state at the previous time step, is the log event feature, and are the weight matrices, is the bias,
[0148] Calculate the abnormal distribution using the probability density estimation KDE:
[0149]
[0150] where p(x L ) is the probability density of the data point x L , is the number of log data samples, and h L is the bandwidth, is the value of sample i, and K(·) is the kernel function,
[0151] Specifically, in step S2, the feature calculations of different modalities are refined. For the text modality, word vector embedding, TFIDF, and BiLSTM are used to extract sentiment and semantic features. For the image and video modalities, CNN and OCR are used to identify target and text information. For the audio modality, MFCC is used for speech feature extraction. For the log modality, LSTM is used to analyze time series patterns to avoid cross-modal confusion;
[0152] In step S3, based on the features of each modality extracted in step S2, cross-modal feature fusion is performed to obtain unified multi-modal features;
[0153] The cross-modal feature fusion method includes:
[0154] Using a cross-modal consistency contrast network to calculate the correlation between different modality features;
[0155] Using a feature alignment method based on the attention mechanism to extract cross-modal important features;
[0156] Using a shared feature encoder to perform unified representation of different modality features;
[0157] The steps for performing cross-modal feature fusion are,
[0158] Using a cross-modal consistency contrast network to calculate the similarity between different modality features in the common feature space. The calculation formula is:
[0159]
[0160] where, represents the Euclidean distance between modality p and modality q, and p, q ∈ {T, I, A, V, L} correspond to the five modalities of text, image, audio, video, and log, is the feature vector extracted for modality p, is the feature vector extracted for modality q, and ||·||^2 represents the square of the second norm,
[0161] Define cross-modal consistency using a similarity function. The function formula is:
[0162]
[0163] where, is the similarity between modality p and modality q, is the temperature parameter of modality p,
[0164] Using the attention mechanism to extract cross-modal important features. The extraction formula is:
[0165]
[0166] in, is the attention weight of modality r, R M is the total number of all modes, is the attention weight matrix of modality r, is the eigenvector extracted from mode r,
[0167] Calculate the cross-modal feature fusion representation, the calculation formula is:
[0168]
[0169] in, It is a cross-modal fusion feature representation;
[0170] The shared feature encoder unifies the features of different modalities to obtain the final multimodal features, which can be expressed as:
[0171]
[0172] in, is the final multimodal feature representation, is the shared feature encoding matrix, is the bias term;
[0173] Specifically, in step S3, through cross-modal consistency calculation, attention mechanism alignment and shared feature encoding, the features of different modalities are fused, and the information of each modality can be represented in the same feature space:
[0174] The correlation between modalities is calculated through Euclidean distance, and the importance weights between modalities are extracted using the attention mechanism. Finally, a shared feature encoder is used to unify the representation of different modalities.
[0175] Step S4, based on the multimodal features generated in step S3, construct an audit model to determine the violation of the data;
[0176] The audit methods of the audit model include:
[0177] Combined with the rule engine for rapid content screening;
[0178] Use deep learning models for end-to-end classification and output violation scores;
[0179] Analyze log behavior data through anomaly detection models to identify potential anomalies;
[0180] Set dynamic risk thresholds and conduct secondary review of high-risk content;
[0181] The steps of constructing an audit model based on the multi-modal features generated in step S3 and determining violations of the data are as follows:
[0182] The audit model calculates the violation score based on deep learning, and the scoring formula is:
[0183]
[0184] Where: is the violation score, is the weight matrix of the output layer, is the bias term, and σ is the activation function.
[0185] Define the basis for violation determination:
[0186]
[0187] Where: is the violation determination result, is the violation determination threshold, I(·) is the indicator function. If then Otherwise
[0188] Calculate the abnormal behavior determination threshold based on dynamic threshold adjustment, and the calculation formula is:
[0189]
[0190] Where: is the dynamic risk threshold, is the mean of historical data, is the adjustment coefficient, is the standard deviation of data features.
[0191] Calculate the data abnormality based on the anomaly detection model, and the calculation formula is:
[0192]
[0193] Where: is the anomaly detection flag, I(·) is the indicator function. When then Otherwise
[0194] Specifically, in step S4, based on the multi-modal features generated in step S3, a deep learning audit model is constructed to determine data violations and identify abnormal behaviors. The audit model uses a multi-layer neural network to calculate the violation score and combines the dynamic threshold adjustment algorithm for anomaly detection. In addition, a violation determination mechanism is set up to enable the model to accurately identify the risk level of content violations.
[0195] Step S5: Based on the review results of Step S4, perform interpretability analysis and generate a review report;
[0196] The review report includes:
[0197] Using visualization technology, mark suspicious areas in images and videos;
[0198] Combine language models to generate review conclusions and evidence of violations;
[0199] Add time indices to the review results of audio and videos;
[0200] The steps of performing interpretability analysis based on the review results of Step S4 and generating a review report are as follows:
[0201] For images and videos, mark suspicious areas, and the marking method is:
[0202]
[0203] where is the set of coordinates of suspicious areas, is the pixel coordinates in the image or video frame, is the pixel violation probability, is the risk threshold of the suspicious area,
[0204] Use a natural language generation model to generate review conclusions
[0205]
[0206] where is the review report text, is based on multimodal features is the generated language description function,
[0207] For audio and videos, add review time indices:
[0208]
[0209] where is the set of time indices in the review report, is the audio violation timestamp, is the video violation timestamp, and is the time step and is the violation determination result of
[0210] Specifically, here the output of the review model is converted into an interpretable review report for manual review;
[0211] In terms of visual analysis, suspicious areas are marked for the illegal content in images and videos, and high-risk areas are determined through pixel-level probability mapping.
[0212] In terms of text review, a natural language generation model is used to generate review conclusions.
[0213] In terms of the review of audio and video data, a time index is added so that reviewers can quickly locate the specific time point when the illegal event occurred.
[0214] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A cloud platform audit method based on multimodal data processing, characterized in that: include, Step S1, collecting storage and transmission data of the cloud platform and performing data preprocessing, wherein the data preprocessing includes format standardization, noise removal, integrity verification and metadata extraction; Step S2, based on the data preprocessed in step S1, feature extraction is performed on the text, image, audio, video and log data to obtain features of each modality; Step S3, based on the features of each modality extracted in step S2, cross-modal feature fusion is performed to obtain a unified multi-modal feature; Step S4, based on the multimodal features generated in step S3, construct an audit model to determine the violation of the data; Step S5, based on the audit result of step S4, perform explainability analysis and generate an audit report.
2. A cloud platform audit method based on multimodal data processing as claimed in claim 1, characterized in that: The stored and transmitted data include text, images, audio, video and log data.
3. A cloud platform audit method based on multimodal data processing as claimed in claim 2, characterized in that: The feature extraction comprises: Use natural language processing models to analyze text and identify keywords, named entities, and sentiment tendencies; Use computer vision models to analyze images and videos to detect target objects, optical characters, and illegal elements; Use automatic speech recognition technology to transcribe the audio and perform semantic, emotional and timbre analysis; Parse log data and extract cloud platform calls, access records, and abnormal traffic characteristics.
4. A cloud platform audit method based on multimodal data processing as claimed in claim 3, characterized in that: The steps of extracting features from text, image, audio, video and log data are as follows: Use natural language processing methods to extract key features from text data and convert text data into word vectors in, represents the global vector representation of the text at time step t, is the number of text words at time step t, is the weight of the i-th word, is the embedding vector corresponding to the word, The keyword weight is calculated based on TFIDF and named entity recognition. The calculation formula is: in, is the score of the jth keyword, is the frequency of occurrence of the keyword in the text, M T is the total number of all texts, is the number of texts containing the keyword, The bidirectional long short-term memory network is used to calculate the emotional tendency, which is expressed as: in, represents the hidden state at time step t, is the hidden state of the previous time step, is the hidden state weight matrix, is the word vector of the current time step t, is the bias term, σ is the activation function; Use computer vision models to extract image and video features, Convolutional neural network is used for target detection, and the detection process is expressed as: in, represents the feature map of the image, is the convolution kernel, For input image data, is the bias term, * represents the convolution operation, Classification calculation based on connection timing: in, is the character probability predicted by OCR, is the set of possible character paths, is the character path at time step t, is the total number of time steps, is the predicted probability of the character corresponding to time step t; Automatic speech recognition is used to extract audio features. The speech features are calculated based on the Mel spectrum coefficients. The calculation formula is: in, For nth A dimensional MFCC coefficients, is the kth frequency component, representing the signal energy after Mel filtering, is the Fourier transform window size, A deep neural network is used to calculate the emotional tendency. The calculation process is as follows: in, represents the hidden state at time step t, is the weight matrix, is the input feature, is the bias term; Analyze cloud platform calls, access records and abnormal traffic characteristics, Perform access modeling and predict LSTM calculations based on time series: in, is the hidden state at the current time step, is the hidden state of the previous time step, is the log event feature, and is the weight matrix, is the bias, Use probability density estimation KDE to calculate abnormal distribution: Among them, p(x L ) is the data point x L The probability density of is the number of log data samples, h L is the bandwidth, is the value of sample i, and K(·) is the kernel function.
5. A cloud platform audit method based on multimodal data processing as claimed in claim 4, characterized in that: The cross-modal feature fusion method includes: A cross-modal consistency comparison network is used to calculate the correlation between features of different modalities; Adopt feature alignment method based on attention mechanism to extract important cross-modal features; A shared feature encoder is used to uniformly represent features of different modalities.
6. A cloud platform audit method based on multimodal data processing as claimed in claim 5, characterized in that: The steps of performing cross-modal feature fusion are: A cross-modal consistency comparison network is used to calculate the similarity of different modal features in the common feature space. The calculation formula is: in, represents the Euclidean distance between modality p and modality q, p,q∈{T,I,A,V,L} corresponds to five modalities: text, image, audio, video and log. is the eigenvector extracted from mode p, is the eigenvector extracted from mode q, ||·||^2 represents the square of the second norm, The cross-modal consistency is defined by the similarity function, and the function formula is: in, is the similarity between mode p and mode q, is the temperature parameter of mode p, The attention mechanism is used to extract important cross-modal features. The extraction formula is: in, is the attention weight of modality r, R M is the total number of all modes, is the attention weight matrix of modality r, is the eigenvector extracted from mode r, Calculate the cross-modal feature fusion representation, the calculation formula is: in, It is a cross-modal fusion feature representation; The shared feature encoder unifies the features of different modalities to obtain the final multimodal features, which can be expressed as: in, is the final multimodal feature representation, is the shared feature encoding matrix, is the bias term.
7. A cloud platform audit method based on multimodal data processing as claimed in claim 6, characterized in that: The audit methods of the audit model include: Combined with the rule engine for rapid content screening; Use deep learning models for end-to-end classification and output violation scores; Analyze log behavior data through anomaly detection models to identify potential anomalies; Set dynamic risk thresholds and conduct secondary review of high-risk content.
8. A cloud platform audit method based on multimodal data processing as claimed in claim 7, characterized in that: The step of constructing an audit model based on the multimodal features generated in step S3 and determining the violation of the data is as follows: The audit model calculates violation scores based on deep learning, and the scoring formula is: in, Score the violation. is the output layer weight matrix, is the bias term, σ is the activation function, Definition of violation criteria: in, The violation determination result is: is the violation judgment threshold, I(·) is the indicator function, if but otherwise The abnormal behavior determination threshold is calculated based on dynamic threshold adjustment. The calculation formula is: in, is the dynamic risk threshold, is the mean of historical data, is the adjustment factor, is the data characteristic standard deviation, The data anomaly is calculated based on the anomaly detection model. The calculation formula is: in, is the abnormal detection mark, I(·) is the indicator function, when hour otherwise 9. A cloud platform audit method based on multimodal data processing as claimed in claim 8, characterized in that: The audit report includes: Use visualization technology to mark suspicious areas in images and videos; Generate audit conclusions and violation evidence in combination with language models; Add time index to audio and video review results.
10. A cloud platform audit method based on multimodal data processing as claimed in claim 9, characterized in that: The steps of performing interpretability analysis based on the audit result of step S4 and generating an audit report are as follows: For images and videos, suspicious areas are marked using the following method: in, is the coordinate set of the suspicious area, are pixel coordinates in an image or video frame, is the pixel violation probability, is the risk threshold of the suspicious area, Generate audit conclusions using natural language generation models in, To review the report text, Based on multimodal features The generated language describes the function, For audio and video, add review time index: in, A collection of time indexes in the audit report. The timestamp for the audio violation, The timestamp of the video violation. and is the time step and Violation determination results.
Citation Information
Patent Citations
Multi-modal sentiment classification method taking text as core
CN113312530A
Multi-mode network content security intelligent auditing system and method thereof
CN118312922A
Rotating apparatus of specimen
KR102347032B1
Cited By
Intelligent green product auditing and verifying method based on cloud platform
CN120875910A
A cloud-based intelligent green product auditing and verification method
CN120875910B
Material auditing and content security filtering method
CN121050984A
Identification method, device and equipment for multi-modal fusion research and judgment of pornographic scene and medium
CN121412730A
Power grid operation and maintenance sensitive operation identification method based on multi-modal fusion and related equipment
CN121502612A