Content control method, device and equipment based on deep integration

By processing the data to be detected into multiple dimensions and using a deep ensemble learning model or a student model generated by model distillation for content detection, the problem of accuracy in content control in multi-dimensional data is solved, and efficient control is achieved on resource-limited devices.

CN120632189AActive Publication Date: 2025-09-12HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202511127978.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-12
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing content control technologies are difficult to achieve high accuracy in multi-dimensional data and are difficult to be effectively applied on resource-poor devices.

Method used

By processing the data to be detected into data of multiple different dimensions, using the trained preprocessing model to extract feature vectors, and combining it with the student model generated by the deep ensemble learning model or model distillation to perform content detection and control, the multimodal detection results are integrated to improve accuracy, and the student model is deployed on resource-poor devices to reduce resource requirements.

Benefits of technology

It improves the accuracy of content detection and expands the scope of content control, especially enabling efficient content control on devices with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632189A_ABST
    Figure CN120632189A_ABST
Patent Text Reader

Abstract

The invention provides a content management and control method, device and equipment based on deep integration. The method comprises the following steps: processing to-be-detected data into K data of different dimensions; for the data of any dimension in the K data of different dimensions, utilizing the trained preprocessing model of the dimension to preprocess the data of the dimension to obtain a feature vector of the dimension; and taking the K feature vectors of different dimensions as input of a trained content detection model, determining a content detection result of the to-be-detected data by using the trained content detection model, and performing content management and control processing on the to-be-detected data according to the content detection result. The method can improve the accuracy of content management and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security and data security, and in particular to a content management method, device and equipment based on deep integration. Background Art

[0002] Content control refers to the use of technical means to detect content from multiple dimensions such as images, text, and voice to identify and filter abnormal content and ensure that the content of the platform, application, or website complies with relevant regulations.

[0003] Content control technology is an important tool to ensure the health and security of the Internet environment. How to accurately control content has become a hot research direction. Summary of the Invention

[0004] In view of this, the present application provides a content management method, apparatus and device based on deep integration.

[0005] Specifically, this application is implemented through the following technical solutions: According to a first aspect of an embodiment of the present application, a content management method based on deep integration is provided, comprising: Obtain flow data for asset identification through flow monitoring; Process the data to be detected into data of K different dimensions; where K ≥ 2; For data of any dimension of the K different dimensions, preprocess the data of the dimension using the trained preprocessing model of the dimension to obtain a feature vector of the dimension; The K feature vectors of different dimensions are respectively used as the input of the trained content detection model, and the trained content detection model is used to determine the content detection results of the data to be detected, and content control processing is performed on the data to be detected based on the content detection results; wherein, the content detection model includes a deep ensemble learning model, or a student model generated by the deep ensemble learning model using a model distillation method; the deep ensemble learning model includes a plurality of basic models of different dimensions, and the content detection results of the data to be detected by the deep ensemble learning model are determined based on the detection results of the feature vectors of the corresponding dimensions by the basic models of different dimensions; for non-first basic models in the deep ensemble learning model, the input of the basic model includes the feature vectors of the corresponding dimensions, and the output of the feature output layer of the previous basic model.

[0006] According to a second aspect of an embodiment of the present application, a content management device based on deep integration is provided, including: The first processing unit is configured to process the data to be detected into data of K different dimensions, wherein K is greater than or equal to 2; A second processing unit is configured to preprocess the data of any dimension of the K different dimensional data using a trained preprocessing model for the dimension to obtain a feature vector for the dimension; A content control unit is used to use the K feature vectors of different dimensions as the input of a trained content detection model, use the trained content detection model to determine the content detection results of the data to be detected, and perform content control processing on the data to be detected based on the content detection results; wherein, the content detection model includes a deep integrated learning model, or a student model generated by the deep integrated learning model using a model distillation method; the deep integrated learning model includes a plurality of basic models of different dimensions, and the content detection results of the data to be detected by the deep integrated learning model are determined based on the detection results of the feature vectors of the corresponding dimensions by the basic models of different dimensions; for non-first basic models in the deep integrated learning model, the input of the basic model includes the feature vectors of the corresponding dimensions, and the output of the feature output layer of the previous basic model.

[0007] According to a third aspect of an embodiment of the present application, there is provided an electronic device, including a processor and a memory, wherein: Memory for storing computer programs; The processor is used to implement the method provided in the first aspect when executing the program stored in the memory.

[0008] According to a fourth aspect of the embodiments of the present application, a computer program product is provided, wherein a computer program is stored in the computer program product, and when the computer program is executed by a processor, the method provided in the first aspect is implemented.

[0009] The content control method based on deep integration of the embodiment of the present application processes the data to be detected into data of K different dimensions, and uses the trained preprocessing model of the corresponding dimension to preprocess the data of K different dimensions respectively to obtain feature vectors of K dimensions, and then, respectively inputs the feature vectors of K different dimensions into the basic models of different dimensions of the deep integration learning model, uses the deep integration learning model to determine the content detection results of the data to be detected, and performs content control processing based on the determined detection results. By fusing multimodal detection results, the accuracy of content detection is improved, thereby improving the accuracy of content control. Alternatively, the feature vectors of K different dimensions can be input into the student model of the deep integration learning model, and the student model can be used to determine the content detection results of the data to be detected, and content control processing is performed based on the determined detection results. By reducing the demand for resources for content control, content control on resource-scarce devices is achieved, and the scope of application of internal control is expanded. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a flowchart of a content management method based on deep integration, shown as an exemplary embodiment of the present application; Figure 2 A schematic diagram of a preprocessing model shown as an exemplary embodiment of the present application; Figure 3 A schematic diagram of a deep ensemble learning model shown as an exemplary embodiment of the present application; Figure 4 A schematic diagram of a model distillation shown as an exemplary embodiment of the present application; Figure 5 This is a schematic structural diagram of a content management device based on deep integration, shown as an exemplary embodiment of the present application; Figure 6 The figure is a schematic diagram of the hardware structure of an electronic device shown as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0011] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, some of the terms involved in the embodiments of the present application are briefly explained below.

[0012] 1. Ensemble learning: In machine learning or deep learning, learning tasks are completed by building and combining multiple learners. For example, a certain algorithm or model generates multiple sub-models, and then these models are combined according to certain methods to solve a specific problem. Usually, these methods include bagging, boosting, stacking, etc.

[0013] 2. Deep ensemble learning: This approach shares the same principle as ensemble learning, integrating multiple sub-models to solve a specific problem. The difference is that it combines deep learning model construction and training differences to propose new ensemble ideas, including dropout, test dataset expansion (TTA), and snapshots.

[0014] 3. Model distillation: Model distillation can distill a larger model into a smaller model. By training the teacher model and the student model, the smaller student model can imitate the teacher model and minimize the difference loss between the two models.

[0015] Exemplarily, model distillation implementation methods may include but are not limited to: online distillation, offline distillation, soft label distillation, intermediate layer feature distillation, adversarial distillation, etc.

[0016] In order to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application are further described in detail below with reference to the accompanying drawings.

[0017] See Figure 1 , is a flow chart of a content management method based on deep integration provided in an embodiment of the present application, such as Figure 1 As shown, the content management method based on deep integration may include the following steps: Step S100: Process the data to be detected into data of K different dimensions; wherein K≥2.

[0018] Exemplarily, the data to be detected may include data requiring content detection so as to perform content control processing on the data.

[0019] For example, the data to be detected may include but is not limited to videos, pictures, recordings, or texts to be published.

[0020] In the embodiment of the present application, in order to improve the accuracy of content control, content detection can be performed on the data to be detected from multiple different dimensions.

[0021] Accordingly, the data to be detected can be processed into data of K (K ≥ 2) different dimensions.

[0022] Exemplarily, the above-mentioned different dimensions may include an image dimension, a text dimension, or a voice dimension.

[0023] Step S110 : For data of any dimension among the K data of different dimensions, use the trained preprocessing model of the dimension to preprocess the data of the dimension to obtain a feature vector of the dimension.

[0024] In an embodiment of the present application, preprocessing models of different dimensions can be pre-trained to preprocess data of different dimensions respectively to obtain feature vectors of corresponding dimensions.

[0025] For example, the preprocessing model may use different models for corresponding dimensions.

[0026] For example, for the text dimension, the preprocessing model may include but is not limited to an LSTM (Long Short-Term Memory) model or a transformer model.

[0027] For the image dimension, the preprocessing model may include but is not limited to a CNN (Convolutional Neural Network) model or a Resnet (Residual Network) model.

[0028] For the speech dimension, the preprocessing model may include but is not limited to: Conformer model (CNN+transformer model, i.e., convolution-enhanced transformer model) or transformer model, etc.

[0029] Correspondingly, for data of any dimension among the above K different dimensions, the trained preprocessing model of the dimension can be used to preprocess the data of the dimension to obtain the feature vector of the dimension.

[0030] For example, assuming that the data to be detected is processed to obtain data of K dimensions including image dimension data and text dimension data, the image dimension data can be preprocessed using a trained image dimension preprocessing model to obtain a feature vector of the image dimension; and the text dimension data can be preprocessed using a trained text dimension preprocessing model.

[0031] For example, in order to improve the processing efficiency of the preprocessing model in processing the data of the corresponding dimension, the data of the corresponding dimension may be encoded before the preprocessing model is used to process the data of the corresponding dimension.

[0032] For example, data in the text dimension, image dimension, and voice dimension can be encoded separately: 1) Text dimension: Use word segmentation and embedding technology to process text data into N*1 vectors; 2) Image Dimension: Image data is processed into an M*M dimensional tensor using resizing (e.g., scaling, cropping, padding) and normalization techniques. 3) Speech dimension: The speech data is converted into MFCC (Mel Frequency Cepstral Coefficients) features, or processed into a P*1 vector using an audio feature extractor.

[0033] Among them, the calculation process of MFCC can include the following steps: Preprocessing: This includes pre-emphasis and framing. Pre-emphasis is used to enhance high-frequency components, while framing divides the continuous speech signal into multiple short time periods (frames) for processing.

[0034] Fast Fourier Transform (FFT): Perform FFT on each frame of speech signal to obtain the spectrum.

[0035] Mel filter bank: Filters the spectrum using a set of Mel filters, where each filter corresponds to a frequency range. The output of these filters represents the energy of different frequency bands.

[0036] Logarithmic energy calculation: Take the logarithm of the output of each filter to obtain the logarithmic energy.

[0037] Discrete Cosine Transform (DCT): Perform discrete cosine transform on the logarithmic energy to obtain MFCC coefficients.

[0038] Exemplarily, for the obtained MFCC coefficients, a part of the MFCC coefficients, such as the coefficients of the preset number, can be used as a feature vector. Step S120, taking K feature vectors of different dimensions as the input of the trained content detection model, using the trained content detection model to determine the content detection results of the data to be detected, and performing content control processing on the data to be detected based on the content detection results; wherein, the content detection model includes a deep ensemble learning model, or, using a model distillation method, a student model generated based on the deep ensemble learning model; the deep ensemble learning model includes a plurality of basic models of different dimensions, and the deep ensemble learning model determines the content detection results of the data to be detected based on the detection results of the feature vectors of the corresponding dimensions by the basic models of different dimensions; for the non-first basic model in the deep ensemble learning model, the input of the basic model includes the feature vectors of the corresponding dimensions, and the output of the feature output layer of the previous basic model.

[0039] In an embodiment of the present application, in order to improve the accuracy of content detection, a deep integrated learning model can be used for content detection.

[0040] Exemplarily, a deep ensemble learning model may include base models of multiple different dimensions.

[0041] For example, the deep ensemble learning model may include at least two of a base model in a text dimension, a base model in an image dimension, and a base model in a speech dimension.

[0042] In one example, the deep ensemble learning model may include a base model in a text dimension, a base model in an image dimension, and a base model in a speech dimension.

[0043] For a basic model of any dimension, it can be used to perform content detection on the data to be detected based on the feature vector of the corresponding dimension.

[0044] For example, for a basic model of the text dimension, content detection can be performed on the data to be detected based on the feature vector of the text dimension.

[0045] The deep ensemble learning model can determine the final content detection results of the data to be detected based on the content detection results of the basic models in each dimension.

[0046] Exemplarily, the content detection result may include normal or abnormal.

[0047] For example, when the content detection result is abnormal, the abnormality type can be further distinguished (the deep integrated learning model needs to be trained accordingly).

[0048] Exemplarily, in order to further improve the accuracy of content detection, for the non-first basic model in the deep integrated learning model, the input of the basic model may include not only the feature vector of the corresponding dimension, but also the output of the feature output layer of the previous basic model. That is, the non-first basic model can also refer to the features output by the previous basic model during the process of performing content detection on the data to be detected.

[0049] In one example, the feature output layer of the base model can be the second-to-last layer of the base model.

[0050] In the embodiments of the present application, it is considered that in actual scenarios, there may be some devices with relatively scarce resources that also have content management needs, such as servers or industrial computers or edge terminal devices with lower resources. For these types of devices, directly deploying the above-mentioned deep integrated learning model may result in insufficient resources.

[0051] Accordingly, for devices with relatively scarce resources and content control needs, the model distillation method can be used to use the above-mentioned deep integrated learning model as the teacher model to generate a student model, and the student model can be deployed on the device for content control. In this way, resource consumption can be reduced while reducing a certain degree of accuracy, thereby meeting the content control needs of resource-scarce devices.

[0052] It can be seen that in Figure 1 In the method flow shown, the data to be detected is processed into data of K different dimensions, and the data of K different dimensions are preprocessed respectively using the trained preprocessing model of the corresponding dimension to obtain feature vectors of K dimensions. Then, the feature vectors of K different dimensions are respectively input into the basic models of different dimensions of the deep integrated learning model, and the deep integrated learning model is used to determine the content detection results of the data to be detected, and content control processing is performed based on the determined detection results. By fusing multimodal detection results, the accuracy of content detection is improved, thereby improving the accuracy of content control. Alternatively, the feature vectors of K different dimensions can be input into the student model of the deep integrated learning model, and the student model is used to determine the content detection results of the data to be detected, and content control processing is performed based on the determined detection results. By reducing the demand for resources for content control, content control on resource-scarce devices is achieved, and the scope of application of internal control is expanded.

[0053] In some embodiments, the K different dimensions include at least two of an image dimension, a text dimension, and a speech dimension; For any dimension among the K different dimensions, if data of this dimension cannot be generated based on the data to be detected, the data of this dimension obtained by processing the data to be detected is empty, and the feature vector of the data to be detected corresponding to this dimension is all 0.

[0054] For example, considering that in actual applications, it may be impossible to generate data of preset dimensions for the data to be detected.

[0055] For example, for image data, it can generate image-dimensional data and text-dimensional data (for example, text information obtained from an image through OCR (optical character recognition)), but generally cannot generate voice-type data. For voice data, it can generate voice-dimensional data and text-dimensional data (for example, text information obtained from voice through voice recognition), but generally cannot generate image-dimensional data.

[0056] In order to be compatible with the above situation, for any dimension of the K different dimensions, if the data of this dimension cannot be generated based on the data to be detected, the data of this dimension obtained by processing the data to be detected is empty, and the feature vector of the data to be detected corresponding to this dimension is all 0.

[0057] Exemplarily, when the data to be detected is image data and the K different dimensions include a speech dimension, the data of the speech dimension obtained by processing the data to be detected is empty, and the feature vector of the data to be detected corresponding to the speech dimension is all 0; When the data to be detected is speech data and the K different dimensions include an image dimension, the data of the image dimension obtained by processing the data to be detected is empty, and the feature vector of the image dimension corresponding to the data to be detected is all 0; When the data to be detected is text data and the K different dimensions include an image dimension, the data of the image dimension obtained by processing the data to be detected is empty, and the feature vector of the image dimension corresponding to the data to be detected is all 0; and / or, when the data to be detected is text data and the K different dimensions include a speech dimension, the data of the speech dimension obtained by processing the data to be detected is empty, and the feature vector of the speech dimension corresponding to the data to be detected is all 0.

[0058] For example, assuming that the data to be detected is text data, and the above K different dimensions include image dimension and speech dimension, then the image dimension data and speech dimension data obtained by processing the data to be detected are both empty, and the feature vector of the image dimension corresponding to the data to be detected and the feature vector of the corresponding speech dimension are both 0.

[0059] In some embodiments, the preprocessing model can be trained by: For any training sample file, process the training sample file into training data of K different dimensions; For the training data of any dimension among the K different dimensions of training data, the training data of the dimension is preprocessed using the preprocessing model of the dimension to be trained to obtain the training feature vector of the dimension; Based on the training feature vector of the dimension, the predicted label of the training data of the dimension is determined, and based on the predicted label and the labeled label of the training data of the dimension, the preprocessing model of the dimension to be trained is feedback-tuned; wherein, the labeled labels of the training data of different dimensions corresponding to the same training sample are the same.

[0060] Exemplarily, in order to implement the training of the preprocessing model, a training sample file for training the preprocessing model may be obtained.

[0061] For example, the training sample files can be labeled by manual labeling, model-assisted labeling, or model-automatic labeling.

[0062] For any training sample file, the training sample file can be processed into data of K different dimensions (which can be called training data), and the K training data of different dimensions are respectively used to train the preprocessing model of the corresponding dimension.

[0063] Exemplarily, the annotation labels of training data of different dimensions corresponding to the same training sample are the same.

[0064] For training data of any dimension among the K different dimensions of training data, the preprocessing model of the dimension to be trained can be used to preprocess the training data of the dimension to obtain a feature vector of the dimension (which can be called a training feature vector).

[0065] During the training process of the preprocessing model, a label prediction part can be set in the preprocessing model, for example, a softmax layer. The label prediction part can determine the predicted label of the training data of the dimension based on the training feature vector obtained in the above manner. Then, the preprocessing model of the dimension to be trained can be fed back and tuned based on the predicted label and the labeled label of the training data of the dimension.

[0066] For example, a training loss (such as a cross-entropy loss) can be determined based on the difference between the predicted label and the labeled label of the training data of the dimension, and the preprocessing model can be feedback-tuned based on the determined training loss.

[0067] It should be noted that in actual applications, when the trained preprocessing model is deployed, the above-mentioned label prediction part does not need to be deployed. That is, when the trained preprocessing model is applied, the trained preprocessing model does not need to perform label prediction, but can obtain the feature vector of the data of the corresponding dimension in the above manner.

[0068] In some embodiments, the deep ensemble learning model can be trained by: For any training sample file, process the training sample file into training data of K different dimensions; For the training data of any dimension among the K different dimensions of training data, the training data of the dimension is preprocessed using the preprocessing model trained for the dimension to obtain the training feature vector of the dimension; K training feature vectors of different dimensions are used as the input of the basic model of the corresponding dimensions of the deep ensemble learning model to be trained. The deep ensemble learning model to be trained is used to determine the predicted labels of the training sample files. Based on the predicted labels and the labeled labels of the training sample files, the deep ensemble learning model to be trained is fed back and tuned.

[0069] Exemplarily, in order to implement the training of the preprocessing model, a training sample file for training the preprocessing model may be obtained.

[0070] For example, the training sample files can be labeled by manual labeling, model-assisted labeling, or model-automatic labeling.

[0071] For any training sample file, the training sample file can be processed into data of K different dimensions (which can be called training data), and the training data of K different dimensions can be used to train the deep integration learning model.

[0072] In one example, the training data used to train the deep ensemble learning model can be the same batch of training data as the training data used to train the preprocessing model.

[0073] For example, K training feature vectors of different dimensions can be used as inputs of the basic models of corresponding dimensions of the deep ensemble learning model to be trained, and the predicted labels of the training sample files can be determined based on the output results of the basic models of different dimensions.

[0074] For example, for any training sample, the K training feature vectors of different dimensions corresponding to the training sample can be input into the basic model of the corresponding dimension of the deep integration learning model to be trained, and the predicted label of the training sample file can be determined by weighted averaging based on the output results of the basic models of different dimensions.

[0075] Exemplarily, feedback tuning can be performed on the deep ensemble learning model to be trained based on the predicted labels of the training sample files and the annotated labels of the training sample files.

[0076] In one example, when the content detection model is a student model generated by a deep ensemble learning model using a model distillation method, the content detection model training further includes: The trained deep ensemble learning model is used as the teacher model. Based on the training feature vectors of K different dimensions and the corresponding annotation labels, the constructed student model is trained using the model distillation method to obtain the trained student model.

[0077] Exemplarily, when the content detection model is a student model of a deep ensemble learning model, and a trained deep ensemble learning model is obtained in the above manner, the trained deep ensemble learning model can also be used as a teacher model. Based on the above K training feature vectors of different dimensions and the corresponding annotation labels, the constructed student model can be trained using the model distillation method to obtain a trained student model.

[0078] As an example, the above-mentioned training feature vectors of K different dimensions and the corresponding annotation labels are used to train the constructed student model using the model distillation method, which may include: Based on the K training feature vectors of different dimensions, the output probability distributions of the deep ensemble learning model and the student model corresponding to the same training samples are determined respectively, and the first loss is determined based on the output probability distribution of the deep ensemble learning model and the output probability distribution of the student model; as well as, Determine the predicted label of the training sample by the student model based on the K training feature vectors of different dimensions, and determine the second loss based on the predicted label and the labeled label of the training sample; Feedback tuning is performed on the student model based on the first loss and the second loss.

[0079] For example, during the model distillation process, the output probability distribution of the teacher model can be used as a "soft label" and the school model can be trained based on the "soft label".

[0080] Exemplarily, during the model distillation process, on the one hand, the output probability distribution of the deep ensemble learning model (i.e., the teacher model) and the student model corresponding to the same training samples can be determined based on K training feature vectors of different dimensions, and the corresponding loss (which can be called the first loss) can be determined based on the output probability distribution of the deep ensemble learning model and the output probability distribution of the student model.

[0081] Exemplarily, the first loss may be KL divergence.

[0082] On the other hand, the student model’s predicted labels for the training samples can be determined based on the K training feature vectors of different dimensions, and the corresponding loss (which can be called the second loss) can be determined based on the predicted labels and the labeled labels of the training samples.

[0083] Exemplarily, the second loss may be a cross entropy loss.

[0084] When the first loss and the second loss are determined, the comprehensive loss of the student model can be determined based on the first loss and the second loss, and the student model can be feedback-tuned based on the comprehensive loss.

[0085] For example, the weighted average of the first loss and the second loss can be determined as the comprehensive loss.

[0086] It should be noted that in the process of determining the output probability distribution of the teacher model, a temperature parameter (T) can also be introduced. During the training process, the T value can be increased to amplify the differences between categories, so that the output probability distribution of the teacher model can convey knowledge more richly.

[0087] In some embodiments, when the base model of the deep ensemble learning model includes a base model of the text dimension, the first base model in the deep ensemble learning model is the base model of the text dimension, and the input of the first base model of the deep ensemble learning model is a feature vector of the text dimension.

[0088] For example, considering that various types of raw data (such as the data to be tested mentioned above) can generally generate data of text dimension, therefore, when the first basic model in the deep ensemble learning model is set to a text model, it can effectively avoid the input of the first basic model of the deep ensemble learning model being empty.

[0089] Since the input of non-first base models in the deep ensemble learning model includes the output of the feature output layer of the previous base model, when the input of the first base model is not empty, the input of each base model will not be empty, thereby improving the content detection accuracy of the deep ensemble learning model.

[0090] Based on this, when the base model of the deep ensemble learning model includes a base model of the text dimension, the first base model in the deep ensemble learning model is the base model of the text dimension, and the input of the base model is the feature vector of the text dimension.

[0091] In some embodiments, when the base model of the deep integrated learning model includes a base model of text dimension, a base model of image dimension, and a base model of speech dimension, the order of the base models of each dimension in the deep integrated learning model is: base model of text dimension, base model of image dimension, and base model of speech dimension.

[0092] Exemplarily, the input of the non-first base model in the deep ensemble learning model includes the output of the feature output layer of the previous base model.

[0093] In addition, in actual scenarios, the raw data that can generate image dimensions is usually more than the raw data that can generate speech dimensions.

[0094] Based on this, when the basic model of the deep integrated learning model includes a basic model of the text dimension, a basic model of the image dimension, and a basic model of the speech dimension, the order of the basic models of each dimension in the deep integrated learning model is: the basic model of the text dimension, the basic model of the image dimension, and the basic model of the speech dimension. Therefore, the amount of input information of the latter basic models in the deep integrated learning model can be effectively improved, and the content detection accuracy of the deep integrated learning model can be improved.

[0095] It should be noted that in an embodiment of the present application, the number of basic models of the same dimension in the deep integrated learning model can be multiple.

[0096] For example, assuming that the deep integrated learning model includes two basic models of text dimensions (which can be denoted as A1 and A2), two basic models of image dimensions (which can be denoted as B1 and B2), and one basic model of speech dimension (which can be denoted as C), then the order of the basic models in the deep integrated learning model can be A1, A2 (the order of A1 and A2 can be adjusted), B1, B2 (the order of B1 and B2 can be adjusted), and C.

[0097] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, the technical solutions provided by the embodiments of the present application are described below with reference to specific examples.

[0098] The embodiment of the present application provides a content control method based on deep learning. In the process of solving the content detection goal, a deep integrated learning framework that can process multimodal data input is designed to train and generate a teacher model. The larger-scale teacher model can be distilled into a smaller-scale student model through model distillation, thereby achieving high accuracy and multimodal data support for content detection and control.

[0099] The following describes the implementation process of the content management solution based on deep integration provided in the embodiment of the present application.

[0100] 1: Multimodal content data processing.

[0101] In this embodiment, training can be performed based on labeled data. The original training sample files include sample files of dimensions such as video, text, pictures, and voice, and are labeled as normal or abnormal (for example, without specifically distinguishing the abnormal type). For each type of original training sample file, data preprocessing is performed from the three dimensions of text, pictures, and voice. Specifically: 1.1. Process various original training sample files into three dimensions: text, image, and speech.

[0102] For example, a video file can generate the above three dimensional data respectively; a picture file can generate data in two dimensions: image and text; and a voice file can generate data in two dimensions: voice and text.

[0103] 1.2. Data encoding: Encode the data in the text dimension, image dimension, and voice dimension respectively for preprocessing model training and test input.

[0104] 1.2.1) Text dimension: Use word segmentation and embedding techniques to process text data into N*1 vectors.

[0105] 1.2.2) Image Dimensionality: Image data is processed into an M*M dimensional tensor using resizing (e.g., scaling, cropping, padding) and normalization techniques.

[0106] 1.2.3) Speech Dimension: Convert the speech data into MFCC features, or use an audio feature extractor to process it into a P*1 vector.

[0107] 1.3. Preprocessing model construction and pre-training.

[0108] Exemplarily, pre-processing model training is performed on the data processed in the text dimension, image dimension, and voice dimension respectively, for use as input for the subsequent deep integration learning model.

[0109] For example, during the training process, the schematic diagram of the preprocessing model can be as follows: Figure 2 In the actual application stage, the trained preprocessing model does not need to output predicted labels, but instead outputs the extracted data features as the input of the content detection model.

[0110] 1.3.1) Text Dimension: Build a text feature extraction model (also called a text preprocessing model, i.e., a preprocessing model for the text dimension). Use the N*1 vectors and corresponding labels obtained in 1.2.1 to train the text preprocessing model. After the model converges, solidify the model parameters to obtain the text preprocessing model Mw.

[0111] Exemplarily, the text preprocessing model may adopt an LSTM or transformer model, wherein the second layer and the penultimate layer of the model are fully connected layers of X neurons, and the output of the penultimate layer of the model (i.e., the fully connected layer) is the extracted text feature vector.

[0112] 1.3.2) Image Dimension: Build an image feature extraction model (also called an image preprocessing model, i.e., an image-dimensional preprocessing model). Use the M*M dimensional tensor and corresponding labels obtained in 1.2.2 to train the image preprocessing model. After the model converges, solidify the model parameters to obtain the image preprocessing model Mp.

[0113] Exemplarily, the image preprocessing model may adopt CNN or Resnet, wherein the second layer and the penultimate layer of the model are fully connected layers of X neurons, and the output of the penultimate layer of the model (i.e., the fully connected layer) is the extracted image feature vector.

[0114] 1.3.3) Speech Dimension: Build a speech feature extraction model (also called a speech preprocessing model, i.e., a preprocessing model for the speech dimension). Use the MFCC features obtained in 1.2.3, or the P*1 vector and corresponding labels, to train the speech preprocessing model. After model convergence, solidify the model parameters to obtain the speech preprocessing model Mv.

[0115] Exemplarily, the speech preprocessing model may adopt a transformer model, wherein the second layer and the penultimate layer of the model are fully connected layers of X neurons, and the output of the penultimate layer of the model (i.e., the fully connected layer) is the extracted speech feature vector.

[0116] 2. Construction and training of deep integrated learning framework.

[0117] Exemplarily, after step one, all original training sample files can be processed into feature vectors of size X*1 for use in the training of the deep integration learning framework in this step.

[0118] 2.1. Basic model construction.

[0119] Exemplarily, a basic model is constructed for data learning in each dimension (speech dimension, text dimension, image dimension).

[0120] 2.2.1) Text Dimension Basic Model: Use the text preprocessing model in step 1, or use a Transformer model with a fully connected layer consisting of X neurons in the second and penultimate layers.

[0121] 2.2.2) Image Dimension Base Model: Use the image preprocessing model in step 1, or use a ResNet model with a second and penultimate fully connected layer consisting of X neurons.

[0122] 2.2.3) Speech Dimension Basic Model: Use the speech preprocessing model in step 1, or use the Conformer model with a second and penultimate fully connected layer consisting of X neurons.

[0123] 2.2. Construction of deep integrated learning framework.

[0124] Exemplarily, a deep integrated learning base model is used to train base models of multiple dimensions simultaneously, and the output of the base model of each dimension is weighted averaged as the final output.

[0125] Exemplarily, each dimension may include one or more basic models.

[0126] For example, when the number of basic models in each dimension is 1, the schematic diagram of the deep ensemble learning model can be as follows: Figure 3 shown.

[0127] 2.2.1) For each original training sample file, process it from three dimensions: text dimension, image dimension, and speech dimension.

[0128] For example, data of different dimensions generated from the same original training sample file (which may be a file of video, text, voice, image, etc.) share the same label (which can be labeled, and the label can be normal or abnormal), and the original data is processed into a feature vector of dimension X*1 through the preprocessing model in step one.

[0129] 2.2.2) Each time during training, a three-dimensional feature vector of an original training sample file is used to train the deep ensemble learning model.

[0130] For example, if a dimension does not exist, it can be filled with 0. For example, if the original training sample file is a picture, the feature vector of the speech dimension can be a vector consisting of X zeros.

[0131] 2.2.3) Take the weighted average of the outputs Fw, Fp, and Fv of each base model to obtain the model prediction value (i.e., the predicted label). Calculate the difference between the predicted value and the true label (i.e., the annotated label) to obtain the loss. Minimize the loss to reversely train the entire deep ensemble learning model.

[0132] 3. Model distillation.

[0133] For example, deep ensemble learning can effectively integrate the data features of the original sample in three dimensions, enabling multimodal feature extraction and content detection. However, this significantly increases the overall model size, resulting in reduced model efficiency. This solution can use model distillation to mitigate the model size issue caused by deep ensemble learning.

[0134] 3.1. Student model construction: The student model can adopt a brand new model (such as the BERT model), or a transformer model.

[0135] For example, taking the transformer model as an example, a fully connected layer with X*3 neurons can be introduced into the input layer as a student model (Mstudent).

[0136] 3.2. Online distillation: Use online distillation to distill the model. Use the training data in step 1 to train the teacher model (i.e. the above-mentioned deep ensemble learning model Mteacher) and the student model at the same time. 3.3. Soft label distillation: Various distillation methods can be used (such as intermediate feature distillation, adversarial distillation, dynamic distillation, etc.).

[0137] In this embodiment, a soft label distillation method is adopted, that is, Mteacher outputs the probability distribution of samples and trains Mstudent to fit the distribution of the teacher model. In the process, a temperature parameter can also be introduced to amplify the difference between categories. The schematic diagram can be shown as follows: Figure 4 shown.

[0138] For example, the loss function can be set as a joint cross entropy loss (determined according to the true label) and KL divergence loss (determined according to the soft label).

[0139] 4. Model Output After the above three steps, the final model is output: 4.1. Mw, Mp, and Mv: used to perform feature processing on the original samples in three dimensions respectively.

[0140] 4.2. Mteacher: used for content management on high-resource servers, such as content management on the central end.

[0141] 4.3. Mstudent: Used for content management and control on servers, industrial computers, or terminals with lower resources, such as edge content management and control.

[0142] The above describes the method provided by this application. The following describes the device provided by this application: See Figure 5, is a structural diagram of a content management device based on deep integration provided in an embodiment of the present application, such as Figure 5 As shown, the content management and control device based on deep integration may include: The first processing unit is configured to process the data to be detected into data of K different dimensions, wherein K is greater than or equal to 2; A second processing unit is configured to preprocess the data of any dimension of the K different dimensional data using a trained preprocessing model for the dimension to obtain a feature vector for the dimension; A content control unit is used to use the K feature vectors of different dimensions as input of a trained content detection model, use the trained content detection model to determine the content detection results of the data to be detected, and perform content control processing on the data to be detected based on the content detection results; wherein, the content detection model includes a deep ensemble learning model, or a student model generated based on the deep ensemble learning model using a model distillation method; the deep ensemble learning model includes a plurality of basic models of different dimensions, and the content detection results of the data to be detected by the deep ensemble learning model are determined based on the detection results of the feature vectors of corresponding dimensions by the basic models of different dimensions; for non-first basic models in the deep ensemble learning model, the input of the basic model includes the feature vectors of corresponding dimensions, and the output of the previous basic model.

[0143] For example, the specific implementation process of the first processing unit, the second processing unit, and the content control unit to implement content control based on deep integration can refer to the relevant description in the above method embodiment.

[0144] An embodiment of the present application also provides an electronic device, including a processor and a memory, wherein the memory is used to store computer programs; the processor is used to implement the content management method based on deep integration described above when executing the program stored in the memory.

[0145] See Figure 6 , is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device may include a processor 601 and a memory 602 storing machine-executable instructions. The processor 601 and the memory 602 may communicate via a system bus 603. Furthermore, by reading and executing the machine-executable instructions corresponding to the deeply integrated content management logic in the memory 602, the processor 601 may execute the deeply integrated content management method described above.

[0146] The memory 602 mentioned herein may be any electronic, magnetic, optical, or other physical storage device that may contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.

[0147] In some embodiments, a machine-readable storage medium is also provided. Figure 6 The memory 602 in the machine-readable storage medium stores machine-executable instructions. When executed by the processor, the machine-executable instructions implement the content management method based on deep integration described above. For example, the machine-readable storage medium can be ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0148] An embodiment of the present application also provides a computer program product that stores a computer program, and when a processor executes the computer program, it prompts the processor to execute the content management method based on deep integration described above.

Claims

1. A content management method based on deep integration, characterized in that: include: Process the data to be detected into data of K different dimensions; where K ≥ 2; For data of any dimension of the K different dimensions, preprocess the data of the dimension using the trained preprocessing model of the dimension to obtain a feature vector of the dimension; The K feature vectors of different dimensions are used as the input of the trained content detection model, and the trained content detection model is used to determine the content detection results of the data to be detected, and content control processing is performed on the data to be detected based on the content detection results; wherein, the content detection model includes a deep integrated learning model, or a student model generated by the deep integrated learning model using a model distillation method; the deep integrated learning model includes a plurality of basic models of different dimensions, and the content detection results of the data to be detected by the deep integrated learning model are determined based on the detection results of the feature vectors of the corresponding dimensions by the basic models of different dimensions; for non-first basic models in the deep integrated learning model, the input of the basic model includes the feature vectors of the corresponding dimensions, and the output of the feature output layer of the previous basic model.

2. The method according to claim 1, characterized in that The K different dimensions include at least two of an image dimension, a text dimension, and a speech dimension; For any dimension of the K different dimensions, if data of the dimension cannot be generated based on the data to be detected, the data of the dimension obtained by processing the data to be detected is empty, and the feature vector of the data to be detected corresponding to the dimension is all 0; Wherein, when the data to be detected is image data and the K different dimensions include a speech dimension, the data of the speech dimension obtained by processing the data to be detected is empty, and the feature vector of the data to be detected corresponding to the speech dimension is all 0; When the data to be detected is speech data and the K different dimensions include an image dimension, data of the image dimension obtained by processing the data to be detected is empty, and a feature vector of the image dimension corresponding to the data to be detected is all zeros; When the data to be detected is text data and the K different dimensions include image dimensions and / or speech dimensions, the data of the image dimensions and / or speech dimensions obtained by processing the data to be detected is empty, and the feature vectors of the image dimensions and / or speech dimensions corresponding to the data to be detected are all 0.

3. The method according to claim 1, characterized in that The preprocessing model is trained in the following way: For any training sample file, process the training sample file into training data of K different dimensions; For the training data of any dimension among the K different dimensions of the training data, preprocess the training data of the dimension using the preprocessing model of the dimension to be trained to obtain a training feature vector of the dimension; Based on the training feature vector of the dimension, the predicted label of the training data of the dimension is determined, and based on the predicted label and the labeled label of the training data of the dimension, the preprocessing model of the dimension to be trained is feedback-tuned; wherein, the labeled labels of the training data of different dimensions corresponding to the same training sample are the same.

4. The method according to claim 1, wherein The deep ensemble learning model is trained in the following way: For any training sample file, process the training sample file into training data of K different dimensions; For the training data of any dimension among the K different dimensions of training data, preprocess the training data of the dimension using the trained preprocessing model of the dimension to obtain a training feature vector of the dimension; The K training feature vectors of different dimensions are respectively used as the input of the basic model of the corresponding dimension of the deep ensemble learning model to be trained, and the predicted label of the training sample file is determined by using the deep ensemble learning model to be trained. Based on the predicted label and the annotated label of the training sample file, the deep ensemble learning model to be trained is feedback-tuned.

5. The method according to claim 4, characterized in that In the case where the content detection model is a student model generated by the deep ensemble learning model using a model distillation method, the content detection model training further includes: The trained deep ensemble learning model is used as the teacher model. According to the K training feature vectors of different dimensions and the corresponding annotation labels, the constructed student model is trained using the model distillation method to obtain the trained student model.

6. The method according to claim 5, characterized in that The constructed student model is trained using a model distillation method based on the K training feature vectors of different dimensions and the corresponding annotation labels, including: Determining, based on the K training feature vectors of different dimensions, the output probability distributions of the deep ensemble learning model and the student model corresponding to the same training sample, and determining a first loss based on the output probability distribution of the deep ensemble learning model and the output probability distribution of the student model; as well as, Determining a predicted label of the student model for the training sample based on the K training feature vectors of different dimensions, and determining a second loss based on the predicted label and the labeled label of the training sample; Feedback tuning is performed on the student model based on the first loss and the second loss.

7. The method according to any one of claims 1 to 6, characterized in that In the case where the base model of the deep ensemble learning model includes a base model of a text dimension, the first base model in the deep ensemble learning model is a base model of a text dimension, and the input of the first base model of the deep ensemble learning model is a feature vector of a text dimension; And / or, when the base model of the deep integrated learning model includes a base model of text dimension, a base model of image dimension and a base model of speech dimension, the order of the base models of each dimension in the deep integrated learning model is: base model of text dimension, base model of image dimension, and base model of speech dimension.

8. A content management and control device based on deep integration, characterized in that: include: The first processing unit is configured to process the data to be detected into data of K different dimensions, wherein K is greater than or equal to 2; A second processing unit is configured to preprocess the data of any dimension of the K different dimensional data using a trained preprocessing model for the dimension to obtain a feature vector for the dimension; A content control unit is used to use the K feature vectors of different dimensions as the input of a trained content detection model, use the trained content detection model to determine the content detection results of the data to be detected, and perform content control processing on the data to be detected based on the content detection results; wherein, the content detection model includes a deep integrated learning model, or a student model generated by the deep integrated learning model using a model distillation method; the deep integrated learning model includes a plurality of basic models of different dimensions, and the content detection results of the data to be detected by the deep integrated learning model are determined based on the detection results of the feature vectors of the corresponding dimensions by the basic models of different dimensions; for non-first basic models in the deep integrated learning model, the input of the basic model includes the feature vectors of the corresponding dimensions, and the output of the feature output layer of the previous basic model.

9. An electronic device, characterized in that: comprising a processor and a memory, wherein, Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 7 when executing a program stored in a memory.

10. A computer program product, characterized in that The computer program product stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Multi-dimensional time sequence anomaly detection method based on characteristic distillation

    CN114996124A

  • Content retrieval and model training method and device, electronic equipment and storage medium

    CN115114395A

  • Video processing model processing method and device, computer equipment and storage medium

    CN116935170A

  • Task prediction method and device based on video multi-modal information

    CN116975615A

  • Unmanned aerial vehicle image target detection method, device, equipment and medium

    CN117456394A