Cross-modal content security detection system and method based on deep learning

Through a cross-modal content security detection system based on deep learning, the modal heterogeneity, semantic gap and real-time problems in cross-modal content security detection are solved, and efficient identification and dissemination of harmful information are achieved.

CN120337155APending Publication Date: 2025-07-18HEBEI LINGHE COMPUTER INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510557765.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There are modal heterogeneity, semantic gap, real-time requirements and data scarcity in cross-modal content security detection, resulting in low accuracy and efficiency of identification of harmful information.

Method used

A cross-modal content security detection system based on deep learning is adopted, and the fusion representation and detection of image, text and audio features are achieved through multimodal data acquisition, image feature extraction, feature representation alignment and content security detection modules, combined with clustering algorithms and GloVe models.

Benefits of technology

It improves the identification accuracy and recognition efficiency of harmful information in cross-modal scenarios, ensures the standardization of the dissemination of cross-modal content, and realizes efficient detection of harmful information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337155A_ABST
    Figure CN120337155A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cross-modal content security detection, in particular to a cross-modal content security detection system and method based on deep learning. The method comprises the following steps: acquiring multi-modal data, carrying out clustering analysis on an image gray value of image data by adopting a clustering algorithm, and determining a stretching range of an equalized image histogram in different clustering spaces according to a proportion of image information contained in each clustering space; carrying out image feature representation extraction on the enhanced image data; performing feature representation alignment on the image feature representation, the text feature representation and the audio feature representation, and fusing the aligned modal representations to obtain a multi-modal fused feature representation; and inputting the multi-modal fusion feature representation into the content security detection model, and outputting a content security detection result, so that the recognition precision and recognition efficiency of harmful information in a cross-modal scene can be improved, and the propagation normalization of cross-modal content is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-modal content security detection, and particularly to a cross-modal content security detection system and method based on deep learning. Background Art

[0002] With the rapid development of the Internet, the generation and dissemination of multi-modal data (such as images, texts, audios, videos, etc.) have become increasingly convenient. This trend has also brought severe challenges in terms of content security, such as the proliferation of harmful content like false information, violence and pornography, terrorist propaganda, online fraud, etc., which pose a serious threat to social order, public safety, and personal rights and interests.

[0003] Currently, the following main challenges exist in cross-modal content security detection: Modal heterogeneity: Data in different modalities have different representation forms and feature spaces. How to effectively fuse this heterogeneous data is a difficult problem.

[0004] Semantic gap: There may be differences in semantic information between different modalities. How to eliminate this semantic gap and achieve cross-modal semantic understanding is the key to cross-modal detection.

[0005] Requirement for real-time performance: In the Internet environment, the content dissemination speed is extremely fast, and it is required that the detection system can process a large amount of data in real-time or near real-time, which poses extremely high requirements for the efficiency of the algorithm.

[0006] Data scarcity: The dataset of cross-modal harmful content is relatively scarce, and the annotation cost is high, which limits the performance of supervised learning methods.

[0007] Therefore, there is an urgent need for a cross-modal content security detection system and method based on deep learning to solve the above problems. Summary of the Invention

[0008] The purpose of the present invention is to provide a cross-modal content security detection system and method based on deep learning: to solve the problem of content security detection in multi-modal data in the prior art, especially the technical problems of low recognition accuracy and low recognition efficiency of harmful information in cross-modal scenarios.

[0009] The purpose of the present invention can be achieved by the following technical solutions: On the one hand, a cross-modal content security detection system based on deep learning, the system includes: A multi-modal data acquisition module, used to acquire multi-modal data, where the multi-modal data includes text data, image data, and audio data; An image feature extraction module, which is used to perform clustering analysis on the gray values of image data by using a clustering algorithm, obtain multiple different clustering spaces, determine the stretching range of the equalized image histogram according to the proportion of the image information contained in each clustering space in different clustering spaces, complete image enhancement, and extract the image feature representation of the enhanced image data; A feature representation alignment module, which is used to extract the text feature representation of text data based on the GloVe model, extract the audio feature representation of audio data, align the image feature representation, text feature representation and audio feature representation, and fuse the aligned representations of each modality to obtain a multi-modal fused feature representation; A content security detection module, which is used to input the multi-modal fused feature representation into a content security detection model and output a content security detection result.

[0010] Further, performing clustering analysis on the gray values of image data by using a clustering algorithm to obtain multiple different clustering spaces specifically includes the following process: Step 1: Calculation of the initial clustering center: Calculate the gray value of the image data and the cosine distance between them: ; where n is the number of pixels in the image data; obtain the minimum value in , and define the cumulative number of the gray value in the image data with a distance of from the gray value as Calculate the contribution rate of the gray value in the image data: ; arrange the contribution rates in descending order, and select the top K gray values with the largest contribution rates as the initial clustering centers; Step 2: Gray value category division: Assign the gray value closest to the clustering center to the same category: For any gray value of the image data, if it satisfies the following formula, it is assigned to the same category: ; where represents the category after m iterations, represents the clustering center with the largest contribution rate, represents the clustering center with a non-maximum contribution rate; Step 3: Iterative accuracy evaluation: Calculate the iterative accuracy of each sample in different categories : ; where, represents the clustering center, is the weight. If belongs to the category with as the clustering center, then takes the value of 1; otherwise takes the value of 0; If the iteration accuracy does not meet the set threshold, update the clustering center; Step Four: Clustering center update: Solve for the mean of different categories and use it as the new clustering center. Repeat Steps Two and Three until the iteration accuracy meets the set threshold to obtain multiple different clustering spaces. Among them, the clustering center of the category in the (m + 1)-th iteration is: ; where, represents the number of gray levels in the category with as the clustering center.

[0011] Furthermore, determine the stretching range of the histogram of the equalized image according to the proportion of the image information contained in each clustering space in different clustering spaces. The specific process of image enhancement includes the following: Assume that the total number of gray levels in the image data P is N, and the number of gray levels contained in a certain clustering space q is , then the histogram proportionality coefficient of the image in this clustering space after equalization is , where represents rounding ; Statistically analyze the histogram proportionality coefficients of all clustering spaces. Construct a rectangular coordinate system with the clustering space number corresponding to the histogram proportionality coefficient as the X-axis and the histogram proportionality coefficient as the Y-axis. Mark all the histogram proportionality coefficients as points in the rectangular coordinate system, connect the adjacent points in the rectangular coordinate system, generate a histogram proportionality coefficient curve, draw perpendicular lines from both ends of the histogram proportionality coefficient curve to the X-axis to obtain two starting and ending line segments. The closed figure formed by the histogram proportionality coefficient curve, the two starting and ending line segments, and the X-axis is marked as the stretching reference coefficient of the image data P; Define the image with a stretching reference coefficient not exceeding the preset stretching threshold as a dark area, and process it using adaptive histogram equalization. Define the image with a stretching reference coefficient exceeding the preset stretching threshold as a bright area, and stretch its gray value using linear correction to complete image enhancement.

[0012] Furthermore, perform image feature extraction on the enhanced image data based on the VGG19 model.

[0013] Furthermore, audio feature representations are extracted based on the openSMILE toolkit.

[0014] Furthermore, the alignment of the image feature representation, text feature representation, and audio feature representation specifically includes the following process: The image feature representation, text feature representation, and audio feature representation are respectively passed through a projection layer , , and mapped to the common space R. The projected features are: the image feature , the text feature , and the audio feature , where is the unprojected image feature representation, is the unprojected text feature representation, and is the unprojected audio feature representation; By maximizing the similarity between matching sample pairs and minimizing the similarity between non-matching sample pairs for the projected features, the features of different modalities are aligned in the common space. The loss function L is as follows: ; where represents the similarity between the image feature and the text feature , represents the similarity between the text feature and the audio feature , represents the similarity between the image feature and the audio feature , and represents the learnable parameters.

[0015] Furthermore, inputting the fused feature representation of multiple modalities into the content security detection model and outputting the content security detection result specifically includes the following process: Input the fused feature representation of the cross-modal content to be detected, and construct a heterogeneous graph based on the fused feature representation; Perform contrastive learning on the heterogeneous graph and output the content security detection result.

[0016] Furthermore, constructing a heterogeneous graph based on the fused feature representation specifically includes the following process: Input: The fused feature representation u, the set Vu of dangerous content factors associated with the fused feature representation u, the number Mu of selected dangerous content factors in the set Vu, the number Z of selected neighbors, where Z represents the number of neighbors selected by each node, the node set N1: containing all nodes in layer L1, initially only containing the fused feature representation u; the number of layers H, where the number of layers H is the number of layers of the graph, indicating the number of layers expanded outward from the fused feature representation u. Output the constructed heterogeneous graph: Embed the fused feature representation u into the initial graph Gu, and the initial node set N1 = {u}; Select Mu dangerous content factor information from Vu and update N1; Construct an edge u→v in Gu and mark the edge type, where v is a dangerous content factor of u; The construction process includes: for each layer L1, from the second layer 2 to H, construct a temporary node set Ntmp1 to store the nodes of the current layer, and traverse the node set Nl−1 of the previous layer: For each node i, construct a temporary node set Ntmp2. If i is a dangerous content factor node and has more than m′ interactive content factors, select m′ latest interactive content factors and add them to Ntmp2. If i is a fused feature representation node and has more than m′ interactive content factors, select m′ latest interactive content factors and add them to Ntmp2. For each neighbor node nneighbor in Ntmp2, construct an edge i→nneighbor in Gu and mark the edge type, update Ntmp1 = Ntmp1 ∪ Ntmp2, and update the node set Nl of the current layer to Nl = Ntmp1; Use the breadth-first search algorithm to obtain a subgraph from Gu, add the subgraph to the overall graph GG, update i = i + 1, and obtain the heterogeneous graph.

[0017] Furthermore, perform contrastive learning on the heterogeneous graph, and the output content security detection result specifically includes the following process: Compare the heterogeneous graph with a preset heterogeneous graph to determine whether they are the same. If so, there is no security hazard content in the cross-modal content. If not, there is security hazard content in the cross-modal content.

[0018] On the other hand, a cross-modal content security detection method based on deep learning, the method includes: Obtain multimodal data, where the multimodal data includes text data, image data, and audio data; The clustering algorithm is used for the image data to perform clustering analysis on the image gray values, obtaining multiple different clustering spaces. In different clustering spaces, the stretching range of the equalized image histogram is determined according to the proportion of the image information contained in each clustering space, completing image enhancement, and extracting the image feature representation of the enhanced image data; Based on the GloVe model, the text feature representation of the text data is extracted, the audio feature representation of the audio data is extracted, and the image feature representation, text feature representation, and audio feature representation are aligned in terms of feature representation, and the aligned representations of each modality are fused to obtain the multi-modal fused feature representation; The multi-modal fused feature representation is input into the content security detection model, and the content security detection result is output.

[0019] Compared with the existing solutions, the beneficial effects achieved by the present invention are as follows: The present invention obtains multi-modal data. The clustering algorithm is used for the image data to perform clustering analysis on the image gray values, obtaining multiple different clustering spaces. In different clustering spaces, the stretching range of the equalized image histogram is determined according to the proportion of the image information contained in each clustering space, completing image enhancement, and extracting the image feature representation of the enhanced image data; based on the GloVe model, the text feature representation of the text data is extracted, the audio feature representation of the audio data is extracted, and the image feature representation, text feature representation, and audio feature representation are aligned in terms of feature representation, and the aligned representations of each modality are fused to obtain the multi-modal fused feature representation; the multi-modal fused feature representation is input into the content security detection model, and the content security detection result is output, which can improve the recognition accuracy and efficiency of harmful information in the cross-modal scenario and ensure the dissemination standardization of cross-modal content.

[0020] Furthermore, by constructing a unified feature representation space, capturing the interaction between modalities, and designing an efficient feature fusion strategy, accurate and efficient detection of cross-modal harmful content is achieved. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is the system block diagram of a cross-modal content security detection system based on deep learning according to an embodiment of the present invention; Figure 2 It is the working flowchart of the first cross-modal content security detection system based on deep learning according to an embodiment of the present invention; Figure 3 It is the workflow diagram of the second cross-modal content security detection system based on deep learning in the embodiments of the present invention; Figure 4 It is the workflow diagram of a cross-modal content security detection method based on deep learning in the embodiments of the present invention. Specific embodiments

[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present invention.

[0024] In addition, the described features, structures or characteristics can be combined in any suitable manner in one or more example embodiments. In the following description, many specific details are provided to give a full understanding of the example embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced by omitting one or more of the specific details, or by using other methods, components, steps, etc. In other cases, well-known structures, methods, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0025] This embodiment provides a cross-modal content security detection system based on deep learning. Figure 1 It is the system block diagram of a cross-modal content security detection system based on deep learning in the embodiments of the present invention. As Figure 1 shown, the system includes: A multi-modal data acquisition module for acquiring multi-modal data, where the multi-modal data includes text data, image data, and audio data; An image feature extraction module for performing clustering analysis on the gray values of images using a clustering algorithm for the image data to obtain multiple different clustering spaces, determining the stretching range of the equalized image histogram according to the proportion of the image information contained in each clustering space in different clustering spaces, completing image enhancement, and extracting the image feature representation of the enhanced image data; A feature representation alignment module for extracting the text feature representation of the text data based on the GloVe model, extracting the audio feature representation of the audio data, aligning the image feature representation, the text feature representation, and the audio feature representation, and fusing the aligned representations of each modality to obtain a multi-modal fused feature representation; A content security detection module for inputting the multi-modal fused feature representation into a content security detection model and outputting a content security detection result.

[0026] In summary, the present invention obtains multi-modal data, performs clustering analysis on the gray values of the image data using a clustering algorithm to obtain multiple different clustering spaces, determines the stretching range of the equalized image histogram according to the proportion of the image information contained in each clustering space in different clustering spaces, completes image enhancement, and extracts the image feature representation of the enhanced image data; extracts the text feature representation of the text data based on the GloVe model, extracts the audio feature representation of the audio data, aligns the image feature representation, the text feature representation, and the audio feature representation, and fuses the aligned representations of each modality to obtain a multi-modal fused feature representation; inputs the multi-modal fused feature representation into the content security detection model, and outputs the content security detection result, which can improve the recognition accuracy and recognition efficiency of harmful information in the cross-modal scenario and ensure the dissemination standardization of cross-modal content.

[0027] In some embodiments, Figure 2 is the flowchart of the first cross-modal content security detection system based on deep learning according to the embodiment of the present invention. As Figure 2 shown, performing clustering analysis on the gray values of the image data using a clustering algorithm to obtain multiple different clustering spaces specifically includes the following processes: Step 1: Calculation of the initial clustering center; Calculate the cosine distance between and : ; where n is the number of pixels of the image data; obtain the minimum value in , define the cumulative number of the gray value in the image data whose distance from the gray value is ; Calculate the contribution rate of the gray value in the image data: ; arrange the contribution rates in descending order, and select the top K gray values with the largest contribution rates as the initial clustering centers; Step 2: Classification of gray value categories; Classify the gray value closest to the clustering center into one category: For any gray value of the image data, if it satisfies the following formula, it is classified into one category: ; where represents the category after m iterations, Represents the clustering center with the largest contribution rate, Represents the clustering center with the non-largest contribution rate; Step 3: Iterative precision evaluation; Calculate the iterative precision of each sample in different categories : ; Among them, Represents the clustering center, Is the weight. If Belongs to the category with As the clustering center, then The value is 1, otherwise The value is 0; If the iterative precision does not meet the set threshold, update the clustering center; Step 4: Clustering center update; Solve the mean values of different categories and use them as the new clustering centers. Repeat steps 2 and 3 until the iterative precision meets the set threshold to obtain multiple different clustering spaces. Among them, the clustering center of the th category in the (m + 1)th iteration Is: ; Among them, Represents the number of gray levels in the category with As the clustering center.

[0028] In some embodiments, determine the stretching range of the equalized image histogram according to the proportion of the image information contained in each clustering space in different clustering spaces. Completing the image enhancement specifically includes the following process: Assume that the total number of gray levels in the image data P is N, and the number of gray levels contained in a certain clustering space q is , then the histogram proportional coefficient of the equalized image in this clustering space is , among which, Represents rounding ; Statistical histogram proportional coefficients of all clustering spaces. Construct a rectangular coordinate system with the clustering space numbers corresponding to the histogram proportional coefficients as the X-axis and the histogram proportional coefficients as the Y-axis. Mark all the histogram proportional coefficients as points in the rectangular coordinate system, connect the adjacent points in the rectangular coordinate system, generate a histogram proportional coefficient curve, draw perpendicular lines from both ends of the histogram proportional coefficient curve to the X-axis to obtain two starting and ending line segments. The closed figure is formed by the histogram proportional coefficient curve, the two starting and ending line segments, and the X-axis. Mark the total area of the closed figure as the stretching reference coefficient of the image data P; Images with a stretching reference coefficient not exceeding a preset stretching threshold are defined as dark regions, which are processed using adaptive histogram equalization. Images with a stretching reference coefficient exceeding the preset stretching threshold are defined as bright regions, and their gray values are stretched using linear correction to complete image enhancement.

[0029] Further, image features are extracted from the enhanced image data based on the VGG19 model, and audio feature representations are extracted based on the openSMILE toolbox.

[0030] In some embodiments, the alignment of the image feature representation, text feature representation, and audio feature representation specifically includes the following process: The image feature representation, text feature representation, and audio feature representation are respectively passed through a projection layer 、 、 and mapped to the common space R. The projected features are: the image feature , the text feature , and the audio feature , where is the unprojected image feature representation, is the unprojected text feature representation, is the unprojected audio feature representation; The projected features are aligned in the common space by maximizing the similarity between matching sample pairs and minimizing the similarity between non-matching sample pairs. The loss function L is as follows: ; where represents the similarity between the image feature and the text feature , represents the similarity between the text feature and the audio feature , represents the similarity between the image feature and the audio feature , represents the learnable parameter.

[0031] In some embodiments, Figure 3 is the flowchart of the second cross-modal content security detection system based on deep learning in the embodiments of the present invention. As shown in Figure 3 , inputting the multi-modal fusion feature representation into the content security detection model and outputting the content security detection result specifically includes the following process: Step S301: Input the fusion feature representation of the cross-modal content to be detected, and construct a heterogeneous graph based on the fusion feature representation; Specifically, constructing a heterogeneous graph based on the fused feature representation specifically includes the following process: Input: The fused feature representation u, the set Vu of dangerous content factors associated with the fused feature representation u, the number Mu of selected dangerous content factors in the set Vu, the number Z of selected neighbors, where Z represents the number of neighbors selected by each node, the node set N1: containing all nodes in layer L1, initially only containing the fused feature representation u; the number of layers H, where the number of layers H is the number of layers of the graph, indicating the number of layers extended outward from the fused feature representation u. Output: The constructed heterogeneous graph: Embed the fused feature representation u into the initial graph Gu, and the initial node set N1 = {u}; Select Mu dangerous content factor information from Vu and update N1; Construct an edge u → v in Gu and mark the edge type, where v is a dangerous content factor of u; The construction process includes: For each layer L1, from the second layer 2 to H, construct a temporary node set Ntmp1 to store the nodes of the current layer, and traverse the node set Nl−1 of the previous layer: For each node i, construct a temporary node set Ntmp2. If i is a dangerous content factor node and has more than m′ interaction content factors, select m′ of the latest interaction content factors and add them to Ntmp2. If i is a fused feature representation node and has more than m′ interaction content factors, select m′ of the latest interaction content factors and add them to Ntmp2. For each neighbor node nneighbor in Ntmp2, construct an edge i → nneighbor in Gu and mark the edge type, update Ntmp1 = Ntmp1 ∪ Ntmp2, and update the node set Nl of the current layer to Nl = Ntmp1; Use the breadth-first search algorithm to obtain a subgraph from Gu, add the subgraph to the total graph GG, update i = i + 1, and obtain the heterogeneous graph.

[0032] Step S302: Perform contrastive learning on the heterogeneous graph and output the content security detection result.

[0033] Specifically, performing contrastive learning on the heterogeneous graph and outputting the content security detection result specifically includes the following process: Compare the heterogeneous graph with a preset heterogeneous graph to determine whether they are the same. If so, there is no security hazard content in this cross-modal content. If not, there is security hazard content in this cross-modal content.

[0034] In some embodiments, the present invention also provides a cross-modal content security detection method based on deep learning. Figure 4It is a flowchart of a cross-modal content security detection method based on deep learning according to an embodiment of the present invention. As Figure 4 shown, the method includes the following steps: Step S401: Obtain multi-modal data, where the multi-modal data includes text data, image data, and audio data; Step S402: Use a clustering algorithm to perform clustering analysis on the image gray values of the image data to obtain multiple different clustering spaces. Determine the stretching range of the equalized image histogram according to the proportion of the image information contained in each clustering space in different clustering spaces, complete image enhancement, and extract the image feature representation of the enhanced image data; Step S403: Extract the text feature representation of the text data based on the GloVe model, extract the audio feature representation of the audio data, align the image feature representation, text feature representation, and audio feature representation, and fuse the aligned representations of each modality to obtain a multi-modal fused feature representation; It should be noted that the process of fusing the aligned representations of each modality is to splice the three representations.

[0035] Step S404: Input the multi-modal fused feature representation into the content security detection model and output the content security detection result.

[0036] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0037] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraint requirements of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0038] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0039] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only for some logical function divisions, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0040] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0041] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A cross-modal content security detection system based on deep learning, characterized in that, The system includes: A multi-modal data acquisition module for acquiring multi-modal data, where the multi-modal data includes text data, image data, and audio data; An image feature extraction module for performing clustering analysis on the gray values of the image data using a clustering algorithm to obtain multiple different clustering spaces, determining the stretching range of the equalized image histogram according to the proportion of the image information contained in each clustering space in different clustering spaces, completing image enhancement, and extracting the image feature representation of the enhanced image data; A feature representation alignment module for extracting the text feature representation of the text data based on the GloVe model, extracting the audio feature representation of the audio data, aligning the image feature representation, text feature representation, and audio feature representation, and fusing the aligned representations of each modality to obtain a multi-modal fused feature representation; A content security detection module for inputting the multi-modal fused feature representation into a content security detection model and outputting a content security detection result.

2. The cross-modal content security detection system based on deep learning according to claim 1, characterized in that Performing clustering analysis on the gray values of the image data using a clustering algorithm to obtain multiple different clustering spaces specifically includes the following process: Step 1: Calculation of the initial clustering center: Calculate the grayscale value of the image data and the cosine distance between : ; where n is the number of pixels in the image data; obtain the minimum value in , and define the cumulative number of the grayscale value in the image data with a distance of from the grayscale value as ; Calculate the contribution rate of the grayscale value in the image data ; : ; Arrange the contribution rates in descending order, and select the top K grayscale values with the largest contribution rates as the initial clustering centers; Step 2: Classification of gray value categories: Group the gray value closest to the cluster center into one category: For any gray value in the image data , if it satisfies the following formula, it is grouped into one category: ; among them, represents the category after m iterations , represents the cluster center with the largest contribution rate, represents the cluster center with a non - largest contribution rate; Step 3: Evaluation of the iteration accuracy: Calculate the iterative precision of each sample in different categories : ; Among them, represents the cluster center, is the weight. If belongs to the category with as the cluster center, then takes the value of 1, otherwise takes the value of 0; If the iteration accuracy does not meet the set threshold, update the clustering center; Step 4: Update of the clustering center: Solve for the means of different categories and use them as new cluster centers. Repeat steps two and three until the iteration accuracy meets the set threshold, obtaining multiple different clustering spaces. Among them, the cluster center of the m + 1-th iteration for the category is: ; among them, represents the number of gray levels in the category with as the clustering center.

3. The cross-modal content security detection system based on deep learning according to claim 1, characterized in that, Determining the stretching range of the equalized image histogram according to the proportion of the image information contained in each clustering space in different clustering spaces to complete image enhancement specifically includes the following process: Assume that the total number of gray levels in the image data P is N, and the number of gray levels contained in a certain clustering space q is , then the histogram proportionality coefficient of the image after equalization of this clustering space is , where means rounding ; Statistical histogram ratio coefficients of all clustering spaces, constructing a rectangular coordinate system with the clustering space numbers corresponding to the histogram ratio coefficients as the X-axis and the histogram ratio coefficients as the Y-axis, marking all the histogram ratio coefficients in the rectangular coordinate system in the form of points, connecting adjacent points in the rectangular coordinate system to generate a histogram ratio coefficient curve, drawing perpendicular lines from both ends of the histogram ratio coefficient curve to the X-axis to obtain two starting and ending line segments, forming a closed figure by the histogram ratio coefficient curve, the two starting and ending line segments, and the X-axis, and marking the total area of the closed figure as the stretching reference coefficient of the image data P; Defining the images with stretching reference coefficients not exceeding the preset stretching threshold as dark regions and processing them using adaptive histogram equalization, and defining the images with stretching reference coefficients exceeding the preset stretching threshold as bright regions and stretching their gray values using linear correction to complete image enhancement.

4. The cross-modal content security detection system based on deep learning according to claim 1, characterized in that Performing image features on the enhanced image data based on the VGG19 model.

5. The cross-modal content security detection system based on deep learning according to claim 1, wherein Extracting the audio feature representation based on the openSMILE toolbox.

6. The cross-modal content security detection system based on deep learning according to claim 1, wherein Aligning the image feature representation, text feature representation, and audio feature representation specifically includes the following process: The image feature representation, text feature representation, and audio feature representation are respectively projected through a projection layer , , onto the common space R. The projected features are: the image feature , the text feature , and the audio feature , where is the unprojected image feature representation, is the unprojected text feature representation, and is the unprojected audio feature representation; Making the projected features align the features of different modalities in a common space by maximizing the similarity between matching sample pairs and minimizing the similarity between non-matching sample pairs. The loss function L is as follows: ; wherein, represents the similarity between the image feature and the text feature ; represents the similarity between the text feature and the audio feature ; represents the similarity between the image feature and the audio feature ; represents learnable parameters.

7. The cross-modal content security detection system based on deep learning according to claim 1, characterized in that Inputting the multi-modal fused feature representation into a content security detection model and outputting a content security detection result specifically includes the following process: Input the fused feature representation of the cross-modal content to be detected, and construct a heterogeneous graph based on the fused feature representation; Perform contrastive learning on the heterogeneous graph and output the content security detection result.

8. The cross-modal content security detection system based on deep learning according to claim 7, wherein Constructing a heterogeneous graph based on the fused feature representation specifically includes the following process: Input: Fused feature representation u, set of dangerous content factors Vu associated with the fused feature representation u, number Mu of selected dangerous content factors in Vu, number Z of selected neighbors, where Z represents the number of neighbors selected by each node, node set N1: contains all nodes in layer L1, initially only contains the fused feature representation u; number of layers H, the number of layers H is the number of layers of the graph, indicating the number of layers extended outward from the fused feature representation u; Output the constructed heterogeneous graph: Embed the fused feature representation u into the initial graph Gu, and the initial node set N1 = {u}; Select Mu dangerous content factor information from Vu and update N1; Construct an edge u→v in Gu and mark the edge type, where v is a dangerous content factor of u; The construction process includes: for each layer L1, from the second layer 2 to H, construct a temporary node set Ntmp1 to store the nodes of the current layer, and traverse the node set Nl−1 of the previous layer: For each node i, construct a temporary node set Ntmp2. If i is a dangerous content factor node and has more than m′ interactive content factors, select m′ latest interactive content factors and add them to Ntmp2. If i is a fused feature representation node and has more than m′ interactive content factors, select m′ latest interactive content factors and add them to Ntmp2. For each neighbor node nneighbor in Ntmp2, construct an edge i→nneighbor in Gu and mark the edge type, update Ntmp1 = Ntmp1∪Ntmp2, and update the node set Nl of the current layer to Nl = Ntmp1; Use the breadth-first search algorithm to obtain a subgraph from Gu, add the subgraph to the total graph GG, update i = i + 1, and obtain the heterogeneous graph.

9. The cross-modal content security detection system based on deep learning according to claim 8, characterized in that Performing contrastive learning on the heterogeneous graph and outputting the content security detection result specifically includes the following process: Compare the heterogeneous graph with a preset heterogeneous graph to determine whether they are the same. If so, there is no security hazard content in this cross-modal content; if not, there is security hazard content in this cross-modal content.

10. A cross-modal content security detection method based on deep learning, characterized in that Applicable to the cross-modal content security detection system based on deep learning described in any one of claims 1 to 9, the method includes: Obtain multimodal data, where the multimodal data includes text data, image data, and audio data; Use a clustering algorithm to perform clustering analysis on the image gray values of the image data to obtain multiple different clustering spaces, determine the stretching range of the equalized image histogram according to the proportion of the image information contained in each clustering space in different clustering spaces, complete image enhancement, and extract the image feature representation of the enhanced image data; Extract the text feature representation of text data based on the GloVe model, extract the audio feature representation of audio data, align the image feature representation, text feature representation, and audio feature representation, and fuse the aligned representations of each modality to obtain a multi-modal fused feature representation; Input the multi-modal fused feature representation into the content security detection model and output the content security detection result.