Data compliance test method and terminal
By building a multimodal sensitive database and pre-training fine-tuning of neural networks, and combining the logistic function to calculate the probability of violation, the problem of low accuracy in identifying variant texts and metaphors in existing technologies is solved, and more efficient data compliance testing is achieved.
Patent Information
- Application Number
- CN202510507181.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-09-12
AI Technical Summary
The existing data compliance testing methods are unable to effectively identify variant texts and metaphors, resulting in a decrease in test accuracy. The existing technology is unable to effectively identify the technical problems that have not yet been solved. The existing data compliance testing methods are unable to effectively identify variant texts and metaphors, resulting in a decrease in test accuracy. In addition, manual review is highly subjective and the review standards are inconsistent.
By acquiring multimodal sensitive data from various regions to build a sensitive data map, using multimodal neural networks for pre-training and fine-tuning, combined with gaming industry data adjustments, calculating the sensitive similarity of text and images, and using the logistic function to calculate the probability of violation, compliance testing is achieved.
It improves the accuracy and coverage of data compliance testing, ensures the timeliness and comprehensiveness of testing, adapts to the multimodal data of the gaming industry, and achieves more accurate compliance judgments.
Smart Images

Figure CN120632473A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data testing, and in particular to a data compliance testing method and terminal. Background Art
[0002] The existing data compliance testing method usually establishes a harmful text library and a harmful image library, analyzes the harmful text coefficient in the text part of the data to be tested, and determines the compliance of the data based on the coefficients obtained from the analysis.
[0003] Alternatively, the compliance of the data can be obtained by inputting the data to be tested into a feature extraction model and performing similarity calculation based on the extracted features and the harmful feature library.
[0004] Therefore, currently, harmful databases are used for comparison and detection, but they cannot identify variant texts and hidden visual metaphors, resulting in a decrease in test accuracy. In addition, if manual review is used, the review is highly subjective, and there are differences in the understanding of non-compliance in different regions, which can easily lead to inconsistent review standards. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a data compliance testing method and terminal, which can improve the accuracy of data compliance testing and expand the coverage of compliance testing.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: A data compliance testing method comprises the following steps: Acquire multimodal sensitive data from various regions, construct a sensitive data map based on the multimodal sensitive data, and construct a multimodal sensitive database based on the multimodal sensitive data and the sensitive data map; Pre-training a multimodal neural network using the multimodal sensitive database, and adjusting the multimodal neural network based on gaming industry data; Inputting the test data into the multimodal neural network, and calculating the text-sensitive similarity of the text information in the test data and the image-sensitive similarity of the image information in the test data in combination with the sensitive data map of the multimodal sensitive database in the multimodal neural network; A violation probability of the data to be tested is obtained according to the image-sensitive similarity and the text-sensitive similarity, and a compliance test result of the data to be tested is obtained according to the violation probability.
[0007] In order to solve the above technical problems, another technical solution adopted by the present invention is: A data compliance test terminal comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, each step of the above-mentioned data compliance test is implemented.
[0008] The beneficial effects of the present invention are as follows: a sensitive data map is established based on the acquired multimodal sensitive data from various regions, thereby obtaining a multimodal sensitive database; a multimodal neural network is pre-trained using the multimodal sensitive database, and the neural network is fine-tuned using gaming industry data; the data to be tested is input into the multimodal neural network, and the sensitive similarity of the text information and image information in the data to be tested is calculated in combination with the sensitive data map, the violation probability of the data to be tested is obtained based on the image sensitive similarity and the text sensitive similarity, and the compliance test results of the data to be tested are obtained based on the violation probability. In this way, the accuracy of data compliance testing in game content review is improved, and the coverage of compliance testing is expanded. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 This is a flow chart of a data compliance testing method according to an embodiment of the present invention; Figure 2 Schematic diagram of a data compliance test terminal according to an embodiment of the present invention; Figure 3 This is an architectural diagram of a data compliance testing method according to an embodiment of the present invention; Figure 4 This is a review flowchart of the data compliance test according to an embodiment of the present invention.
[0010] Description of labels: 1. A data compliance testing terminal; 2. A memory; 3. A processor. DETAILED DESCRIPTION
[0011] To illustrate the technical content, achieved objectives and effects of the present invention in detail, the following description is given in conjunction with the embodiments and accompanying drawings.
[0012] Please refer to Figure 1 , an embodiment of the present invention provides a data compliance testing method, comprising the steps of: Acquire multimodal sensitive data from various regions, construct a sensitive data map based on the multimodal sensitive data, and construct a multimodal sensitive database based on the multimodal sensitive data and the sensitive data map; Pre-training a multimodal neural network using the multimodal sensitive database, and adjusting the multimodal neural network based on gaming industry data; Inputting the test data into the multimodal neural network, and calculating the text-sensitive similarity of the text information in the test data and the image-sensitive similarity of the image information in the test data in combination with the sensitive data map of the multimodal sensitive database in the multimodal neural network; A violation probability of the data to be tested is obtained according to the image-sensitive similarity and the text-sensitive similarity, and a compliance test result of the data to be tested is obtained according to the violation probability.
[0013] From the above description, it can be seen that the beneficial effects of the present invention are: establishing a sensitive data map based on the multimodal sensitive data obtained from various regions, and then obtaining a multimodal sensitive database; using the multimodal sensitive database to pre-train a multimodal neural network, and using game industry data to fine-tune the neural network; inputting the data to be tested into the multimodal neural network, and combining the sensitive data map to calculate the sensitive similarity of the text information and image information in the data to be tested, respectively, and obtaining the violation probability of the data to be tested based on the image sensitive similarity and the text sensitive similarity, and obtaining the compliance test results of the data to be tested through the violation probability. In this way, the accuracy of data compliance testing in game content review is improved, and the coverage of compliance testing is expanded.
[0014] Furthermore, multimodal sensitive data of each region is obtained, a sensitive data map is constructed based on the multimodal sensitive data, and a multimodal sensitive database is constructed based on the multimodal sensitive data and the sensitive data map, including: Acquire multimodal sensitive data of various regional cultures in real time through crawlers, including text sensitive information and image sensitive information; Performing feature extraction on the image sensitive information to obtain a first sensitive concept and its sensitive attribute, and performing natural language processing on the text sensitive information to obtain a second sensitive concept and its sensitive attribute; Taking the first sensitive concept and the second sensitive concept as nodes, labeling the nodes with corresponding sensitive attributes, and adding corresponding regional information to the nodes; Establishing semantic relationships between nodes based on the region information corresponding to each node, and using a graph database to manage the nodes and their semantic relationships to obtain a sensitive data map; A multimodal sensitive database is constructed based on the image sensitive information, the text sensitive information and the sensitive data map.
[0015] From the above description, we can see that the crawler can obtain cultural multimodal sensitive data of various regions in real time, and extract sensitive concepts and sensitive attributes from each modality of sensitive data, and then establish a sensitive data map to facilitate subsequent cross-module data comparison.
[0016] Furthermore, the multimodal neural network is adjusted according to the gaming industry data, including: Adjusting and training the text encoder in the multimodal neural network using game dialogue data and game description text from game industry data; Adjusting and training the image encoder in the multimodal neural network using game screenshot data and game 3D model data from the game industry data; The contrastive learning module in the multimodal neural network is adjusted and trained using game screenshot data and game description text in the game industry data.
[0017] From the above description, it can be seen that by fine-tuning using game industry data, the multimodal neural network can better understand the specific terms and context in game content, improve the accuracy and applicability of the review, and the fine-tuned multimodal neural network can better process the multimodal data in the game, achieve more accurate compliance judgments, and ensure the timeliness and comprehensiveness of the multimodal neural network.
[0018] Furthermore, combining the sensitive data map of the multimodal sensitive database in the multimodal neural network, calculating the text-sensitive similarity of the text information in the test data and the image-sensitive similarity of the image information in the test data, including: Acquiring first region information of the data to be tested; In the multimodal neural network, an image sensitivity similarity between the image information in the test data and the image sensitive information corresponding to the first region information in the multimodal sensitive database is calculated, and a text sensitivity similarity between the text information in the test data and the text sensitive information corresponding to the first region information in the multimodal sensitive database is calculated; According to the association between the image information and the text information in the sensitive data map, the image sensitivity similarity and the text sensitivity similarity are adjusted: Query the shortest association path between text concepts and image concepts through the graph database; When the path length does not exceed the preset length, the similarity is corrected according to the association weight W. The calculation formula of the association weight is: W=α×(1 / d)+β×C+γ×S; Where d is the path length, C is the normalized co-occurrence frequency, S is the graph embedding semantic similarity; α is the path length coefficient, β is the frequency coefficient, and γ is the semantic similarity coefficient; Corrected image-sensitive similarity S ' img and text-sensitive similarity S ' text for: S ' text=min( S text +W×δ,ζ); S ' img =min( S img +W×δ,ζ); Where δ is the adjustment coefficient and ζ is the upper limit of similarity.
[0019] From the above description, it can be seen that after calculating the sensitive similarity of each mode of the data to be tested, the sensitive similarity is further adjusted according to the correlation relationship in the sensitive data map, so as to ensure more accurate compliance judgment.
[0020] Furthermore, obtaining the violation probability of the data to be tested according to the image-sensitive similarity and the text-sensitive similarity includes:
[0021] Where, P represents the probability of violation, k A constant that represents the steepness of the curve, S img represents image sensitive similarity, S text Indicates text sensitive similarity, θ Indicates the violation threshold.
[0022] From the above description, it can be seen that calculating the violation probability through the logistic function makes the probability calculation process computationally inexpensive and easy to implement.
[0023] Please refer to Figure 2 Another embodiment of the present invention provides a data compliance testing terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the various steps of the above-mentioned data compliance test are implemented.
[0024] The data compliance testing method and terminal described above are suitable for improving the accuracy of data compliance testing and expanding the coverage of compliance testing. The following describes a specific implementation method: Please refer to Figure 1 and Figure 4 , embodiment 1 of the present invention is: A data compliance testing method comprises the following steps: S1. Acquire multimodal sensitive data from various regions, construct a sensitive data map based on the multimodal sensitive data, and construct a multimodal sensitive database based on the multimodal sensitive data and the sensitive data map.
[0025] S11. Acquire multimodal sensitive data of the culture of each region in real time through a crawler, wherein the multimodal sensitive data includes text sensitive information and image sensitive information.
[0026] Specifically, crawler technology is used to obtain cultural multimodal sensitive data from various regions in real time. The attributes of sensitive data include religious taboos, historical events, regional taboos, etc., and the data modalities include text and images.
[0027] The collected data is cleaned and standardized to remove noise and redundant information, and natural language processing (NLP) operations such as word segmentation, part-of-speech tagging, and entity recognition are performed on sensitive text information.
[0028] In some embodiments, the results of text sensitive information after being translated into variants, homophones, or multiple languages may also be stored, and the results of image sensitive information after being rotated may also be stored.
[0029] S12. Perform feature extraction on the image sensitive information to obtain a first sensitive concept and its sensitive attribute, and perform natural language processing on the text sensitive information to obtain a second sensitive concept and its sensitive attribute.
[0030] S13. Take the first sensitive concept and the second sensitive concept as nodes, label the nodes with corresponding sensitive attributes, and add corresponding regional information to the nodes.
[0031] Specifically, sensitive concepts in image sensitive information and text sensitive information are extracted, sensitive concepts (such as "pig") are used as nodes, their attributes (such as "religious taboos") are labeled, and related regional information (such as "Middle East") is added to the nodes.
[0032] S14. Establish semantic relationships between nodes based on the region information corresponding to each node, and use a graph database to manage the nodes and their semantic relationships to obtain a sensitive data map.
[0033] Specifically, semantic analysis technology is used to establish semantic relationships between concept nodes (such as the relationship between "pig" and "religious taboo"), and a graph database (such as Neo4j) is used to store and manage these nodes and relationships to obtain a sensitive data graph.
[0034] In some embodiments, the graph is optimized through machine learning algorithms (such as graph embedding) to improve query efficiency and accuracy, and the graph is updated regularly to ensure its timeliness and comprehensiveness.
[0035] S15. Construct a multimodal sensitive database based on the image sensitive information, the text sensitive information, and the sensitive data map.
[0036] Specifically, the data dimensions of the multimodal sensitive database in this embodiment are shown in Table 1.
[0037] Table 1 Data dimensions of the multimodal sensitive database
[0038] S2. Pre-train a multimodal neural network using the multimodal sensitive database, and adjust the multimodal neural network according to gaming industry data.
[0039] S21. Adjust and train the text encoder in the multimodal neural network using game dialogue data and game description text in game industry data.
[0040] Specifically, we fine-tuned the text encoder to better understand the specific terminology and context of the gaming industry. We trained the model on gaming industry data (such as game dialogues and descriptions) to improve the model's semantic understanding of game text.
[0041] S22. Adjust and train the image encoder in the multimodal neural network using game screenshot data and game 3D model data in the game industry data.
[0042] Specifically, we fine-tune the image encoder to more accurately identify visual elements in games (such as characters, scenes, and props). We also train the model on gaming industry data (such as screenshots and 3D models) to improve its ability to recognize game images.
[0043] S23. Adjust and train the contrastive learning module in the multimodal neural network using game screenshot data and game description text in the game industry data.
[0044] Specifically, we fine-tuned the contrastive learning module to better adapt it to the multimodal data of the gaming industry (such as the association between text and images). We also trained on gaming industry data (such as game screenshots and their corresponding descriptions) to improve the model's ability to analyze associations in multimodal data.
[0045] Therefore, by fine-tuning the model using gaming industry data, it can better understand the specific terminology and context within gaming content, improving the accuracy and applicability of reviews. The fine-tuned model can better handle multimodal data in games (such as the correlation between text and images), enabling more precise compliance assessments. The fine-tuning process supports dynamic updates, enabling timely adaptation to changes in the gaming industry and emerging sensitive content, ensuring the model's timeliness and comprehensiveness.
[0046] S3. Input the data to be tested into the multimodal neural network, and calculate the text-sensitive similarity of the text information in the data to be tested and the image-sensitive similarity of the image information in the data to be tested in combination with the sensitive data map of the multimodal sensitive database in the multimodal neural network.
[0047] Specifically, first region information of the data to be tested is obtained, and in the multimodal neural network, image sensitivity similarity between image information in the data to be tested and image sensitive information corresponding to the first region information in the multimodal sensitive database is calculated, and text sensitivity similarity between text information in the data to be tested and text sensitive information corresponding to the first region information in the multimodal sensitive database is calculated; The image sensitive similarity and the text sensitive similarity are adjusted according to the association relationship between the image information and the text information in the sensitive data map.
[0048] In this embodiment, the text of the data to be tested in the game resource file is first compared with the text in the sensitive database, and the similarity between the text and the sensitive text database is calculated. S text ; Compare the image of the test data with the image in the sensitive database and calculate the similarity between them S img .
[0049] Perform natural language processing (NLP) on the text of the test data to extract key entities and concepts, search for these entities and concepts in the graph, obtain their attributes and associations, and determine the sensitive content and related regions that may be involved in the text based on the search results.
[0050] By querying the sensitive data map, semantic association analysis between text and images can be achieved. Based on the query results, the similarity between text and images can be adjusted to determine their compliance.
[0051] For example, if the text and image are highly semantically related (e.g. the text description "celebration" and the image "wine glass" appear at the same time), the probability of violation will increase significantly. Through the sensitive data map, further adjustments can be made. S img and S text weights to ensure that semantic associations are fully considered.
[0052] According to the association between the image information and the text information in the sensitive data map, the image sensitivity similarity and the text sensitivity similarity are adjusted, specifically: Query the shortest association path between text concepts and image concepts through the graph database; When the path length is ≤3, the similarity is corrected according to the association weight W. The calculation formula of the association weight is: W=α×(1 / d)+β×C+γ×S Where d is the path length, C is the normalized co-occurrence frequency, S is the graph embedding semantic similarity; α is the path length coefficient, β is the frequency coefficient, γ is the semantic similarity coefficient, and α + β + γ = 1. Preferably, α = 0.6, β = 0.3, and γ = 0.1.
[0053] The corrected similarity satisfies: S ' text =min( S text +W×δ, ζ) and S ' img =min( S img +W×δ, ζ) Wherein, δ is the adjustment coefficient, preferably 0.2; ζ is the upper limit of similarity, preferably 0.95.
[0054] Finally, according to the adjusted S ' img and S ' text , calculate the comprehensive violation probability and determine whether the content is compliant.
[0055] S4. Obtain a violation probability of the data to be tested according to the image-sensitive similarity and the text-sensitive similarity, and obtain a compliance test result of the data to be tested according to the violation probability.
[0056] Specifically, the calculation formula for violation probability is as follows:
[0057] Where, P represents the probability of violation, k A constant that represents the steepness of the curve, S img represents image sensitive similarity, S text Indicates text sensitive similarity, θ Represents the violation judgment threshold. This formula is actually a form of logic function, called logistic function, which maps the linear combination of inputs to a probability value in the (0,1) interval through a nonlinear function. S img + S text Less than θ When , the violation probability is close to 0; when S img + Stext Much greater than θ When , the violation probability is close to 1.
[0058] Set a violation threshold. If the violation probability is greater than the violation threshold, a risk report is generated. Otherwise, the data to be tested is marked as compliant.
[0059] Specifically, the violation threshold for strict areas is 0.74, the violation threshold for ordinary areas is 0.82, and the violation threshold for loose areas is 0.90; In addition, the violation threshold θ' is adjusted monthly based on the proportion of newly added sensitive patterns DR: θ'=θ0×(1+0.1×(DR-5%)) Where θ0 is the optimal threshold determined by confusion matrix analysis, preferably, θ0 = 0.82; When the similarity difference between the image and the text exceeds 0.3, the threshold is reduced by 0.05×| S img - S text |.
[0060] Please refer to Figure 3 In this embodiment, the resource parser unpacks game files and separates resources such as textures, models, and text. The multimodal database stores sensitive image features, text lexicons, and semantic graphs. The CLIP analysis engine performs cross-modal image / text comparisons. The report generator outputs a risk-level audit report. Specifically, the resource parser inputs data into the CLIP analysis engine, which then compares it with the multimodal database. The report generator then summarizes the results.
[0061] The application scenarios of this embodiment are as follows: (1) Game version review in the Middle East Detection target: Alcohol-related elements (icons, text descriptions, 3D models) Image detection: Recognize the shape of a wine bottle (even if the image is replaced with a juice package); Text detection: Arabic slang words with implicit alcohol connotations (such as " "literally "soul drink"); Correlation analysis: The simultaneous appearance of "glass cup + liquid splashing effect" triggers an in-depth review.
[0062] (2) Historical strategy game review Detection target: Inappropriate historical representations (such as names of disputed regions) Map texture analysis: Compare the boundary coordinates of sensitive areas with in-game map data; Event description review: Detect biased vocabulary in narrative texts (e.g., "since ancient times" → "disputed area").
[0063] Please refer to Figure 2 , the second embodiment of the present invention is: A data compliance testing terminal 1 includes a memory 2, a processor 3, and a computer program stored in the memory 2 and executable on the processor 3. When the processor 3 executes the computer program, each step of a data compliance testing method of embodiment 1 is implemented.
[0064] In summary, the present invention provides a data compliance testing method and terminal that uses a crawler to obtain updated culturally sensitive data from various regions in real time, automatically expanding the database to ensure its timeliness and comprehensiveness. By constructing a sensitive data map, semantic association analysis between text and images is achieved, further improving the accuracy and coverage of detection. Based on the pre-trained CLIP model, fine-tuning is performed using gaming industry data to improve the model's accuracy and applicability in game content review. By defining a logistic function, the combined violation probability of images and text is calculated, achieving more accurate compliance judgments.
[0065] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's description and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A data compliance testing method, characterized in that: Including steps: Acquire multimodal sensitive data from various regions, construct a sensitive data map based on the multimodal sensitive data, and construct a multimodal sensitive database based on the multimodal sensitive data and the sensitive data map; Pre-training a multimodal neural network using the multimodal sensitive database, and adjusting the multimodal neural network based on gaming industry data; Inputting the test data into the multimodal neural network, and calculating the text-sensitive similarity of the text information in the test data and the image-sensitive similarity of the image information in the test data in combination with the sensitive data map of the multimodal sensitive database in the multimodal neural network; A violation probability of the data to be tested is obtained according to the image-sensitive similarity and the text-sensitive similarity, and a compliance test result of the data to be tested is obtained according to the violation probability.
2. A data compliance testing method according to claim 1, characterized in that: Acquiring multimodal sensitive data from various regions, constructing a sensitive data map based on the multimodal sensitive data, and constructing a multimodal sensitive database based on the multimodal sensitive data and the sensitive data map, including: Acquire multimodal sensitive data of various regional cultures in real time through crawlers, including text sensitive information and image sensitive information; Performing feature extraction on the image sensitive information to obtain a first sensitive concept and its sensitive attribute, and performing natural language processing on the text sensitive information to obtain a second sensitive concept and its sensitive attribute; Taking the first sensitive concept and the second sensitive concept as nodes, labeling the nodes with corresponding sensitive attributes, and adding corresponding regional information to the nodes; Establishing semantic relationships between nodes based on the region information corresponding to each node, and using a graph database to manage the nodes and their semantic relationships to obtain a sensitive data map; A multimodal sensitive database is constructed based on the image sensitive information, the text sensitive information and the sensitive data map.
3. A data compliance testing method according to claim 1, characterized in that: The multimodal neural network was tuned based on gaming industry data, including: Adjusting and training the text encoder in the multimodal neural network using game dialogue data and game description text from game industry data; Adjusting and training the image encoder in the multimodal neural network using game screenshot data and game 3D model data from the game industry data; The contrastive learning module in the multimodal neural network is adjusted and trained using game screenshot data and game description text in the game industry data.
4. A data compliance testing method according to claim 2, characterized in that: Calculating the text-sensitive similarity of the text information in the test data and the image-sensitive similarity of the image information in the test data in combination with the sensitive data map of the multimodal sensitive database in the multimodal neural network, including: Acquiring first region information of the data to be tested; In the multimodal neural network, an image sensitivity similarity between the image information in the test data and the image sensitive information corresponding to the first region information in the multimodal sensitive database is calculated, and a text sensitivity similarity between the text information in the test data and the text sensitive information corresponding to the first region information in the multimodal sensitive database is calculated; According to the association between the image information and the text information in the sensitive data map, the image sensitivity similarity and the text sensitivity similarity are adjusted: Query the shortest association path between text concepts and image concepts through the graph database; When the path length does not exceed the preset length, the similarity is corrected according to the association weight W. The calculation formula of the association weight is: W=α×(1 / d)+β×C+γ×S; Where d is the path length, C is the normalized co-occurrence frequency, S is the graph embedding semantic similarity; α is the path length coefficient, β is the frequency coefficient, and γ is the semantic similarity coefficient; Corrected image-sensitive similarity S ' img and text-sensitive similarity S ' text for: S ' text =min( S text +W×δ, ζ); S ' img =min( S img +W×δ, ζ); Where δ is the adjustment coefficient and ζ is the upper limit of similarity.
5. A data compliance testing method according to claim 1, characterized in that: Obtaining the violation probability of the data to be tested according to the image sensitive similarity and the text sensitive similarity includes: Where, P represents the probability of violation, k A constant that represents the steepness of the curve, S img represents image sensitive similarity, S text Indicates text-sensitive similarity, θ Indicates the violation threshold.
6. A data compliance testing terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the following steps are implemented: Acquire multimodal sensitive data from various regions, construct a sensitive data map based on the multimodal sensitive data, and construct a multimodal sensitive database based on the multimodal sensitive data and the sensitive data map; Pre-training a multimodal neural network using the multimodal sensitive database, and adjusting the multimodal neural network based on gaming industry data; Inputting the test data into the multimodal neural network, and calculating the text-sensitive similarity of the text information in the test data and the image-sensitive similarity of the image information in the test data in combination with the sensitive data map of the multimodal sensitive database in the multimodal neural network; A violation probability of the data to be tested is obtained according to the image-sensitive similarity and the text-sensitive similarity, and a compliance test result of the data to be tested is obtained according to the violation probability.
7. A data compliance testing terminal according to claim 6, characterized in that: Acquiring multimodal sensitive data from various regions, constructing a sensitive data map based on the multimodal sensitive data, and constructing a multimodal sensitive database based on the multimodal sensitive data and the sensitive data map, including: Acquire multimodal sensitive data of various regional cultures in real time through crawlers, including text sensitive information and image sensitive information; Performing feature extraction on the image sensitive information to obtain a first sensitive concept and its sensitive attribute, and performing natural language processing on the text sensitive information to obtain a second sensitive concept and its sensitive attribute; Taking the first sensitive concept and the second sensitive concept as nodes, labeling the nodes with corresponding sensitive attributes, and adding corresponding regional information to the nodes; Establishing semantic relationships between nodes based on the region information corresponding to each node, and using a graph database to manage the nodes and their semantic relationships to obtain a sensitive data map; A multimodal sensitive database is constructed based on the image sensitive information, the text sensitive information and the sensitive data map.
8. A data compliance testing terminal according to claim 6, characterized in that: The multimodal neural network was tuned based on gaming industry data, including: Adjusting and training the text encoder in the multimodal neural network using game dialogue data and game description text from game industry data; Adjusting and training the image encoder in the multimodal neural network using game screenshot data and game 3D model data from the game industry data; The contrastive learning module in the multimodal neural network is adjusted and trained using game screenshot data and game description text in the game industry data.
9. A data compliance testing terminal according to claim 7, characterized in that: Calculating the text-sensitive similarity of the text information in the test data and the image-sensitive similarity of the image information in the test data in combination with the sensitive data map of the multimodal sensitive database in the multimodal neural network, including: Acquiring first region information of the data to be tested; In the multimodal neural network, an image sensitivity similarity between the image information in the test data and the image sensitive information corresponding to the first region information in the multimodal sensitive database is calculated, and a text sensitivity similarity between the text information in the test data and the text sensitive information corresponding to the first region information in the multimodal sensitive database is calculated; According to the association between the image information and the text information in the sensitive data map, the image sensitivity similarity and the text sensitivity similarity are adjusted: Query the shortest association path between text concepts and image concepts through the graph database; When the path length does not exceed the preset length, the similarity is corrected according to the association weight W. The calculation formula of the association weight is: W=α×(1 / d)+β×C+γ×S; Where d is the path length, C is the normalized co-occurrence frequency, S is the graph embedding semantic similarity; α is the path length coefficient, β is the frequency coefficient, and γ is the semantic similarity coefficient; Corrected image-sensitive similarity S ' img and text-sensitive similarity S ' text for: S ' text =min( S text +W×δ, ζ); S ' img =min( S img +W×δ, ζ); Where δ is the adjustment coefficient and ζ is the upper limit of similarity.
10. A data compliance testing terminal according to claim 6, characterized in that: Obtaining the violation probability of the data to be tested according to the image sensitive similarity and the text sensitive similarity includes: Where, P represents the probability of violation, k A constant that represents the steepness of the curve, S img represents image sensitive similarity, S text Indicates text-sensitive similarity, θ Indicates the violation threshold.