Risk assessment method and device for digital content
Through multi-channel feature extraction and cross-modal fusion technology, combined with the decision forest model and weight configuration rules, the problem of missed detection of cross-modal combination risks in multimedia security analysis is solved, the deep fusion and collaborative analysis of multimodal content is achieved, and the accuracy and reliability of the audit are improved.
Patent Information
- Application Number
- CN202510844327.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-03
AI Technical Summary
Existing multimedia security analysis technologies are unable to effectively identify cross-modal combination risks, resulting in missed detection of multimodal illegal content, especially variant attacks in text and visual fields that are difficult to identify.
Multi-channel feature extraction technology is used to extract structured features from visual, audio and text modalities, and the features are aligned and fused through cross-modal fusion technology. The decision forest model is used for multimodal risk assessment, and a comprehensive risk score is generated in combination with weight configuration rules.
It has achieved deep integration and collaborative analysis of multimodal information, significantly improved the ability to identify complex variant illegal content, and enhanced the accuracy and reliability of digital content security audits.
Smart Images

Figure CN120751184A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of large language models, and in particular to a risk assessment method and device for digital content. Background Art
[0002] The explosive growth of the current digital content industry poses significant challenges to multimedia security analysis technology. Mainstream audit systems generally employ rule-based filtering mechanisms, relying on predefined sensitive word libraries and image feature templates for content matching. While these solutions offer basic identification capabilities for overtly illegal content, they suffer from fundamental flaws when faced with increasingly sophisticated variant attacks.
[0003] In text scenarios, circumvention methods such as homophonic replacement, pinyin abbreviations, and metaphorical expressions can easily bypass keyword detection; in the visual field, illegal images that have been processed with local blurring and color shifting are difficult to be captured by static templates; more importantly, the single-modal independent analysis model is completely unable to cope with cross-modal combination risks, such as normal videos with banned audio lyrics, or harmless text descriptions with suggestive illustrations. Multi-modal illegal content is almost completely missed by existing systems. Summary of the Invention
[0004] The present application provides a risk assessment method and device for digital content, which is used to solve the audit problems caused by the missed detection of combined risks and the unexplainable decision-making process in the current multimodal content audit.
[0005] In a first aspect, the present application provides a risk assessment method for digital content, comprising: Use multi-channel feature extraction technology to extract structural features of digital content; Using cross-modal fusion technology to align and fuse structured features to determine the fusion features after the fusion of structured features; Based on the preset decision forest model, a multimodal risk assessment is performed on the fusion features to determine the modal risk score corresponding to each modality; Based on the preset weight configuration rules, the risk score of each modality is calculated to determine the comprehensive risk score corresponding to the digital content.
[0006] In a second aspect, the present application provides a risk assessment device for digital content, comprising: a structured feature determination module configured to extract structured features of digital content using a multi-channel feature extraction technique; A fusion feature determination module is configured to align and fuse the structured features using a cross-modal fusion technology to determine a fusion feature after the structured features are fused; A modality risk score determination module is configured to perform a multimodal risk assessment on the fused features based on a preset decision forest model to determine a modality risk score corresponding to each modality; The comprehensive risk score determination module is configured to calculate the risk score of each modality based on a preset weight configuration rule to determine the comprehensive risk score corresponding to the digital content.
[0007] In a third aspect, the present application provides a readable medium comprising execution instructions. When a processor of an electronic device executes the execution instructions, the electronic device executes any method described in the first aspect.
[0008] In a fourth aspect, the present application provides an electronic device comprising a processor and a memory storing execution instructions. When the processor executes the execution instructions stored in the memory, the processor executes any method described in the first aspect.
[0009] This application provides a risk assessment method for digital content. It uses multi-channel feature extraction technology to extract structured features from three modalities: vision, audio, and text. It then fuses the structured features at the feature level based on a preset fusion strategy to generate a unified fusion feature. It introduces a decision forest model composed of multiple sub-classifiers to analyze the fusion features in terms of display content, legal risks, ethical norms, and other dimensions, and outputs a modal risk score. It also performs weighted integration of the scores of each modality based on the weight configuration rules obtained from historical data statistics to generate a quantifiable comprehensive risk score. This method achieves deep fusion and collaborative analysis of multimodal information, effectively improves the ability to identify complex variant illegal content, solves the problem of missed detection in single modality analysis, and significantly enhances the accuracy and reliability of digital content security audits.
[0010] The further effects of the above-mentioned non-conventional preferred embodiment will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application or the existing technical solutions, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0012] Figure 1 A flowchart of a risk assessment method for digital content provided in one embodiment of the present application; Figure 2 A flowchart of another risk assessment method for digital content provided in one embodiment of the present application; Figure 3 A flowchart of another risk assessment method for digital content provided in one embodiment of the present application; Figure 4 This is a schematic structural diagram of a risk assessment device for digital content provided by an embodiment of the present application; Figure 5 A schematic structural diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0013] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0014] The explosive growth of the current digital content industry poses significant challenges to multimedia security analysis technology. Mainstream audit systems generally employ rule-based filtering mechanisms, relying on predefined sensitive word libraries and image feature templates for content matching. While these solutions offer basic identification capabilities for overtly illegal content, they suffer from fundamental flaws when faced with increasingly sophisticated variant attacks.
[0015] In text scenarios, circumvention methods such as homophonic replacement, pinyin abbreviations, and metaphorical expressions can easily bypass keyword detection; in the visual field, illegal images that have been processed with local blurring and color shifting are difficult to be captured by static templates; more importantly, the single-modal independent analysis model is completely unable to cope with cross-modal combination risks, such as normal videos with banned audio lyrics, or harmless text descriptions with suggestive illustrations. Multi-modal illegal content is almost completely missed by existing systems.
[0016] To address this issue, the present invention proposes a risk assessment method for digital content, which aims to address the current multimodal content review issues such as missed combined risk detection and unexplainable decision-making processes caused by the fragmentation of cross-modal features. In this embodiment, a risk assessment method for digital content includes: Step 101: Extract structural features of digital content using multi-channel feature extraction technology.
[0017] Digital content includes visual channels, audio channels and text channels. The multi-channel feature extraction technology is used to extract the structured features of the digital content, including: using image processing technology to extract image motion features and color distribution features corresponding to the visual channel; using spectrum analysis technology to extract Mel spectrum features and cepstral outlier features corresponding to the audio channel; using semantic analysis technology to extract semantic graph structure features and topic distribution features corresponding to the text channel; and determining structured features based on image motion features and color distribution features, Mel spectrum features and cepstral outlier features, semantic graph structure features and topic distribution features.
[0018] Digital content refers to electronic information entities stored, transmitted, and presented through binary encoding. Its core characteristic is that it can be parsed by computing devices and converted into multimodal signals that users can perceive. Technically, it encompasses three main categories: static content (such as text, images, and vector graphics), time-series streaming content (such as audio, video, and live data streams), and interactive content (such as web pages, games, and virtual reality scenes).
[0019] Multi-channel feature extraction technology uses three independent, parallel processing pipelines to extract key risk features from visual, audio, and text modalities. In the visual channel, the system first analyzes the motion patterns of objects in a video frame sequence. By calculating the brightness gradient of pixels between adjacent frames, it constructs an optical flow tensor, or image motion feature, that describes the direction and speed of the object's movement. This image motion feature can quantify and track sudden violent actions or unusual displacements within the image.
[0020] The system also converts video frames to a hue-saturation-value color space and generates a color distribution histogram. This is done by measuring the degree of difference between the current frame's color distribution and a safe baseline template. This is combined with the results of optical flow mutation detection (triggered when an object's speed exceeds a preset threshold) to generate a color distribution signature. This combined analysis of object motion trajectory and color mutations allows for the precise identification of sudden, bloody patches or unusually moving objects.
[0021] For audio channels, the system uses nonlinear spectrum conversion technology that conforms to the auditory characteristics of the human ear: the original sound signal is first converted into a spectral energy distribution, and then frequency compression is performed through a Mel-scale filter group that simulates the auditory sensitivity of the human ear, and finally a Mel-scale spectrum feature is generated.
[0022] To detect hidden illegal audio, the system dynamically analyzes fluctuations in sound features within a time window, calculates the degree to which the sound feature values deviate from the historical average, and uses the standard deviation as a normalization benchmark. Ultimately, the maximum deviation within the window is selected as the cepstral outlier feature. This method is effective in capturing prohibited vocalizations or sudden explosions masked by background music.
[0023] In text channel processing, the system constructs a semantic relationship network through grammatical structure analysis. Using natural language processing techniques, it identifies the grammatical dependencies between the subject, predicate, and object in a sentence. Using entity words as nodes and grammatical relationships as edges, it forms a weighted directed semantic graph, known as a semantic graph structural feature. This graph can reveal the implicit connections of evasive code words.
[0024] At the same time, a topic probability model is used to analyze the text content to generate probability distribution vectors representing different risk topics, namely topic distribution features. This vector can capture the semantic risk association pattern across sentences.
[0025] These features not only independently carry risk signals for each modality, but their collaborative design also provides a basis for subsequent cross-modal verification: for example, the sudden acceleration of an object can be corroborated by the sound of an explosion, while the distribution of text topics can help interpret the meaning of sensitive symbols in the image.
[0026] Step 102: Use cross-modal fusion technology to align and fuse the structured features to determine the fusion features after the structured features are fused.
[0027] Because visual, audio, and text data differ in dimensionality, scale, and semantic expression in their original feature spaces, directly combining them can easily lead to feature redundancy and semantic conflicts, which in turn affect the model's discriminative capabilities. Therefore, it is necessary to design a fusion mechanism that allows features from different modalities to be collaboratively expressed in a unified semantic space.
[0028] To this end, a feature mapping and alignment strategy can be employed to map image motion and color distribution features, mel-spectrogram features and cepstral outlier features, semantic graph structure features, and topic distribution features into a pre-defined unified feature space. During this mapping process, features from each modality are normalized to eliminate differences in numerical scale and distribution morphology between modalities. Furthermore, to reduce interference caused by redundant information, feature dimensions are compressed through dimensionality reduction, thereby improving fusion efficiency and enhancing model generalization.
[0029] After completing modality alignment and preprocessing, a fusion strategy can be introduced to perform feature-level fusion on the mapped modal features. This fusion strategy can be combined with linear transformations, projection matrices, attention mechanisms, or other deep fusion networks, and can be flexibly configured according to actual application requirements.
[0030] Step 103: Based on the preset decision forest model, a multimodal risk assessment is performed on the fusion features to determine the modal risk score corresponding to each modality.
[0031] The fused feature, as a unified expression vector that integrates multimodal information such as vision, audio, and text, needs to be further input into the decision model for risk assessment. To this end, a multimodal risk scoring mechanism based on a preset decision forest model can be used to independently analyze and quantitatively express potential risk factors in different modalities. Based on the attributes of the feature vector, the system inputs it into a dynamically constructed decision forest. The forest consists of multiple structured branches, each of which corresponds to a different risk discriminator, processing different types of potential risk features respectively. The model has good scalability and interpretability, and can mine high-dimensional risk patterns in multimodal data through classification paths.
[0032] After the fused features are input into the decision forest, each decision tree undergoes a series of judgments based on feature partitioning criteria. These partitioning criteria are typically set based on the features' ability to distinguish sample risk outcomes in different modalities. For example, a decision forest model has multiple subtrees: Tree 1 is a display content classifier, which primarily identifies whether visual elements such as images and videos contain illegal features; Tree 2 is a legal risk classifier, which focuses on detecting whether content involves specific legal risk areas; and Tree 3 is an ethics assessor, which primarily evaluates potential ethical conflicts in text or multimodal content.
[0033] Based on the unified fusion feature input, these sub-classifiers output risk scores for the corresponding modality or dimension. After receiving the fusion feature, each decision tree independently performs splitting, judgment, and classification path calculation, ultimately outputting its own risk result, namely the modality risk score.
[0034] Step 104: Calculate the risk score of each modality based on the preset weight configuration rules to determine the comprehensive risk score corresponding to the digital content.
[0035] Based on historical risk score information, the weight parameters corresponding to each modality are determined; based on the risk scores and weight parameters of each modality, a weighted calculation is performed to determine the comprehensive risk score.
[0036] To accurately assess the overall risk level of digital content, after completing the risk scores for each modality, it is necessary to further integrate these modality scores into a unified, comprehensive risk score. To this end, the system introduces a weighted scoring mechanism based on preset rules. This applies differentiated weights to each modality's risk score based on its risk contribution in specific application scenarios, effectively aggregating risk information.
[0037] The setting of weight parameters does not rely on human experience, but rather on statistical learning and dynamic configuration based on historical risk scoring information. This historical data is derived from the model's extensive scoring records and final processing results from actual audit tasks. Through statistical analysis, the importance of each risk dimension in the decision-making process can be quantified. By comparing the deviation distribution between actual risk results and modal scores, the system can derive the weight parameters for each modality.
[0038] After determining the weighting parameters, the system combines the risk scores of each modality with their corresponding weights to form a comprehensive risk score for the digital content. This weighting process can employ linear combinations or incorporate nonlinear strategies such as exponential adjustments to enhance sensitivity to hidden risks. This comprehensive score not only reflects the overall strength of risk signals across modalities but also, to a certain extent, reflects the cross-influence between risks, ensuring a more representative and practical assessment result.
[0039] After generating a comprehensive risk score, the structural information and content summary information of the digital content can also be extracted to generate basic report information; based on the comprehensive risk score, the risk level corresponding to the digital content can be determined; and based on the comprehensive risk level and basic report information, a risk assessment report corresponding to the digital content can be generated.
[0040] After calculating the comprehensive risk score for digital content, the system proceeds to the risk assessment report generation phase. The core purpose of this phase is to convert the model's quantitative output into readable and interpretable textual risk results for manual review, compliance filing, or automated response mechanisms.
[0041] The system first extracts structural information and summary information related to digital content as the foundation of the report. Structural information includes metadata such as the source platform, publishing terminal, content type (such as images, short videos, live broadcast clips, etc.), publication time, and content length. Summary information uses natural language processing technology to extract core semantic fragments from text descriptions, audio transcriptions, or video subtitles, quickly presenting the content's themes and key points.
[0042] Based on the calculated comprehensive risk score, the system categorizes digital content into different risk levels, such as low risk, medium risk, and high risk, according to pre-set risk grading criteria. Risk grading not only considers the absolute value of the score but also incorporates dynamic parameters such as historical score distribution, platform risk tolerance thresholds, and current review policies for correction, resulting in a grading judgment that is more tailored to business realities.
[0043] The system aggregates structural information, content summaries, comprehensive risk scores, modality score details, and determined risk levels, and outputs them as structured or natural language reports based on different application scenarios. These reports can include visual charts (such as risk histograms for each modality), location information for high-risk segments, and score trends, helping reviewers quickly grasp the overall risk profile of the content. The system also supports inter-system interface calls to trigger subsequent processing processes such as automatic blocking, rate limiting, and manual review.
[0044] Through the above technical solution, it can be seen that the beneficial effects of this embodiment are: The embodiment of the present application provides a risk assessment method for digital content. It uses multi-channel feature extraction technology to extract the structured features of digital content; uses cross-modal fusion technology to align and fuse the structured features to determine the fusion features after the fusion of the structured features; based on a preset decision forest model, it performs a multi-modal risk assessment on the fusion features to determine the modal risk score corresponding to each modality; based on a preset weight configuration rule, it calculates the risk score of each modality to determine the comprehensive risk score corresponding to the digital content. It achieves deep fusion and collaborative analysis of multimodal information, effectively improves the ability to identify complex variant illegal content, solves the problem of missed detection in single modality analysis, and significantly enhances the accuracy and reliability of digital content security audits.
[0045] Figure 1 What is shown is only a basic embodiment of a risk assessment method for digital content of the present application. By performing certain optimization and expansion on this basis, other preferred embodiments of a risk assessment method for digital content can be obtained.
[0046] like Figure 2 FIG. 1 shows another specific embodiment of a risk assessment method for digital content according to the present application.
[0047] In this embodiment, a risk assessment method for digital content includes the following steps: Step 201: Extract structural features of digital content using multi-channel feature extraction technology.
[0048] Step 202: Use cross-modal fusion technology to align and fuse the structured features to determine fusion features after the structured features are fused.
[0049] Step 203: Map each structured feature to a unified feature space.
[0050] The system maps the structured features extracted from different modal channels into a unified feature space. In their original state, visual, audio, and text features have significant differences in dimensional distribution, numerical scale, and semantic encoding. Directly concatenating them can cause feature conflicts and noise enhancement.
[0051] To this end, the system uses a trained projection matrix or nonlinear mapping network to project the features of each modality into a low-dimensional representation space with a shared semantic distribution. This process not only achieves initial semantic alignment but also creates structural compatibility for subsequent fusion operations.
[0052] Step 204: In the feature space, normalize and compress the dimensions of each structured feature to generate a mapping feature.
[0053] After mapping is complete, the system performs normalization and dimensionality reduction on each modality within the unified feature space. This normalization process eliminates extreme deviations in the numerical distribution of each feature dimension, ensuring that the contributions of different modalities to the fusion process are comparable.
[0054] Dimensionality compression removes redundant feature components through principal component analysis, feature selection, or deep compression networks, retaining the feature dimensions that are most representative of the discrimination task, thereby reducing computational complexity and improving model generalization capabilities.
[0055] Step 205: Based on a preset fusion strategy, perform feature-level fusion on the mapped features to generate fused features.
[0056] The system performs feature-level fusion of mapped features based on a preset fusion strategy, which can be linear weighting, concatenation, attention mechanism, or multi-channel fusion network. Different fusion structures can be selected based on actual needs.
[0057] For example, a weighting mechanism can be used to highlight the importance of specific modalities, or an attention mechanism can be used to dynamically capture the complementary relationships between different modalities. After fusion is complete, the resulting fused features retain the effective information of each modality while maintaining a unified vector structure, allowing them to be directly input into decision-making models for risk assessment.
[0058] Step 206: Based on the preset decision forest model, perform a multimodal risk assessment on the fusion features to determine the modal risk score corresponding to each modality.
[0059] Step 207: Calculate the risk score of each modality based on the preset weight configuration rule to determine the comprehensive risk score corresponding to the digital content.
[0060] The above technical solution demonstrates the beneficial effects of this embodiment: after fusion, the resulting fused features retain the effective information of each modality while maintaining a unified vector structure, enabling direct input into the decision-making model for risk assessment. This fusion process significantly improves the consistency of cross-modal content expression and model trainability.
[0061] like Figure 3 FIG. 1 is another specific embodiment of a risk assessment method for digital content of the present application. This embodiment further describes the above embodiments.
[0062] In this embodiment, a risk assessment method for digital content includes the following steps: Step 301: Extract structural features of digital content using multi-channel feature extraction technology.
[0063] Step 302: Use cross-modal fusion technology to align and fuse the structured features to determine fusion features after the structured features are fused.
[0064] Step 303: Based on the preset decision forest model, a multimodal risk assessment is performed on the fusion features to determine the modal risk score corresponding to each modality.
[0065] Step 304: Statistics are collected on the classification paths of the fusion features in each decision tree to generate risk judgment results corresponding to each modality.
[0066] The decision forest model consists of multiple specialized decision trees, each of which is trained for a specific risk dimension - for example, the legal compliance assessment tree focuses on identifying normative features, the ethics assessment tree detects interactive behavior features, and the content security tree is responsible for content semantic association analysis.
[0067] When the fused features are input into the forest, each decision tree performs classification based on its specialized dimension: the legal assessment tree analyzes the correlation between logos, audio terms, and textual clauses in the image; the ethical assessment tree tracks the behavioral logic of multimodal interactions. Each tree generates a complete classification path from the root node to the leaf node, with the branch nodes in the path recording the judgment labels for the multimodal features under that dimension.
[0068] Step 305: Based on the preset measurement rules, each risk judgment result is quantified to generate a modality risk score corresponding to each modality.
[0069] The system aggregates multi-dimensional risk assessment results by counting the classification paths of all decision trees in the forest. For example, for legal risk, if more than 60% of legally relevant decision trees (including legal feature classifiers and some general classifiers) mark at least three "high-risk" nodes in their paths, and the patent description feature in the text mode is repeatedly triggered, the legal risk assessment result is generated as requiring manual review.
[0070] Subsequently, the preset measurement rules convert such qualitative judgments into quantitative scores: based on the dimension weight coefficient, path confidence and cross-tree consistency index, the score of each modality, namely the modal risk score, is calculated.
[0071] The training of the decision forest model constructs multiple training subsets based on preset labeled samples; each training subset is used to train a training tree; during the training process, the sample features in the training subset are split at nodes according to the preset feature partitioning criteria; the splitting threshold of the training tree is adjusted based on the preset loss function; when the splitting performance of the training tree meets the preset training conditions, the training tree is determined to be a decision tree; and the decision trees are combined to determine the decision forest model.
[0072] Multiple training subsets are constructed based on pre-set annotated samples. Each training subset contains representative samples from different modalities and diverse risk types within digital content. By partitioning the training dataset, the model's generalization and overfitting resistance are effectively enhanced, ensuring coverage of a wider range of feature spaces and risk scenarios during training. Each training subset is then used to independently train a training tree, resulting in a diverse set of decision trees. During node splitting, feature split points are selected based on the criterion of maximizing the Gini impurity gain.
[0073] To further optimize model performance, the training process also dynamically adjusts the splitting threshold of the training tree using a preset loss function. This balances classification accuracy with model complexity and avoids overfitting caused by excessive splitting. For example, the Gini gain of all candidate feature split points is calculated, and the feature and threshold with the largest gain are selected for splitting. To improve model adaptability, the splitting threshold is dynamically adjusted based on the cross-entropy loss function. After each round of training, the threshold is updated based on the gradient direction of misclassified samples, and a historical threshold mean constraint is introduced to prevent oscillation.
[0074] Once a trained tree achieves the predetermined splitting performance metric and meets the set training conditions, it is officially designated as a decision tree. Ultimately, all trained and validated decision trees are combined to construct a decision forest model. By aggregating the judgment results of multiple decision trees, the decision forest can integrate the strengths of each tree to achieve more robust and accurate multimodal risk assessment, improving the overall model's noise resistance and generalization performance.
[0075] Step 306: Calculate the risk score of each modality based on the preset weight configuration rule to determine the comprehensive risk score corresponding to the digital content.
[0076] Through the above technical solutions, it can be seen that the beneficial effect of this embodiment is that the system can scientifically and quantitatively characterize the risk level of each modality, laying a solid foundation for subsequent multimodal fusion and comprehensive scoring.
[0077] like Figure 4 The figure shows a specific embodiment of a risk assessment device for digital content in the present application. This embodiment is a risk assessment device for digital content, which is used to perform Figures 1-3 A physical device for a risk assessment method for digital content. The technical solution is essentially the same as that of the above embodiment, and the corresponding descriptions in the above embodiment are also applicable to this embodiment. In this embodiment, a risk assessment device for digital content includes: The structured feature determination module 401 is configured to extract the structured features of the digital content using a multi-channel feature extraction technique; The fusion feature determination module 402 is configured to align and fuse the structured features using a cross-modal fusion technology to determine a fusion feature after the structured features are fused; The modality risk score determination module 403 is configured to perform a multimodal risk assessment on the fused features based on a preset decision forest model to determine a modality risk score corresponding to each modality; The comprehensive risk score determination module 404 is configured to calculate the risk score of each modality based on a preset weight configuration rule to determine a comprehensive risk score corresponding to the digital content.
[0078] Figure 5 : This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include internal memory, such as high-speed random access memory (RAM), and may also include non-volatile memory (non-volatile memory), such as at least one disk storage. Of course, the electronic device may also include hardware required for other services.
[0079] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0080] Memory is used to store execution instructions. Specifically, execution instructions are computer programs that can be executed. Memory can include internal memory and non-volatile memory, and provides execution instructions and data to the processor.
[0081] In one possible implementation, a processor reads corresponding execution instructions from a non-volatile memory into a memory and then executes them. Alternatively, the processor may obtain corresponding execution instructions from another device to logically form a risk assessment device for digital content. The processor executes the execution instructions stored in the memory to implement a risk assessment method for digital content provided in any embodiment of the present application.
[0082] The above application Figure 4 The method implemented by a digital content risk assessment device in the illustrated embodiment can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the method can be completed by hardware integrated logic circuits within the processor or by software instructions. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor.
[0083] The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0084] The embodiment of the present application further proposes a readable medium, which stores an execution instruction. When the stored execution instruction is executed by the processor of the electronic device, the electronic device can execute a risk assessment method for digital content provided in any embodiment of the present application, and is specifically used to execute the following: Figure 1 or Figure 2 or Figure 3 The method shown.
[0085] The electronic device in each of the aforementioned embodiments may be a computer.
[0086] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods or computer program products. Therefore, the present application may adopt a completely hardware embodiment, a completely software embodiment, or a combination of software and hardware.
[0087] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiments.
[0088] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0089] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A risk assessment method for digital content, characterized in that: include: Use multi-channel feature extraction technology to extract structural features of digital content; Using cross-modal fusion technology to align and fuse the structured features to determine fusion features after the structured features are fused; Based on a preset decision forest model, a multimodal risk assessment is performed on the fusion features to determine a modality risk score corresponding to each of the modalities; Based on a preset weight configuration rule, each modal risk score is calculated to determine a comprehensive risk score corresponding to the digital content.
2. The method according to claim 1, characterized in that The digital content includes a visual channel, an audio channel, and a text channel, and the method of extracting the structured features of the digital content using the multi-channel feature extraction technology includes: Using image processing technology, extracting image motion features and color distribution features corresponding to the visual channel; Utilizing spectrum analysis technology, extracting Mel spectrum features and cepstrum outlier features corresponding to the audio channel; Using semantic analysis technology, extracting semantic graph structure features and topic distribution features corresponding to the text channel; The structured feature is determined based on the image motion feature and the color distribution feature, the Mel spectrum feature and the cepstral outlier feature, the semantic graph structure feature and the topic distribution feature.
3. The method according to claim 2, characterized in that The aligning and fusing the structured features using the cross-modal fusion technology to determine the fused features after the structured features are fused includes: Mapping each of the structured features to a unified feature space; In the feature space, normalizing and dimensionally compressing each of the structured features to generate mapping features; Based on a preset fusion strategy, the mapping features are fused at a feature level to generate the fused features.
4. The method according to claim 3, characterized in that The decision forest model includes multiple decision trees. Then, based on the preset decision forest model, the multimodal risk assessment of the fusion features is performed to determine the modal risk score corresponding to each modality, including: Performing statistics on the classification paths of the fusion features in each of the decision trees to generate risk judgment results corresponding to each of the modalities; Based on preset measurement rules, each of the risk judgment results is quantified to generate the modality risk score corresponding to each of the modalities.
5. The method according to claim 4, characterized in that Also includes: Construct multiple training subsets based on preset labeled samples; Each of the training subsets is used to train a training tree; During the training process, the sample features in the training subset are node-split according to the preset feature partitioning criteria; Adjusting the splitting threshold of the training tree based on a preset loss function; When the splitting performance of the training tree meets the preset training condition, determining the training tree to be the decision tree; The decision trees are combined to determine the decision forest model.
6. The method according to claim 5, characterized in that The step of calculating each modality risk score based on a preset weight configuration rule to determine a comprehensive risk score corresponding to the digital content includes: Determining weight parameters corresponding to each of the modalities based on historical risk score information; A weighted calculation is performed based on each of the modality risk scores and the weight parameter to determine the comprehensive risk score.
7. The method according to any one of claims 1 to 6, characterized in that: Also includes: Extracting structural information and content summary information of the digital content to generate basic report information; Determining a risk level corresponding to the digital content based on the comprehensive risk score; The risk level and the report basic information are combined to generate a risk assessment report corresponding to the digital content.
8. A risk assessment device for digital content, characterized in that: include: a structured feature determination module configured to extract structured features of digital content using a multi-channel feature extraction technique; a fusion feature determination module, configured to align and fuse the structured features using a cross-modal fusion technology to determine a fusion feature after the structured features are fused; a modality risk score determination module, configured to perform a multimodal risk assessment on the fused features based on a preset decision forest model to determine a modality risk score corresponding to each of the modalities; The comprehensive risk score determination module is configured to calculate the risk score of each modality based on a preset weight configuration rule to determine the comprehensive risk score corresponding to the digital content.
9. A computer-readable storage medium storing a computer program, characterized in that: The computer program is used to execute the risk assessment method for digital content described in any one of claims 1 to 7.
10. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the risk assessment method for digital content according to any one of claims 1 to 7.
Citation Information
Cited By
Supplier dynamic risk assessment method and device
CN121481278A
Media content risk early warning method based on multi-modal data and dynamic entropy weight model
CN121707336A