Electronic publication content intelligent auditing system based on AI
Through the AI-based electronic publication content intelligent review system, multimodal data acquisition and deep learning technology, the diversity and copyright issues of electronic publication content are solved, efficient and accurate content review is achieved, adapting to the needs of different audiences, and ensuring content compliance and copyright security.
Patent Information
- Application Number
- CN202510553407.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The diversity and copyright issues of electronic publications have led to inefficient and insufficient accuracy of traditional manual reviews, making it difficult to keep up with the pace of rapid updates of electronic publications, and copyright infringement is concealed, making it difficult to accurately judge the existing technology.
The AI-based electronic publication content intelligent review system, including a multimodal data acquisition module, a semantic feature extraction module, a violation content detection module, a copyright matching module and a dynamic audit strategy generation module, uses a multimodal fusion coding algorithm, a deep learning classification model, a knowledge graph embedding model and a reinforcement learning decision-making framework to achieve comprehensive and intelligent review of electronic publication content.
It improves the efficiency and accuracy of content review of electronic publications, can accurately identify violations and infringements in various modal content, adapt to the needs of different audiences, ensure that the content will not have adverse effects during the dissemination process, and supports the healthy and orderly development of the electronic publishing industry.
Smart Images

Figure CN120067278A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic publication review, and specifically to an intelligent content review system for electronic publications based on AI. Background Art
[0002] With the full advancement of the digital age, electronic publications have occupied an increasingly important position in people's lives and studies. From various online literary works, e-magazines, to online course materials containing rich multimedia content, the forms and quantities of electronic publications have shown explosive growth. However, behind this prosperous scene, there are many problems that need to be solved urgently, bringing unprecedented challenges to the content review work of electronic publications.
[0003] The diversity of electronic publication content is the primary problem. It covers various modalities such as text, images, audio and video, and each modality has its unique manifestation form and review key points. For example, text content may have problems such as sensitive words, grammar errors or logical confusion; images may involve risks such as copyright infringement and bad guidance; audio and video may contain inappropriate remarks and illegal information. Traditional manual review methods are overwhelmed when faced with such complex and diverse content. Manual review is not only inefficient and difficult to keep up with the rapid update rhythm of electronic publications, but also easily affected by the subjective factors of reviewers, resulting in inconsistent review standards and thus misjudgments and omissions.
[0004] Copyright issues are also a major problem in the electronic publishing industry. In the Internet environment, the information dissemination speed is extremely fast, and copyright infringement of electronic publications becomes easier and more concealed. Some creators may use fragments of others' works without authorization, or modify and piece together original content before publishing, which brings huge losses to copyright owners. Existing copyright review technologies often rely on simple text comparison or image recognition, and it is difficult to accurately judge complex infringement forms such as semantic understanding-based content plagiarism and cross-modal creative reference, making copyright protection face severe challenges.
[0005] In addition, electronic publications have a wide range of dissemination, and the audience is large and complex. Users in different regions, different age groups, and different cultural backgrounds have different acceptance levels and value judgments of electronic publication content. Some content that is considered normal in some regions or groups may cause disputes or even violate local laws, regulations and moral norms in other places. Therefore, a review mechanism that can adapt to different audience needs and is comprehensive and accurate is needed to ensure that the content of electronic publications will not have an adverse impact during the dissemination process.
[0006] With the continuous development of technology, the creation and dissemination methods of electronic publications are also constantly innovating. Emerging creation tools and platforms are emerging continuously, making the production of content more convenient, but at the same time increasing the difficulty of review. For example, for some content generated by artificial intelligence, its review criteria and methods are not yet perfect, and traditional review means are difficult to cope with. Facing these increasingly complex situations, the existing review system can no longer meet the rapid development needs of the electronic publishing industry, and there is an urgent need for an innovative and intelligent review system to ensure the healthy and orderly development of the industry. Summary of the Invention
[0007] The purpose of the present invention is to provide an intelligent review system for electronic publication content based on AI to solve the problems raised in the above background technology.
[0008] To achieve the above purpose, the present invention provides the following technical solutions: An intelligent review system for electronic publication content based on AI, the system includes: A multi-modal data acquisition module, used to obtain text, image, and audio-video content data sets in electronic publications and generate structured review data; A semantic feature extraction module, used to map the structured review data into a semantic feature space and generate a semantic feature vector cluster based on a multi-modal fusion coding algorithm, the semantic feature vector cluster includes semantic relevance, context logic, and cross-modal consistency features; A violation content detection module, used to dynamically calibrate the confidence of the semantic feature vector cluster according to preset sensitive information recognition rules, divide the content risk level through a deep learning classification model, and extract the violation feature set within each level; A copyright matching module, used to construct a copyright verification unit including a knowledge graph embedding model, and perform similarity iterative calculation on the knowledge graph embedding model using the violation feature set to generate a copyright infringement determination result; A dynamic review policy generation module, used to perform multi-dimensional rule fusion on the copyright infringement determination result according to a preset reinforcement learning decision framework and output an intelligent review instruction for electronic publication content.
[0009] Preferably, generating the semantic feature vector cluster based on the multi-modal fusion coding algorithm includes: Extract text paragraphs, image regions, and audio-video segments in the structured review data to construct a multi-modal tensor; Use a sliding context segmentation algorithm to perform local feature extraction on the multi-modal tensor to generate a set of context sub-tensors; Adopt a multi-head attention mechanism to perform cross-modal correlation analysis on the set of context sub-tensors and eliminate irrelevant noise features; The remaining sub-tensors are aggregated in a graph structure through a graph convolutional network to generate a cluster of feature vectors with semantic relevance.
[0010] Preferably, the dynamic confidence calibration includes: Calculating the mean and standard deviation of the confidence of violation features within each level according to the historical review result distribution of the content risk level; Based on the time decay factor algorithm, sliding window weighted correction is performed on the mean and standard deviation to generate a dynamic confidence benchmark; Using a probabilistic graph model to model the uncertainty of the dynamic confidence benchmark to generate an adaptive confidence interval.
[0011] Preferably, the division of the content risk level by the deep learning classification model includes: Defining the loss function as the joint minimization of the risk level misjudgment rate and the feature confusion degree; Setting the constraint conditions as the semantic coherence threshold, the copyright similarity upper limit, and the cross-modal alignment error range; Optimizing the loss function through the gradient descent algorithm combined with adversarial training, and outputting the optimal risk level division scheme.
[0012] Preferably, the similarity iterative calculation of the knowledge graph embedding model using the set of violation features includes: Inputting the set of violation features into the entity alignment layer of the knowledge graph embedding model, and using the TransR algorithm to extract the copyright entity vectors; Performing matching degree sorting on the copyright entity vectors through cosine similarity calculation to generate an infringement suspicion sequence; Adopting negative sampling technology combined with contrastive learning to optimize the entity vector space and update the decision-making parameters of the copyright verification unit.
[0013] Preferably, the multi-modal data acquisition module preprocesses the content data set according to a preset multi-modal content cleaning strategy; The multi-modal content cleaning strategy includes: Identifying duplicate text fragments and similar image regions in the content data set, and performing deduplication based on the hash fingerprint comparison rule; Detecting missing or damaged media data, and using a generative adversarial network to repair the missing data; Performing normalization processing on the repaired data to generate structured review data with unified dimensions.
[0014] Preferably, the reinforcement learning decision-making framework includes: Constructing a multi-objective decision-making model based on Q-learning to quantify the impact of review instructions on content compliance, copyright protection, and user experience; Generate a state-action sample set through an experience replay pool, and calculate the long-term cumulative reward for each action; Select actions with reward values higher than a preset threshold as the basis for generating optimized review instructions.
[0015] Preferably, the multi-dimensional rule fusion includes: Integrate the semantic rule library, copyright regulation library, and user feedback data to construct a heterogeneous rule fusion matrix; Use a matrix decomposition algorithm to perform feature decoupling on the heterogeneous rule fusion matrix, and extract the core rule feature vector; Input the core rule feature vector into the dynamic review policy generation module to generate review instructions that integrate multi-dimensional rules.
[0016] Preferably, the system further includes: Construct a warning module including a content quality evaluation function, perform correlation analysis on the review instructions and the content quality score, and output a hierarchical review suggestion; Trigger a preset content correction mechanism according to the hierarchical review suggestion to generate text revision or media replacement instructions.
[0017] Preferably, the present invention further includes an electronic device, and the device includes: At least one processor; and a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the operations of the AI-based intelligent content review system for electronic publications as described above.
[0018] Compared with the prior art, the beneficial effects of the present invention are: The AI-based intelligent content review system for electronic publications proposed by the present invention has brought significant positive impacts to the electronic publishing industry from multiple dimensions.
[0019] In the data collection and preprocessing stage, the multimodal data collection module plays a key role. It can comprehensively obtain the text, image, audio and video content data sets in electronic publications, and preprocess them through the preset multimodal content cleaning strategy. This processing method can not only remove duplicate text fragments and similar image areas to avoid the waste of audit resources, but also use generative adversarial networks to repair missing or damaged media data, greatly improving the integrity and availability of audit data. Taking an electronic textbook containing a large number of illustrations as an example, the multimodal data collection module can accurately identify and delete duplicate illustrations, and repair some damaged images due to file transfer problems, so that subsequent audit work can be based on complete and high-quality data, providing a solid foundation for accurately judging content quality and compliance.
[0020] The semantic feature extraction module generates semantic feature vector clusters through a multimodal fusion coding algorithm. This process achieves a deep understanding of the content of electronic publications. Unlike traditional single-modal analysis methods, it can capture semantic relevance, contextual logic, and cross-modal consistency features. When reviewing electronic novels with both pictures and text, this module can explore the potential connection between text descriptions and illustrations, understand the author's complete intentions expressed through different modalities, and more accurately determine whether the content is in violation of regulations or infringement. This multimodal fusion analysis method effectively avoids the understanding bias caused by modal separation, greatly improving the accuracy and comprehensiveness of the review.
[0021] The dynamic confidence calibration and deep learning classification model of the illegal content detection module are important guarantees for ensuring the accuracy of the audit. The dynamic confidence calibration is based on the distribution of historical audit results and the time decay factor algorithm, which fully considers the timeliness and uncertainty of the data, making the audit results closer to the actual situation. The deep learning classification model is optimized by defining a reasonable loss function and setting strict constraints, which can accurately divide the content risk level and extract the illegal feature set. In practical applications, for some sensitive information with fuzzy boundaries, the module can accurately judge its risk level and reduce the occurrence of misjudgment and missed judgment. For example, when reviewing electronic short articles posted on social media, even if the sensitive information is expressed in an obscure way, the illegal content detection module can accurately identify it with its powerful analysis ability and give the corresponding risk level, providing a strong basis for subsequent processing.
[0022] The copyright matching module significantly improves the accuracy and efficiency of copyright infringement determination by using a knowledge graph embedding model and similarity iterative calculation. By extracting copyright entity vectors through the TransR algorithm and optimizing the entity vector space by combining cosine similarity calculation and negative sampling technology, it can quickly and accurately identify potential copyright infringement behaviors. When processing a large number of academic literature electronic publications, this module can quickly locate whether there are infringement problems in the cited content, effectively protecting the legitimate rights and interests of copyright owners, maintaining the copyright order in the electronic publishing industry, and promoting the legal dissemination and innovation of knowledge.
[0023] The dynamic audit policy generation module generates intelligent audit instructions based on a reinforcement learning decision framework and multi-dimensional rule fusion, realizing the intelligent and dynamic adjustment of audit policies. It comprehensively considers various factors such as content compliance, copyright protection, and user experience, and can automatically generate the most appropriate audit instructions according to different content risk levels and types. For low-risk daily news electronic publications, the system can adopt a fast audit policy to improve audit efficiency; while for electronic publications involving important copyrights or sensitive content, strict and detailed audits will be carried out to ensure content security. This audit method of optimizing resource allocation not only improves audit efficiency but also ensures audit quality, making the audit work more scientific and reasonable.
[0024] The early warning module and content correction mechanism in the system further improve the audit process. The early warning module outputs graded audit suggestions based on audit instructions and content quality scores, and the content correction mechanism generates text revision or media replacement instructions according to the suggestions. This not only helps to timely discover potential problems in electronic publications but also provides specific solutions to the problems, achieving all-round quality control of the content of electronic publications. Through this mechanism, the creators and publishers of electronic publications can timely modify and improve the content, enhancing the overall quality of electronic publications, providing users with higher-quality and legal content, and promoting the healthy and sustainable development of the electronic publishing industry. Description of the Drawings
[0025] Figure 1 It is the working principle diagram of the AI-based intelligent content audit system for electronic publications described in the present invention; Figure 2 It is a schematic diagram of generating a semantic feature vector cluster by a multi-modal fusion coding algorithm; Figure 3 It is a step diagram of similarity iterative calculation of the knowledge graph embedding model; Figure 4 It is a flow chart of multi-dimensional rule fusion. Detailed Implementation Modes
[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] Please refer to Figures 1-4 , the present invention provides an intelligent content review system for e-publications based on AI. The system aims to comprehensively, efficiently, and intelligently review the content of e-publications to ensure its compliance, copyright security, and content quality. Its overall implementation plan is as follows: The multimodal data acquisition module is responsible for obtaining the text, image, and audio-video content data sets in e-publications. After obtaining the data, it will preprocess the content data set according to the preset multimodal content cleaning strategy. Specifically, it will identify duplicate text fragments and similar image regions in the content data set, and perform deduplication operations based on the hash fingerprint comparison rule; detect missing or damaged media data, and use the generative adversarial network to repair the content of the missing data; perform normalization processing on the repaired data, and finally generate structured review data to provide a standardized and unified data basis for the subsequent review process.
[0028] The semantic feature extraction module receives the structured review data generated by the multimodal data acquisition module and maps it into the semantic feature space. Through a series of operations, a semantic feature vector cluster is generated based on the multimodal fusion coding algorithm. This cluster contains semantic relevance, context logic, and cross-modal consistency features, which are crucial for accurately understanding and analyzing the content of e-publications and are the key basis for the subsequent review work.
[0029] The violation content detection module processes the semantic feature vector cluster generated by the semantic feature extraction module according to the preset sensitive information recognition rules. During the processing, dynamic confidence calibration will be performed first, and then the content risk level will be divided through a deep learning classification model, and the violation feature set within each level will be extracted to clarify the content part with risks in the e-publication.
[0030] The copyright matching module constructs a copyright verification unit that includes a knowledge graph embedding model. Using the violation feature set extracted by the violation content detection module, similarity iterative calculation is performed on the knowledge graph embedding model, and then a copyright infringement determination result is generated to determine whether there is a copyright infringement problem in the e-publication.
[0031] The dynamic audit policy generation module performs multi-dimensional rule fusion on the copyright infringement determination results generated by the copyright matching module according to the preset reinforcement learning decision framework. It integrates rules and data from multiple aspects and finally outputs an intelligent audit instruction for the content of e-publications, which is used to guide the audit operation of e-publication content and achieve intelligent and accurate auditing.
[0032] The following further describes the implementation of the present invention in combination with Embodiments 1 to 6. Embodiment 1
[0033] After the multi-modal data acquisition module obtains the content data set of the e-publication, it starts to execute the multi-modal content cleaning strategy. When identifying duplicate text fragments, it will perform chunking on the text content and generate a unique hash fingerprint for each text block. A hash fingerprint is a fixed-length code generated based on the characteristics of the text content, and the hash fingerprints generated by different text contents are likely to be different. By comparing the hash fingerprints, it is possible to quickly and accurately determine which text fragments are duplicates and then remove them. For the identification of similar image regions, it is also based on the hash fingerprint comparison rule. First, the features of the image are extracted and converted into data in a specific format, and then the hash fingerprint is generated for comparison to achieve deduplication of similar image regions.
[0034] When detecting missing or damaged media data, a generative adversarial network (GAN) is used for content repair. The generative adversarial network consists of a generator and a discriminator. The generator is responsible for generating the missing data according to the existing data features, and the discriminator judges the authenticity of the generated data. In the continuous adversarial training process, the data generated by the generator gets closer and closer to the real data, thus completing the repair of the missing data. After the repair is completed, in order to make the data have a unified format and scale for subsequent processing, it is necessary to normalize the data. For example, for image data, its pixel values are mapped to the interval [0, 1] or [-1, 1]; for text data, corresponding normalization operations are performed according to different feature representation methods. Finally, structured audit data with unified dimensions is generated.
[0035] When the semantic feature extraction module generates a semantic feature vector cluster, it first extracts text paragraphs, image regions, and audio-video segments from the structured audit data to construct a multi-modal tensor. A multi-modal tensor is a mathematical structure that can represent multiple modalities of data at the same time. It integrates data of different modalities for convenient subsequent unified processing. Then, the sliding context segmentation algorithm is used to extract local features from the multi-modal tensor. The sliding context segmentation algorithm slides on the multi-modal tensor according to a certain window size and step length, and extracts a local region each time. These local regions contain rich context information. In this way, a set of context sub-tensors is generated.
[0036] After that, a multi-head attention mechanism is adopted to perform cross-modal correlation analysis on the set of context sub-tensors. The multi-head attention mechanism can focus on and analyze data from different perspectives. By performing parallel calculations with multiple attention heads, it can capture more comprehensive information. During this process, irrelevant noise features will be identified and removed, making the data cleaner and more conducive to subsequent analysis. Finally, the remaining sub-tensors are aggregated through a graph convolutional network. The graph convolutional network can perform convolutional operations on graph-structured data, mine the internal relationships between data, and aggregate the sub-tensors into a cluster of feature vectors with semantic relevance. This cluster contains the semantic relevance, context logic, and cross-modal consistency features of various modal data such as text, images, and audio-visual, providing an important data basis for subsequent review work. Embodiment 2
[0037] This embodiment focuses on the dynamic confidence calibration process in the violation content detection module. In the violation content detection module, dynamic confidence calibration is a key link. First, the confidence mean and standard deviation of violation features within each level are calculated based on the historical review result distribution of content risk levels. Assume that the content risk levels are divided into three levels: low, medium, and high, represented by , , . For each level, a large amount of historical review data is collected, and relevant data such as the frequency of occurrence of violation features is statistically analyzed. Let the number of historical review samples of a certain violation feature in level ( ) be , and the confidence of this violation feature in each sample be ( ). Then the confidence mean of the violation feature within this level is calculated by the formula: . The standard deviation is calculated by the formula: .
[0038] Based on the time decay factor algorithm, the mean and standard deviation are corrected with weighted sliding windows to generate a dynamic confidence benchmark. The time decay factor algorithm takes into account the timeliness of review data. The closer the data is to the current time, the higher its reference value for the current review. Let the time decay factor be ( ), and the size of the sliding window be . For the current moment , when calculating the dynamic confidence benchmark, different weights are assigned to the historical data within the window. The weight calculation formula is ( ), where represents the position of the data within the sliding window. The corrected mean and standard deviation respectively as follows: , .
[0039] The probability graph model is used to model the uncertainty of the dynamic confidence benchmark and generate an adaptive confidence interval. The probability graph model can intuitively represent the probability relationship between variables. By constructing a suitable probability graph model, combining the corrected mean and standard deviation, considering the uncertainty factors of the data, an adaptive confidence interval is calculated. For example, using a Gaussian probability graph model, according to the characteristics of the Gaussian distribution, with the corrected mean as the center, the upper and lower limits of the confidence interval are determined in combination with the standard deviation. Set the confidence level to , then the confidence interval is , where is the quantile of the standard normal distribution, corresponding to the confidence level . In this way, the confidence interval can be dynamically adjusted according to historical data and current situations, improving the accuracy and reliability of detecting illegal content. Embodiment 3
[0040] This embodiment details the specific process of dividing the content risk level through a deep learning classification model.
[0041] In the illegal content detection module, the deep learning classification model optimizes the model by defining a suitable loss function and setting constraint conditions, and using the gradient descent algorithm combined with adversarial training, so as to achieve an accurate division of the content risk level.
[0042] Define the loss function as the joint minimization of the risk level misjudgment rate and the feature confusion degree. Let the risk level misjudgment rate be , and the feature confusion degree be , the expression of the loss function is: , where and are weight coefficients, used to balance the importance of the risk level misjudgment rate and the feature confusion degree in the loss function, and . The risk level misjudgment rate represents the proportion of samples in which the risk level predicted by the classification model is inconsistent with the actual risk level. The feature confusion degree is used to measure the confusion degree between features of different risk levels. The higher the feature confusion degree, the worse the ability of the model to distinguish different risk levels.
[0043] Set the constraint conditions as the semantic coherence threshold, the upper limit of copyright similarity, and the cross-modal alignment error range. The semantic coherence threshold is used to ensure that when the model divides the risk level, it takes into account the semantic coherence of the text content. If the semantic coherence of a piece of text is lower than the set threshold, it indicates that there may be problems with this piece of text and its risk level needs to be increased. Let the semantic coherence threshold be , and by calculating the semantic coherence index of the text , when , adjust the result of the risk level division. The upper limit of copyright similarity is used to control the similarity between the content of e-publications and the existing copyrighted content to prevent copyright infringement. Let the upper limit of copyright similarity be , when the calculated copyright similarity is greater than , increase the risk level. The cross-modal alignment error range is used to ensure the consistency of multi-modal data in the feature extraction and analysis process. Let the cross-modal alignment error range be , when the cross-modal alignment error is not within this range, adjust the model.
[0044] Optimize the loss function through the gradient descent algorithm combined with adversarial training, and output the optimal risk level division scheme. The gradient descent algorithm is a commonly used optimization algorithm. By continuously calculating the gradient of the loss function with respect to the model parameters and updating the parameters along the opposite direction of the gradient, the loss function is gradually reduced. During the adversarial training process, adversarial samples are introduced, which are generated by making small perturbations to the original data, aiming to enhance the robustness of the model. During the training process, continuously adjust the model parameters until the loss function converges to a smaller value. At this time, the obtained risk level division scheme is the optimal scheme. In this way, the accuracy and reliability of the deep learning classification model for dividing the content risk level can be improved, providing more powerful support for the review of e-publication content. Example 4
[0045] This example details the process of using the knowledge graph embedding model for similarity iterative calculation in the copyright matching module.
[0046] The copyright verification unit constructed by the copyright matching module contains a knowledge graph embedding model, and uses the set of violation features to perform similarity iterative calculation on this model, thereby generating a copyright infringement determination result.
[0047] Input the set of illegal features into the entity alignment layer of the knowledge graph embedding model, and use the TransR algorithm to extract the copyright entity vectors. A knowledge graph is a semantic network used to represent the relationships between entities. In the knowledge graph embedding model, entities and relationships need to be mapped into a low-dimensional vector space for calculation and analysis. The TransR algorithm is a commonly used knowledge graph embedding algorithm that performs projection operations in the entity space and the relationship space respectively, and can better capture the semantic information of entities and relationships. Let the set of illegal features be and the set of entities in the knowledge graph be , and the set of relationships be. In the entity alignment layer, each feature in the set of illegal features is matched and aligned with the entities in the knowledge graph. For each entity , calculate its corresponding copyright entity vector through the TransR algorithm. The specific calculation process is as follows: First, define the projection matrix, where represents the relationship. For a triple ( represents the head entity, represents the tail entity), in the projection space of the relationship , the projection vectors of the head entity and the tail entity are respectively , . In this way, project the features in the set of illegal features into the copyright entity vector space to obtain the copyright entity vectors.
[0048] Sort the matching degrees of the copyright entity vectors through cosine similarity calculation to generate an infringement suspicion sequence. Cosine similarity is a commonly used method to measure the similarity between two vectors, and its calculation formula is: , where and are two copyright entity vectors, represents the dot product of the vectors, and respectively represent the norms of the vectors. Calculate the cosine similarity between the copyright entity vectors corresponding to the set of illegal features and the existing copyright entity vectors in the knowledge graph, sort the similarities from high to low to generate an infringement suspicion sequence. The higher the similarity, the more similar the content corresponding to the illegal feature is to the existing copyright content, and the greater the infringement suspicion.
[0049] The negative sampling technique is combined with contrastive learning to optimize the entity vector space and update the decision parameters of the copyright verification unit. The negative sampling technique randomly selects a part of samples from a large number of negative samples for training to improve the training efficiency. In this embodiment, some non-copyright-related entities are randomly selected from the knowledge graph as negative samples. Contrastive learning differentiates between positive samples (i.e., copyright entities related to violation features) and negative samples, enabling the model to better distinguish different entities. During the contrastive learning process, a contrastive loss function is defined, such as the InfoNCE loss function: , where is the positive sample vector, is the query vector, is the negative sample vector, is the number of negative samples, is the temperature parameter, which is used to control the difficulty of contrastive learning. By continuously performing negative sampling and contrastive learning, the entity vector space is optimized, enabling the copyright entity vector to more accurately represent copyright information. At the same time, the decision parameters of the copyright verification unit are updated to improve the accuracy of copyright infringement determination. Embodiment 5
[0050] This embodiment elaborates in detail the specific implementation process of the reinforcement learning decision framework and multi-dimensional rule fusion.
[0051] In the dynamic audit policy generation module, multi-dimensional rule fusion is performed on the copyright infringement determination results based on a preset reinforcement learning decision framework to output intelligent audit instructions for e-publication content.
[0052] A multi-objective decision-making model based on Q-learning is constructed to quantify the impact of audit instructions on content compliance, copyright protection, and user experience. Q-learning is a model-free reinforcement learning algorithm that selects the optimal action by learning the state-action value function (Q function). In this embodiment, the state represents the current audit state of the e-publication, including information such as copyright infringement determination results and content risk levels; the action represents various possible audit instructions, such as passing the audit, passing the audit after partial modification, rejecting the audit, etc. Let the state space be , the action space be , and the Q function be , where , . Define the reward function , which is used to quantify the impact of audit instructions on content compliance, copyright protection, and user experience. For example, if the audit instruction can ensure content compliance, protect copyright, and not affect the user experience, a higher reward is given; otherwise, a lower reward is given. When calculating the reward, multiple factors need to be comprehensively considered. Let the weight of content compliance be , the weight of copyright protection be The weight of the user experience is and . Then the reward function can be expressed as: , where , , respectively represent the reward values of the review instruction in terms of content compliance, copyright protection, and user experience.
[0053] Generate a state-action sample set through the experience replay pool, and calculate the long-term cumulative reward of each action. The experience replay pool is used to store historical state-action pairs and their corresponding rewards and the next state. During training, a batch of samples is randomly drawn from the experience replay pool for learning. Let the experience replay pool be , and each time samples are drawn from samples ( ), where is the current state, is the action taken, is the reward, is the next state. According to the update formula of the Q-learning algorithm: , where is the learning rate, is the discount factor, which is used to measure the importance of future rewards. By continuously updating the Q function, calculate the long-term cumulative reward of each action.
[0054] Select the actions with reward values higher than the preset threshold as the basis for generating the optimized review instruction. The preset threshold is set according to actual needs and experience. After training is completed, for the current review state , calculate the Q values of all actions , and select actions as the basis for generating the optimized review instruction. If there are multiple actions that meet the conditions, the final review instruction can be determined according to a certain strategy (such as random selection or selecting the action with the largest Q value).
[0055] In terms of multi-dimensional rule fusion, integrate the semantic rule library, copyright regulation library, and user feedback data to construct a heterogeneous rule fusion matrix. The semantic rule library contains rules regarding the legality and normativity of text semantics, the copyright regulation library contains laws and regulations related to copyright, and the user feedback data reflects users' views and opinions on the content of e-publications. Integrate these different types of data and present them in the form of a matrix. Each row in the matrix represents a rule, and each column represents a feature of the rule.
[0056] The singular value decomposition (SVD) algorithm is used to decouple the features of the heterogeneous rule fusion matrix and extract the core rule feature vectors. The matrix decomposition algorithm can decompose a large matrix into the product of multiple small matrices, thereby reducing the data dimension while retaining the key information. By analyzing the magnitudes of the singular values, the main features to be retained can be determined. The column vectors corresponding to the larger singular values are selected to form the core rule feature vectors. These core rule feature vectors contain the key information in semantic rules, copyright regulations, and user feedback data.
[0057] The core rule feature vectors are input into the dynamic audit policy generation module to generate audit instructions that integrate multi-dimensional rules. After receiving the core rule feature vectors, the dynamic audit policy generation module combines the results obtained from the reinforcement learning decision framework and comprehensively generates audit instructions. For example, if the core rule feature vectors indicate semantic non-compliance and possible copyright issues in a certain e-publication, and at the same time the action rewards in the reinforcement learning decision framework suggest strict auditing, then the generated audit instructions may require a detailed review of the relevant content of the e-publication and provide modification suggestions. The audit instructions generated in this way integrate multi-dimensional rules and can more comprehensively and accurately guide the audit work of e-publication content, improving the quality and efficiency of auditing. Example 6
[0058] This example mainly describes the working process of the warning module in the system and the triggering and execution process of the content correction mechanism.
[0059] The system also includes a warning module that constructs a content quality evaluation function. This module performs correlation analysis on the audit instructions and the content quality scores and outputs graded audit suggestions. The content quality evaluation function is constructed based on multiple factors, including but not limited to the grammatical correctness of the text, semantic integrity, image clarity, audio-visual fluency, etc. For example, for text content, the text quality score can be calculated by checking indicators such as the number of grammar errors and the richness of vocabulary; for image content, it can be evaluated according to indicators such as image resolution and color restoration.
[0060] After receiving the audit instructions output by the dynamic audit policy generation module, the warning module performs correlation analysis on them and the content quality scores. If the audit instructions indicate a certain risk and the content quality score is low, then the warning module will output more stringent graded audit suggestions. For example, when the audit instructions point out minor violation features and the content quality score is at a medium to low level, the warning module may recommend a secondary review of the relevant content, focusing on the parts that may have problems. If the audit instructions show serious violation problems, even if the content quality score is high, the warning module will issue a high-risk warning, suggesting immediately stopping the release and conducting a comprehensive rectification.
[0061] Trigger the preset content correction mechanism according to the hierarchical review suggestions to generate text revision or media replacement instructions. When the hierarchical review suggestions output by the early warning module require content correction, the content correction mechanism is triggered. If the problem lies in the text part, the system will generate text revision instructions based on the violation characteristics and review suggestions. For example, if sensitive words are detected in the text, the system will give suggestions on replacing words or modifying the sentence structure according to the specific situation of the words and the context. For text with grammar errors, the system will use natural language processing technology to automatically identify the error type and generate corresponding correction instructions.
[0062] When the problem involves media content such as images, audio, and video, the system will generate media replacement instructions. For example, if there are copyright issues or the clarity of an image does not meet the requirements, the system will search for a qualified alternative image in the media resource library and generate a replacement instruction. For audio and video content, if there are issues such as infringement or freezing, the system will find legal and higher-quality audio and video resources and guide the operator to make replacements. Through such an early warning module and content correction mechanism, the system can timely detect problems in e-publications and take effective measures to handle them, ensuring the quality and compliance of e-publications.
[0063] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0064] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An AI-based intelligent review system for electronic publication content, characterized in that: The system comprises: Multimodal data collection module, used to obtain text, image, audio and video content data sets in electronic publications and generate structured review data; A semantic feature extraction module, used for mapping the structured audit data into a semantic feature space, generating a semantic feature vector cluster based on a multimodal fusion coding algorithm, wherein the semantic feature vector cluster includes semantic relevance, contextual logic and cross-modal consistency features; An illegal content detection module is used to dynamically calibrate the confidence of the semantic feature vector cluster according to preset sensitive information identification rules, divide the content risk level through a deep learning classification model, and extract illegal feature sets within each level; A copyright matching module, which is used to construct a copyright verification unit including a knowledge graph embedding model, and use the violation feature set to perform iterative similarity calculation on the knowledge graph embedding model to generate a copyright infringement judgment result; The dynamic audit strategy generation module is used to perform multi-dimensional rule fusion on the copyright infringement judgment results according to a preset reinforcement learning decision framework, and output intelligent audit instructions for electronic publication content.
2. The AI-based electronic publication content intelligent review system according to claim 1, characterized in that: The generating of semantic feature vector clusters based on the multimodal fusion coding algorithm includes: Extracting text paragraphs, image regions, and audio and video clips from the structured review data to construct a multimodal tensor; Using a sliding context segmentation algorithm to extract local features from the multimodal tensor to generate a set of context sub-tensors; A multi-head attention mechanism is used to perform cross-modal correlation analysis on the context sub-tensor set to remove irrelevant noise features; The remaining sub-tensors are aggregated through graph convolutional networks to generate feature vector clusters with semantic associations.
3. The AI-based electronic publication content intelligent review system according to claim 2, characterized in that: The dynamic confidence calibration comprises: Based on the distribution of historical audit results of content risk levels, calculate the confidence mean and standard deviation of violation features within each level; Based on the time decay factor algorithm, the mean and standard deviation are subjected to sliding window weighted correction to generate a dynamic confidence benchmark; The dynamic confidence benchmark is modeled with a probabilistic graphical model to generate an adaptive confidence interval.
4. The AI-based electronic publication content intelligent review system according to claim 1, characterized in that: The method of dividing the content risk level by the deep learning classification model includes: The loss function is defined as the joint minimization of the risk level misjudgment rate and the feature confusion; The constraints are set as the semantic coherence threshold, the copyright similarity upper limit, and the cross-modal alignment error range; The loss function is optimized by combining the gradient descent algorithm with adversarial training to output the optimal risk level classification scheme.
5. The AI-based electronic publication content intelligent review system according to claim 1, characterized in that: The using the violation feature set to iteratively calculate the similarity of the knowledge graph embedding model includes: Inputting the illegal feature set into the entity alignment layer of the knowledge graph embedding model, and extracting the copyright entity vector using the TransR algorithm; The copyright entity vectors are sorted by matching degree through cosine similarity calculation to generate a suspected infringement sequence; Negative sampling technology is combined with contrastive learning to optimize the entity vector space and update the judgment parameters of the copyright verification unit.
6. The AI-based electronic publication content intelligent review system according to claim 1, characterized in that: The multimodal data collection module pre-processes the content data set according to a preset multimodal content cleaning strategy; The multimodal content cleaning strategy includes: Identify repeated text segments and similar image regions in the content data set, and remove duplicates based on hash fingerprint comparison rules; Detect missing or damaged media data and use generative adversarial networks to repair the missing data; The repaired data is normalized to generate structured audit data with unified dimensions.
7. The AI-based electronic publication content intelligent review system according to claim 1, characterized in that: The reinforcement learning decision-making framework includes: Construct a multi-objective decision-making model based on Q-learning to quantify the impact of review instructions on content compliance, copyright protection, and user experience; Generate a state-action sample set through the experience replay pool and calculate the long-term cumulative reward of each action; Actions with reward values higher than the preset threshold are selected as the basis for generating optimized review instructions.
8. The AI-based electronic publication content intelligent review system according to claim 1, characterized in that: The multi-dimensional rule fusion includes: Integrate semantic rule base, copyright law base and user feedback data to build a heterogeneous rule fusion matrix; Using a matrix decomposition algorithm to perform feature decoupling on the heterogeneous rule fusion matrix and extracting a core rule feature vector; The core rule feature vector is input into the dynamic audit strategy generation module to generate an audit instruction integrating multi-dimensional rules.
9. The AI-based electronic publication content intelligent review system according to any one of claims 1 to 8, characterized in that: The system further comprises: Construct an early warning module including a content quality assessment function, perform correlation analysis between the review instructions and the content quality score, and output graded review suggestions; The preset content correction mechanism is triggered according to the graded review suggestions to generate text revision or media replacement instructions.
10. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; In which, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the operations of the AI-based electronic publication content intelligent review system as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Multi-mode network content security intelligent auditing system and method thereof
CN118312922A
Method and device for auditing content created by network platform user and medium
CN118350376A
Multi-agent deep reinforcement learning path planning method based on improved A*heuristic
CN118759846A
Intelligent access control management method and system based on multi-mode identification and Internet of Things technology
CN118968665A
Bidding and tendering data intelligent analysis method and system based on AI technology and storage medium
CN119850315A
Cited By
Multi-modal content compliance auditing method and system
CN120611053A
Safety improvement system and method based on content risk control engine
CN120892649A
Artificial intelligence information auditing system and method based on distributed micro-service architecture
CN120929208A
Artificial intelligence information auditing system and method based on distributed micro-service architecture
CN120929208B
Multi-modal digital textbook content examination method and device
CN121144951A