AIGC text credibility dynamic evaluation system and method based on reinforcement learning

By using a reinforcement learning-based dynamic evaluation system that combines multi-dimensional features and user feedback, the system addresses the issues of rapid iteration and cross-domain adaptability in AIGC text evaluation. It achieves efficient identification and accurate cross-domain evaluation of logically contradictory nested sentences, meeting the real-time needs of scenarios such as financial risk control.

CN120849585APending Publication Date: 2025-10-28CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510916121.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing AIGC text credibility assessment methods cannot adapt to rapidly iterating generative models, have difficulty identifying novel semantic-level deceptive content, and have limited cross-domain applications, failing to meet real-time assessment requirements.

Method used

A dynamic evaluation system based on reinforcement learning is adopted. Through feature extraction, reinforcement learning, context awareness and user feedback modules, a multi-dimensional feature fusion architecture is constructed to adjust the credibility scoring strategy in real time. Combined with deep Q network (DQN) and Markov decision process (MDP), it can achieve rapid response and accurate evaluation.

Benefits of technology

It significantly improves the ability to identify complex deception tactics such as logically contradictory nested sentences, shortens the assessment response time, and achieves a cross-domain assessment accuracy rate of over 92%, meeting the needs of real-time scenarios such as financial risk control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849585A_ABST
    Figure CN120849585A_ABST
Patent Text Reader

Abstract

The invention relates to an AIGC text credibility dynamic evaluation system and method based on reinforcement learning, and belongs to the technical field of artificial intelligence text evaluation. In order to solve the problems that a traditional method is poor in adaptability and generalization ability and cannot dynamically evaluate AIGC text credibility, a solution driven by reinforcement learning is provided, a Markov decision process is constructed through a deep Q network, and a credibility scoring strategy is dynamically adjusted; fusing the multi-dimensional features of the text, the context scene information and the real-time feedback of the user; a balance mechanism is designed, explored and utilized to accelerate algorithm convergence. The method has the technical effects that the dynamic adaptive capacity to novel false contents is remarkably improved; the objectivity and comprehensiveness of an evaluation result are enhanced; the large-scale text processing efficiency is optimized; and the generalization ability of a cross-domain application scene is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence text evaluation technology, and relates to an AIGC text credibility dynamic evaluation system and method based on reinforcement learning. Background Technology

[0002] AI-generated content (AIGC) text generation technology has been widely applied in core areas such as news gathering and editing, intelligent customer service, and advertising copywriting. According to authoritative industry reports (Gartner 2023), the global annual growth rate of AIGC text generation reached 152%, with 35.7% of the content related to business decision support scenarios. However, the rapid iteration of generation models such as GPT-4 and Claude has led to a significant increase in the semantic complexity of the generated text, posing a risk of systemic failure to traditional credibility assessment methods.

[0003] The current mainstream evaluation system suffers from fundamental bottlenecks, specifically manifested in three technological gaps:

[0004] (1) Relying on a predefined keyword library (such as a list of features for false information) and a grammatical rule engine, its static rule library requires manual updates of 28% of its entries every quarter, making it unable to identify new semantic-level deceptive content. For example, when faced with logically contradictory nested sentences, the misjudgment rate is as high as 61.3%.

[0005] (2) Algorithms such as Support Vector Machine (SVM) and Random Forest require a retraining cycle of more than 72 hours to address the data distribution drift problem. Industry test data shows that the evaluation accuracy of such models decreases by 27.6 percentage points in new scenarios.

[0006] (3) The existing scheme ignores the differences in the authority of the publishing platform (such as the credibility difference between academic journals and social media by 42%) and real-time user feedback signals, and the consistency coefficient between the evaluation results and manual review is only 0.48.

[0007] Industry-based empirical research (ICML 2023 benchmark testing) has revealed structural flaws in the existing technology:

[0008] Traditional methods cannot keep up with the weekly iteration pace of AIGC technology, and the failure rate for recognizing ironic metaphors generated by GPT-4 is as high as 68.4%.

[0009] Machine learning model updates require relabeling tens of thousands of samples, and the processing of millions of texts takes more than 18 minutes, which cannot meet the needs of real-time scenarios such as financial risk control.

[0010] The lack of inclusion of contextual features (such as the requirement for FDA certification for medical texts) and user feedback in the evaluation system limits cross-domain applications.

[0011] The core technical issues that urgently need to be addressed in this field include:

[0012] Establish a dynamic scoring strategy optimization mechanism, requiring a response latency of less than 5 seconds to adapt to the rapid iteration of AIGC technology;

[0013] Construct a multi-dimensional feature fusion architecture that integrates deep semantic features of text (lexical richness, information density), contextual scene weights, and user feedback signals.

[0014] Design a reinforcement learning-driven convergence acceleration engine to improve the efficiency of processing millions of texts to within 2 minutes;

[0015] Breaking through the bottleneck of scenario generalization, the accuracy rate of assessment in vertical fields such as news, healthcare, and finance has reached over 92%. Summary of the Invention

[0016] In view of this, the purpose of this invention is to provide a dynamic evaluation system and method for AIGC text credibility based on reinforcement learning. This invention aims to overcome the shortcomings of existing AIGC text credibility evaluation methods, such as poor adaptability and weak generalization ability, and proposes a dynamic evaluation system and method based on reinforcement learning. In existing technologies, traditional manual rules and simple statistical feature methods are difficult to adapt to the rapid evolution of AIGC text generation technology, while traditional machine learning methods are slow to update when faced with large-scale dynamically changing text data, failing to effectively identify new types of false or misleading content. This invention uses reinforcement learning modeling, combined with multi-dimensional text features, contextual information, and user feedback, to dynamically adjust the credibility scoring strategy in real time, thereby improving the accuracy and reliability of the evaluation results. Simultaneously, an exploration-utilization balancing mechanism is introduced to accelerate algorithm convergence, improve learning efficiency, and ensure that the system can quickly respond to evaluation needs. Furthermore, training on a large-scale labeled dataset enhances the model's generalization ability and adaptability, achieving broad and accurate evaluation of various types of AIGC text.

[0017] To achieve the above objectives, the present invention provides the following technical solution:

[0018] A dynamic evaluation system for AIGC text credibility based on reinforcement learning, comprising:

[0019] The feature extraction module is used to extract multi-dimensional features from AIGC text (AI-generated content) and generate feature vectors.

[0020] The reinforcement learning module uses a deep Q-network (DQN) to construct a Markov decision process (MDP) and dynamically adjusts the credibility scoring strategy.

[0021] The context-aware module is used to collect information about the text source, domain background, and publishing scenario;

[0022] The user feedback module receives user feedback and encodes it into a feedback vector.

[0023] The evaluation results output module generates a credibility evaluation report.

[0024] The output of the feature extraction module is connected to the input of the reinforcement learning module. The reinforcement learning module adjusts the scoring strategy based on the information from the context-aware module and the user feedback module. The evaluation result output module receives the output of the reinforcement learning module.

[0025] Furthermore, the features extracted by the feature extraction module include lexical richness, semantic consistency, and information density, and irrelevant symbols and stop words are removed and segmented through preprocessing.

[0026] Furthermore, the action space of the reinforcement learning module includes increasing the credibility score, decreasing the credibility score, or keeping the credibility score unchanged;

[0027] The step size for increasing or decreasing the score is a fixed percentage Δ.

[0028] Furthermore, the context-aware module adjusts the feature vector weights using weight coefficients:

[0029]

[0030] Where λ is the adjustment parameter, relevance i Let represent the correlation between the i-th type of contextual information and credibility, and n represent the number of types of contextual information.

[0031] A dynamic evaluation method for AIGC text credibility based on reinforcement learning includes the following steps:

[0032] Step 1: Preprocess the text and calculate lexical richness, semantic consistency, and information density to generate feature vectors;

[0033] Step 2: Use the feature vector as the state vector;

[0034] Step 3: Define the action space for increasing the credibility score, decreasing the credibility score, or keeping the credibility score unchanged;

[0035] Step 4: Design a reward function based on the evaluation accuracy, stability, and task importance;

[0036] Step 5: Train a deep Q-network (DQN) to dynamically adjust the scoring strategy;

[0037] Step 6: Integrate contextual information to update the state vector;

[0038] Step 7: Process user feedback and reassess credibility.

[0039] Furthermore, step 1 includes:

[0040] Vocabulary richness calculation:

[0041]

[0042] The number of distinct words refers to the number of unique words in the text, and the total number of words refers to the total number of words after preprocessing.

[0043] Furthermore, in step 4, the reward function is:

[0044] r=(accuray+stability)×task_importance

[0045] Here, accuracy indicates the consistency between the evaluation result and the true label; stability indicates that it is 1 when the fluctuation range of N consecutive evaluations is ≤ the threshold θ, otherwise it is 0; task_importance indicates the task scenario weight.

[0046] Furthermore, step 5 includes:

[0047] Deep Q-Network (DQN) updates parameters according to the Q-learning formula:

[0048]

[0049] Where α is the learning rate and γ is the discount factor.

[0050] Furthermore, step 6 includes:

[0051] The state vector update method is as follows:

[0052] s new =[V×w context ,S×w context ,I×w context ]

[0053] Where V, S, and I are the original values ​​of lexical richness, semantic consistency, and information density, respectively, and w context These are the weighting coefficients.

[0054] Furthermore, step 7 includes:

[0055] User feedback is encoded as a feedback vector f, and the state vector update formula is:

[0056] s new =s×(1-β)+f×β

[0057] Where β is the feedback fusion coefficient, and its value ranges from 0 to 1.

[0058] The beneficial effects of this invention are as follows:

[0059] (1) By using reinforcement learning modeling to achieve real-time dynamic adjustment of credibility scoring strategies, the evaluation system can adapt to the rapid iterative evolution of AIGC text generation technology, effectively overcoming the technical rigidity drawbacks of traditional rule systems and static machine learning models. The Markov Decision Process (MDP) driven by Deep Q-Network (DQN) can autonomously identify novel semantic-level false content, significantly improving the ability to identify complex deception methods such as logically contradictory nested sentences.

[0060] (2) Integrate multi-dimensional deep features of text, contextual information, and real-time user feedback to construct a three-dimensional integrated evaluation system:

[0061] The feature extraction module integrates lexical richness, semantic consistency, and information density to form a basic feature vector;

[0062] The context-aware module dynamically adjusts feature importance based on domain weight coefficients;

[0063] The user feedback module corrects evaluation biases in real time using feedback vectors.

[0064] This system breaks through the limitations of traditional methods that rely on a single dimension, enabling the evaluation results to more objectively reflect the level of authenticity and credibility of the text.

[0065] (3) Introducing an exploration-based optimization of the training process of Deep Q-Network (DQN) using a balancing mechanism:

[0066] Experience replay buffer enables efficient sample utilization;

[0067] Q-learning and updating formulas drive strategies for rapid convergence;

[0068] The reward function design incorporates stability constraints and task importance weights.

[0069] This mechanism significantly shortens model response time, ensuring the system's real-time evaluation capability in high-concurrency, large-scale text scenarios.

[0070] (4) Train a deep Q-network (DQN) model based on a large-scale labeled dataset to enable cross-domain transfer capabilities:

[0071] In the news field, the emphasis should be placed on strengthening the weight of the authority of the source;

[0072] In the medical field, the focus is on testing the consistency of professional terminology;

[0073] The financial sector is focusing on verifying the authenticity of data.

[0074] This adaptive mechanism significantly expands the applicability of the evaluation system in diverse scenarios.

[0075] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0076] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0077] Figure 1 This is a system composition diagram of the present invention;

[0078] Figure 2 This is a flowchart of the text feature extraction process in this invention;

[0079] Figure 3 This is a flowchart illustrating the action space definition in this invention.

[0080] Figure 4 This is a flowchart illustrating the design of the reward function in this invention;

[0081] Figure 5 This is a flowchart of the training process of the Deep Q-Network (DQN) in this invention;

[0082] Figure 6 This is a flowchart of the context information fusion process in this invention;

[0083] Figure 7 This is a flowchart of the user feedback processing in this invention;

[0084] Figure 8 This is a flowchart illustrating the overall workflow of the present invention. Detailed Implementation

[0085] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0086] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0087] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0088] I. System Modules of the Invention

[0089] like Figure 1 As shown, the system involved in this invention mainly includes the following key modules:

[0090] 1. Feature Extraction Module: Responsible for extracting multi-dimensional features from AIGC text, such as lexical richness, semantic consistency, and information density, and converting the text into feature vector representations to provide basic data for subsequent evaluation.

[0091] 2. Reinforcement Learning Module: The module uses a deep Q-network (DQN) as its core architecture to construct a Markov decision process (MDP), which includes elements such as state space, action space, and reward function, to achieve dynamic adjustment of the credibility scoring strategy.

[0092] 3. Context-aware module: Collects contextual information about the text, such as the source of the text, the domain background, and the publication scenario, analyzes the impact of this information on the credibility of the text, and incorporates it into the evaluation process.

[0093] 4. User Feedback Module: Receives user feedback, such as evaluations of text credibility and satisfaction with the assessment results, and adjusts the credibility score in real time to make the assessment results better meet user needs.

[0094] 5. Evaluation Result Output Module: Based on the decision results of the reinforcement learning module, the final text credibility evaluation report is generated, including credibility score, evaluation basis, etc., and presented to the user in an intuitive way.

[0095] The feature extraction module provides input data for the reinforcement learning module. The reinforcement learning module dynamically adjusts the evaluation strategy based on the information from the context awareness module and the user feedback module. The evaluation result output module presents the evaluation results to the user.

[0096] II. Specific implementation steps of the present invention

[0097] Step 1: Text feature extraction

[0098] 1.1. Text preprocessing: When receiving the AIGC text to be evaluated, first preprocess the text. The specific operations include removing irrelevant symbols (such as punctuation marks, special characters, etc.) and stop words (such as common auxiliary words like "de", "le", "zai", etc.) in the text, and performing word segmentation to cut the text into individual lexical units. For example, for the text "Today's weather is very good and suitable for going out to play.", the lexical sequence obtained after preprocessing may be "Today", "weather", "very good", "suitable", "go out", "play". As Figure 2 shown.

[0099] 1.2. Calculation of lexical richness: Count the number of different words and the total number of words in the text, and calculate the lexical richness according to the formula.

[0100] The specific formula is:

[0101] Lexical richness = Number of different words / Total number of words × 100%.

[0102] For example, if a certain text has 1000 words after preprocessing and 600 different words among them, then the lexical richness = 600 / 1000 × 100% = 60%.

[0103] 1.3. Calculation of semantic consistency: Invoke the pre-trained language model (BERT) to calculate the semantic similarity of each pair of adjacent sentences in the text. Input each pair of adjacent sentences into the model respectively to obtain the semantic similarity score (ranging from 0 to 1) output by the model. Then, calculate the average value of the semantic similarities of all adjacent sentences as the semantic consistency index of the text.

[0104] Suppose there are 5 pairs of adjacent sentences in the text, and the calculated semantic similarity scores are: 0.8, 0.6, 0.7, 0.9, 0.8

[0105] Then the semantic consistency = (0.8 + 0.6 + 0.7 + 0.9 + 0.8) / 5 = 0.76.

[0106] 1.4. Calculation of information density: Identify the key information in the text, such as entities (persons, places, organizations, etc.), events (actions, behaviors, etc.). Count the number of key information and the length of the text (in terms of the number of characters or words),

[0107] The information density is calculated using the formula:

[0108] Information density = Number of key information items / Text length × 100%.

[0109] For example, if a text is 500 characters long and contains 50 key information points after key information identification, then the information density is 50 / 500 × 100% = 10%.

[0110] After completing the above calculations, the specific values ​​of the text in the three dimensions of lexical richness, semantic consistency, and information density are obtained, which are 60%, 0.76, and 10%, respectively. These three values ​​are combined into a vector to form the feature vector representation of the text:

[0111] s = [60%, 0.76, 10%].

[0112] Step 2: State Space Construction

[0113] The text feature vectors extracted in step 1 are directly used as the state vectors s in reinforcement learning. That is, each state in the state space consists of three feature dimensions: lexical richness, semantic consistency, and information density.

[0114] For example, the feature vector s = [60%, 0.76, 10%] calculated above represents the state of the current text during the reinforcement learning process. The state space encompasses all possible combinations of text feature vectors, comprehensively reflecting the differences and characteristics of different texts at the feature level, and providing an information basis for subsequent reinforcement learning decisions.

[0115] Step 3: Define the motion space

[0116] The adjustment strategy for credibility scoring is defined as the action space, which specifically includes three types of actions:

[0117] 3.1 Improve Credibility Score: Increase the current credibility score by a fixed step value Δ. For example, if Δ = 5%, and the current credibility score is 50%, after performing this action, the credibility score becomes 50% + 5% = 55%.

[0118] 3.2 Reduce Credibility Score: Decrease the current credibility score by a fixed step value Δ. Again, using Δ = 5% as an example, if the current credibility score is 50%, after performing this action, the credibility score becomes 50% - 5% = 45%.

[0119] 3.3 Maintain Credibility Score: The credibility score remains unchanged and retains its original value. For example, if the current credibility score is 50%, performing this action will keep the score at 50%.

[0120] like Figure 3 As shown, these three actions constitute the action space A = {increase score (+Δ), decrease score (-Δ), remain unchanged (0)}, which provides a selectable operation scheme for reinforcement learning algorithms to dynamically adjust the credibility score according to different states.

[0121] Step 4: Reward Function Design

[0122] 4.1 Accuracy Calculation: The evaluation result is compared with the manually labeled true credibility tag. If the evaluation result matches the true tag, then accuracy = 1; otherwise, accuracy = -1. For example, assuming the manually labeled text has a true credibility tag of "credible", and the evaluation result also determines it to be "credible", then accuracy = 1; if the evaluation result is "unreliable", then accuracy = -1.

[0123] 4.2 Stability Judgment: This checks the fluctuation of N consecutive evaluation results and their closeness to the true label. If the fluctuation range of N consecutive evaluation results is less than or equal to the set threshold θ, and the evaluation results are consistent with the true label, then stability = 1; otherwise, stability = 0. For example, setting N = 3 and θ = 3%, if the last three evaluation results are 50%, 52%, and 49% (fluctuation range ≤ 3%), and the evaluation results are consistent with the true label, then stability = 1; if the fluctuation range exceeds 3% or the evaluation results are inconsistent with the true label, then stability = 0.

[0124] 4.3 task_importance setting: Set the corresponding weight according to the importance of text credibility assessment in different task scenarios.

[0125] For example, task_importance is set to 1.2 for text evaluation in the news domain and 1.0 for general text evaluation in the social media domain.

[0126] Reward Value Calculation: Taking into account the above three factors, the final reward value is calculated according to the formula:

[0127] r=(accuray+stability)×task_importance

[0128] like Figure 4 As shown.

[0129] For example, in a specific evaluation, if the evaluation result is correct (accuracy = 1) and the results are stable for three consecutive evaluations (stability = 1), and the task is a news field evaluation (task_importance = 1.2), then the reward value r = (1+1) × 1.2 = 2.4; if the evaluation result is correct but unstable (stability = 0), then r = (1+0) × 1.2 = 1.2; if the evaluation result is incorrect (accuracy = -1), then r = (-1+0) × 1.2 = -1.2.

[0130] Step 5: Training a Deep Q-Network (DQN)

[0131] 5.1 Network Structure Initialization

[0132] A Deep Q-Network (DQN) is constructed, consisting of an input layer, hidden layers, and an output layer. The input layer has 3 neurons, corresponding to the dimension of the state space, and is used to receive the text's feature vectors (lexical richness, semantic consistency, and information density). The hidden layer contains 64 neurons, using ReLU (Rectified Linear Unit) as the activation function, for non-linear transformation and feature extraction of the input features. The output layer has 3 neurons, corresponding to the three actions in the action space (increase the score, decrease the score, and remain unchanged), representing the Q-value for each action. The network weights are initialized using the Xavier initialization method to ensure a good numerical distribution of the parameters in the initial state.

[0133] For example, the weight matrix from the input layer to the hidden layer:

[0134] W inputhidden ∈R 3 × 64

[0135] According to Xavier's initialization formula:

[0136] W inputhidden ~U(-√ 6 / ( 3 +6 4) ,√ 6 / ( 3 + 64 ))

[0137] That is, they are uniformly distributed within the interval (-√(6 / (3+64)),√(6 / (3+64))); the hidden layer bias B_hidden∈R 64 It is initialized to 0.

[0138] 5.2 Experience Collection and Storage

[0139] Training is performed using a large-scale labeled AIGC text dataset. In each training step, a batch of text samples (e.g., 32 samples) is randomly selected from the dataset. For each sample, its state s (i.e., the feature vectors obtained in steps 1 and 2) is extracted first. Then, an action A (selected from the action space) is executed according to the current confidence scoring strategy, yielding a reward r and a new state s' (a new text feature vector, which may change due to contextual information or user feedback). These experiences (s, A, r, s') are stored in an experience replay buffer. For example, in the early stages of training, the experience replay buffer gradually accumulates experience data until it reaches a set storage capacity (e.g., 1000 experiences).

[0140] 5.3 Network Parameter Update

[0141] Once the number of experiences in the experience replay buffer reaches a certain level, a batch of experiences (e.g., 64 experiences) is randomly sampled from the buffer to update the network parameters of DQN. Stochastic gradient descent (SGD) is used as the optimization algorithm, with a learning rate set to 0.001.

[0142] According to the update formula of the Q-learning algorithm:

[0143]

[0144] Where α is the learning rate (0.001) and γ is the discount factor (0.9). The specific operation is as follows:

[0145] For each sampled experience (s,A,r,s'), input the state s into the current DQN and calculate the current Q value Q(s,A).

[0146] Input the new state s' into DQN, obtain the Q value corresponding to all possible actions A', and take the maximum value maxQ(s',A').

[0147] Calculate the new Q value according to the updated formula:

[0148] Q new (s,a)=Q(s,a)+0.001×[r+0.9×maxQ(s′,a′)-Q(s,a)]

[0149] The mean squared error loss between the predicted Q value and the actual Q value (Q_new(s,A) calculated according to the update formula) is calculated, and the weight parameters of the network are updated through the backpropagation algorithm so that the output of the network gradually approaches the true Q value.

[0150] For example, suppose the Q-value of the current state s is 8, the reward r after performing action A is 2.4, and the maximum Q-value of the new state s' is 10.

[0151] but:

[0152] Q new (s,a)Q new (s,a)=8+0.001×[2.4+0.9×10-8]=8+0.001×3.4=8.0034

[0153] Through continuous iterative training, DQN can learn strategies for selecting the optimal action under different states, maximizing long-term rewards and thus improving the accuracy and reliability of text credibility assessment. Figure 5 As shown.

[0154] Step 6: Contextual Information Fusion

[0155] Contextual Information Collection and Relevance Determination: The context-aware module collects contextual information about the text in real time, including its source (e.g., news websites, academic journals, social media platforms), domain background (e.g., science and technology, culture, sports), and publication context (e.g., news reports, academic research, personal blogs). Simultaneously, based on expert annotations or historical data statistics, the relevance of each type of contextual information to the text credibility assessment is determined. i .

[0156] For example, statistical analysis has found that the credibility of news texts is highly correlated with the authority of the publishing platform, and its relevance... i It might be 0.8; while the credibility of user-generated content on social media is highly correlated with the publisher's past credit history, relevance i It could be 0.6.

[0157] 6.1 Calculation of weighting coefficients:

[0158] Based on the relevance of the collected contextual information, the weight coefficient for each type of contextual information is calculated using the following formula:

[0159]

[0160] Where λ is an adjustment parameter used to control the degree of influence of contextual information relevance on the weights, and can be set to 0.5; relevance i Let represent the correlation between the i-th type of contextual information and the text credibility; n represents the number of types of contextual information.

[0161] For example, assuming we are considering three types of contextual information: news relevance 0.8, academic relevance 0.6, and social media relevance 0.4, then:

[0162] News domain weight:

[0163] w_news = (1 + 0.5 × 0.8) / [(1 + 0.5 × 0.8) + (1 + 0.5 × 0.6) + (1 + 0.5 × 0.4)]

[0164] =1.4 / (1.4+1.3+1.2) =1.4 / 3.9≈0.359

[0165] Academic field weight:

[0166] w_academic = (1 + 0.5 × 0.6) / 3.9 = 1.3 / 3.9 ≈ 0.333

[0167] Social media domain weight:

[0168] w_social = (1 + 0.5 × 0.4) / 3.9 = 1.2 / 3.9 ≈ 0.308

[0169] 6.2 Eigenvector weight adjustment and state vector recalculation

[0170] The weights of each feature in the feature vector are adjusted based on the calculated context information weights.

[0171] Assuming the original state vector is s = [V, s, I], and the corresponding default weight vector is [w v ,w s ,w i If the weights are [0.3, 0.4, 0.3] (e.g., [0.3, 0.4, 0.3]), then the adjusted weight vector is:

[0172] [w v ×w context ,w s ×w context ,w i ×w context ]

[0173] Where w context This corresponds to the weight of the context information. For example, if the text originates from the news field, its context information weight w context =0.359, then the adjusted weight vector is:

[0174] [0.3×0.359,0.4×0.359,0.3×0.359]=[0.1077,0.1436,0.1077].

[0175] Then, the original state vector is weighted and summed using the adjusted weight vector to obtain a new state vector:

[0176] s new =V×(w v ×w context )+S×(w s ×w context)+I×(w i ×w context )

[0177] For example, the original state vector s = [60%, 0.76, 10%], and the adjusted weight vector is [0.1077, 0.1436, 0.1077].

[0178] but:

[0179] s new =60%×0.1077+0.76×0.1436+10%×0.1077≈6.462%+10.91%+1.077%≈18.449%.

[0180] Furthermore, weighted summation compresses the three-dimensional state vector into a single value, potentially leading to information loss. Therefore, in practical applications, the dimensions of the state vector can be kept constant, only the weights of the features in each dimension can be adjusted. This allows the contribution of different features to the Q-value to vary according to the importance of the context information during subsequent Q-network calculations. In other words, the new state vector remains three-dimensional, but each feature value is multiplied by a corresponding weight adjustment coefficient.

[0181] For example, the new state vector:

[0182] s new = [60% × 0.359, 0.76 × 0.359, 10% × 0.359] ≈ [21.54%, 0.272, 3.59%].

[0183] Thus, in s new When input into DQN, the network can perceive the differences in importance of different features in the current context, thereby making more reasonable action choices, such as... Figure 6 As shown.

[0184] Step 7: User Feedback Processing

[0185] like Figure 7 As shown.

[0186] 7.1 Feedback Information Collection and Coding

[0187] The user feedback module receives user feedback on the text credibility assessment results. Feedback can be a user's subjective evaluation of the text credibility (e.g., "credible," "unreliable," "uncertain"), or specific opinions on satisfaction with the assessment results (e.g., "rated too high," "rated too low," "rated reasonable"). User feedback is encoded into a feedback vector f, whose dimensions are the same as the state vector, i.e., the three dimensions correspond to lexical richness, semantic consistency, and information density, respectively. The values ​​of the feedback vector are determined according to the feedback type, for example:

[0188] If a user believes the credibility score should be improved, the feedback vector is f = [0.1, 0.1, 0.1] (i.e., increase by 0.1 in each feature dimension).

[0189] If a user believes the credibility score should be lowered, the feedback vector is f = [-0.1, -0.1, -0.1] (i.e., reduced by 0.1 in each feature dimension).

[0190] 7.2 Feedback Fusion Coefficient Settings

[0191] A feedback fusion coefficient β is determined to control the influence of user feedback on the state vector update. The value of β ranges from 0 to 1 and can be adjusted according to the credibility and importance of the feedback. For example, β can be set to 0.7 for high-credibility feedback from domain experts, and to 0.3 for feedback from ordinary users. Initially, β is assumed to be set to 0.5, indicating that the influence of user feedback is comparable to that of the original state vector.

[0192] 7.3 State Vector Update

[0193] Calculate the new state vector s after fusion based on the feedback fusion formula. new :

[0194] s new =s×(1-β)+f×β

[0195] For example, given the original state vector s = [60%, 0.76%, 10%], the feedback vector f = [0.1, 0.1, 0.1] (the user believes the credibility score should be improved), and β = 0.5, then:

[0196] s n ew=[60%×(1-0.5)+0.1×0.5, 0.76×(1-0.5)+0.1×0.5, 10%×(1-0.5)+0.1×0.5]

[0197] = [30% + 5%, 0.38 + 0.05, 5% + 5%]

[0198] = [35%, 0.43%, 10%]

[0199] 7.4Q Network Reassessment and Action Selection

[0200] The updated state vector s n The input is fed into a deep Q-network (DQN). The DQN recalculates the Q-value for each action based on the current network parameters and selects the action with the largest Q-value as the new evaluation strategy, thereby adjusting the credibility score of the text.

[0201] For example, after calculation, if in sn In the ew state, the Q-value of the action to increase the rating is the largest, so the action to increase the rating is executed, increasing the current credibility rating by Δ (e.g., 5%). In this way, user feedback can be effectively integrated into the evaluation process, allowing the evaluation results to adapt to users' personalized needs and subjective opinions in a timely manner, further improving the reliability and satisfaction of the evaluation results.

[0202] The reinforcement learning module can dynamically adjust the evaluation strategy according to the context of different texts, thereby improving the accuracy and adaptability of the evaluation results.

[0203] Figure 8 This is a flowchart illustrating the overall workflow of the present invention.

[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A dynamic evaluation system for AIGC text credibility based on reinforcement learning, characterized in that: include: The feature extraction module is used to extract multi-dimensional features from AIGC text (AI-generated content) and generate feature vectors. The reinforcement learning module uses a deep Q-network (DQN) to construct a Markov decision process (MDP) and dynamically adjusts the credibility scoring strategy. The context-aware module is used to collect information about the text source, domain background, and publishing scenario; The user feedback module receives user feedback and encodes it into a feedback vector. The evaluation results output module generates a credibility evaluation report. The output of the feature extraction module is connected to the input of the reinforcement learning module. The reinforcement learning module adjusts the scoring strategy based on the information from the context-aware module and the user feedback module. The evaluation result output module receives the output of the reinforcement learning module.

2. The AIGC text credibility dynamic evaluation system based on reinforcement learning according to claim 1, characterized in that: The features extracted by the feature extraction module include lexical richness, semantic consistency, and information density, and irrelevant symbols and stop words are removed and words are segmented through preprocessing.

3. The AIGC text credibility dynamic evaluation system based on reinforcement learning according to claim 1, characterized in that: The action space of the reinforcement learning module includes increasing the credibility score, decreasing the credibility score, or keeping the credibility score unchanged. The step size for increasing or decreasing the score is a fixed percentage Δ.

4. The AIGC text credibility dynamic evaluation system based on reinforcement learning according to claim 1, characterized in that: The context-aware module adjusts the feature vector weights using weighting coefficients: Where λ is the adjustment parameter, relevance i Let represent the correlation between the i-th type of contextual information and credibility, and n represent the number of types of contextual information.

5. A dynamic evaluation method for AIGC text credibility based on reinforcement learning, characterized in that: The steps include: Step 1: Preprocess the text and calculate lexical richness, semantic consistency and information density to generate feature vectors; Step 2: Use the feature vectors as state vectors; Step 3: Define the action space for increasing the credibility score, decreasing the credibility score, or keeping the credibility score unchanged; Step 4: Design a reward function based on the evaluation accuracy, stability, and task importance; Step 5: Train a deep Q-network (DQN) to dynamically adjust the scoring strategy; Step 6: Integrate contextual information to update the state vector; Step 7: Process user feedback and reassess credibility.

6. The AIGC text credibility dynamic evaluation method based on reinforcement learning according to claim 5, characterized in that: Step 1 includes: Vocabulary richness calculation: The number of distinct words refers to the number of unique words in the text, and the total number of words refers to the total number of words after preprocessing.

7. The AIGC text credibility dynamic evaluation method based on reinforcement learning according to claim 5, characterized in that: In step 4, the reward function is: r=(accuray+stability)×task_importance Here, accuracy indicates the consistency between the evaluation result and the true label; stability indicates that it is 1 when the fluctuation range of N consecutive evaluations is ≤ the threshold θ, otherwise it is 0; task_importance indicates the task scenario weight.

8. The AIGC text credibility dynamic evaluation method based on reinforcement learning according to claim 5, characterized in that: Step 5 includes: Deep Q-Network (DQN) updates parameters according to the Q-learning formula: Where α is the learning rate and γ is the discount factor.

9. The AIGC text credibility dynamic evaluation method based on reinforcement learning according to claim 5, characterized in that: Step 6 includes: The state vector update method is as follows: s new =[V×w context ,S×w context ,I×w context ] Where V, S, and I are the original values ​​of lexical richness, semantic consistency, and information density, respectively, and w context These are the weighting coefficients.

10. The AIGC text credibility dynamic evaluation method based on reinforcement learning according to claim 5, characterized in that: Step 7 includes: User feedback is encoded as a feedback vector f, and the state vector update formula is: s new =s×(1-β)+f×β Where β is the feedback fusion coefficient, and its value ranges from 0 to 1.