Value-oriented content analysis and intervention method in ideological and political education

By constructing a multimodal value knowledge graph and a multimodal understanding model, combined with reinforcement learning and counterfactual reasoning, the problem of the difficulty in identifying and intervening in implicit values ​​in ideological and political education has been solved, the accurate identification and quantitative evaluation of explicit and implicit values ​​have been achieved, personalized education intervention strategies have been generated, and the pertinence and adaptability of education have been improved.

CN120764804AInactive Publication Date: 2025-10-10QINGDAO HUANGHAI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510848377.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-10
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120764804A_ABST
    Figure CN120764804A_ABST
Patent Text Reader

Abstract

The invention provides a value-oriented content analysis and intervention method in ideological and political education, and belongs to the technical field of education and teaching. A pre-trained multi-modal value understanding model is combined with a recurrent neural network to extract dominant value features and capture context dependence, implicit value expressions are recognized through an attention mechanism and comparative learning, and a tendency index is generated based on the similarity between feature vector calculation and core value. A reinforcement learning training content intervention agent is utilized to generate an education intervention strategy, and finally anti-fact reasoning is adopted to evaluate different intervention strategy effects and optimize an intervention scheme, so that a complete technical closed loop from content analysis to intervention implementation to effect evaluation is formed. The technical problem that the implicit value in the ideological and political education content is difficult to quantify to accurately identify and implement effective intervention is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of education and teaching technology, and specifically relates to a content analysis and intervention method for value orientation in ideological and political education. Background Art

[0002] Ideological and political education is a key component of cultivating correct values ​​in college students. Traditionally, this approach relies primarily on teacher experience and textbook content to guide values. Existing technologies analyze value education content based on methods such as keyword matching, sentiment analysis, and topic modeling, identifying value expressions within text using pre-set vocabulary or rules. In practical applications, these methods are used in textbook review, curriculum design, and teaching evaluation, presenting the distribution of values ​​through statistical analysis.

[0003] However, traditional techniques have significant limitations. First, rule-based and vocabulary-based approaches can only identify explicit expressions of values, with low recognition rates for implicit expressions such as euphemism, metaphor, and irony. Second, single-modality analysis cannot handle the multimodal expressions of values ​​found in contemporary educational resources, such as images, text, and videos. Third, existing intervention methods are often static and lack adaptive adjustment mechanisms for individual differences, resulting in unstable intervention effects.

[0004] Faced with the complex and ever-changing forms of value expression, especially implicit value content, existing technologies have difficulty achieving accurate identification and quantitative evaluation. Simultaneously, there is a lack of scientific evidence-based intervention strategy generation and effect evaluation mechanisms. This seriously restricts the pertinence and effectiveness of ideological and political education. There is an urgent need for a technical solution that can accurately identify multimodal implicit values ​​and implement intelligent intervention. In other words, existing technologies have the technical problem of accurately identifying and effectively intervening in the difficult-to-quantify implicit values ​​in ideological and political education content. Summary of the Invention

[0005] In view of this, the present invention provides a content analysis and intervention method for value orientation in ideological and political education, which can solve the technical problem in the existing technology that it is difficult to accurately identify and effectively intervene in the implicit values ​​in the content of ideological and political education.

[0006] The present invention is achieved as follows: The present invention provides a value-oriented content analysis and intervention method in ideological and political education, including: constructing a multimodal value knowledge graph and integrating it into the mainstream social value system as a basic semantic network; using a multimodal value understanding model to extract explicit value features and using a recursive neural network to capture contextual dependencies; applying an attention mechanism to analyze the connection between key words and visual elements in the content to identify implicit value expressions; constructing a positive and negative sample pair learning value expression difference enhancement model through a comparative learning method to recognize rhetorical techniques; calculating the cosine similarity between the content and the core value vector based on the value feature vector to generate a value tendency index; using reinforcement learning technology to train a content intervention agent to generate an education intervention strategy based on the value tendency index; using a counterfactual reasoning method to evaluate the impact of different intervention strategies on the ideological and political education effect and optimize the intervention plan.

[0007] Among them, the multimodal value knowledge graph is a structured knowledge network constructed by the value concepts, expressions and associations in various media forms, which includes entity nodes, attributes and relationship edges, and can represent the mapping relationship and semantic connection between value expressions in different modalities.

[0008] Among them, implicit value expression is the transmission of values ​​and positions through indirect rhetoric techniques such as euphemism, metaphor, and irony, or non-explicit expressions. It is usually necessary to comprehensively judge the true tendency based on the context, cultural background and expression method.

[0009] Among them, the value orientation index is a quantitative score obtained by calculating the similarity between the content feature vector and the preset value vector. It is used to measure the degree of fit between content and values. The value range is 0 to 1. The higher the value, the higher the fit.

[0010] Among them, the counterfactual reasoning module is a computing unit built based on the principle of causal inference. By simulating possible results under different intervention strategies, it evaluates the effectiveness and applicability of various intervention plans and thus selects the optimal intervention path.

[0011] Among them, the multimodal values ​​comprehension model is a deep neural network model that is trained with large-scale corpus and multimodal content and is used to understand and analyze explicit and implicit value expressions in various media. It uses an encoder-decoder architecture to integrate cross-modal information and has a culturally sensitive adaptation mechanism.

[0012] Among them, the specific structure of the multimodal value understanding model is a deep neural network based on the stacking of multi-head self-attention mechanism and cross-modal interaction layers, which includes four core components: text encoding module, visual encoding module, audio encoding module and cross-modal fusion module. The text encoding module adopts a hierarchical bidirectional converter structure to process text information from character to sentence to paragraph level. The visual encoding module uses visual converters to extract spatial features and temporal relationships in images and video frames. The audio encoding module extracts emotional and intonation features in sound information through one-dimensional convolutional networks and long short-term memory networks. The cross-modal fusion module adopts a sparse multi-head attention mechanism to realize the dynamic weight integration of different modal information.

[0013] Among them, the sparse factor parameters in the sparse multi-head attention mechanism are dynamically adjusted according to the semantic density of the value knowledge graph, the complexity of the content and the contribution of each segmented text clustering center. The number of attention heads is determined by the difficulty of identifying implicit value expressions and the complexity of rhetorical techniques. The final output layer of the model generates a value tendency feature vector through nonlinear activation function mapping.

[0014] Among them, the contribution of the cluster center refers to the weight of the influence of each cluster center on the overall value expression when cluster analysis is performed after text information segmentation. It is quantified by calculating the semantic relevance of each cluster center with the value standard vector and its coverage in the text. Cluster centers with high contribution will obtain more attention resources allocation in the sparse attention mechanism.

[0015] Among them, the steps for establishing a training data set for the multimodal values ​​​​understanding model include collecting educational texts, images and video resources with clear value labels from multi-source databases, manually annotating all collected resources to determine the value category and intensity, dividing the data into two categories of explicit expression and implicit expression according to the content expression method and performing detailed annotation, constructing a sample set containing positive and negative values ​​for comparison, and stratifying the data according to cultural background and contextual factors to ensure that the training set covers multiple expression scenarios.

[0016] This method, by constructing a multimodal values ​​knowledge graph and a values ​​comprehension model, combined with techniques such as attention mechanisms and contrastive learning, enables accurate identification and quantitative assessment of explicit and implicit value expressions. This method automatically calculates the degree of alignment between content and core values, generating a value orientation index that provides an objective basis for educational intervention.

[0017] In technical implementation, the application breaks through the bottleneck of identifying implicit values in the prior art, captures context dependency through a recurrent neural network, analyzes key element connections through an attention mechanism, and significantly improves the recognition accuracy of indirect expression methods such as euphemism and metaphor. At the same time, the intervention agent based on reinforcement learning and the counterfactual reasoning evaluation method realize the intelligent generation and optimization of intervention strategies, making the intervention measures more targeted and adaptive.

[0018] The present application solves the technical problem of accurately identifying and effectively intervening in the current difficulty of quantifying the implicit values in the ideological and political education content by closely combining multi-modal content analysis with intelligent intervention strategies. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 The flowchart of the method of the present application. DETAILED DESCRIPTION

[0020] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application.

[0021] As Figure 1 shown is a flowchart of a value-oriented content analysis and intervention method in ideological and political education provided by the present application, the method includes the following steps:

[0022] S01, construct a multi-modal value knowledge graph and integrate a mainstream value system as a basic semantic network;

[0023] S02, use a pre-trained multi-modal value understanding model to extract explicit value features in text and video and use a recurrent neural network to capture context dependency;

[0024] S03, apply an attention mechanism to analyze the relationship between key words and visual elements in the content to identify implicit value expressions;

[0025] S04, construct a positive and negative sample pair learning value expression difference enhancement model through a contrast learning method to enhance the ability to identify rhetorical devices;

[0026] S05, calculate the content and core value vector cosine similarity to generate a value orientation index;

[0027] S06, use reinforcement learning technology to train a content intervention agent to generate an education intervention strategy according to the value orientation index;

[0028] S07, use a counterfactual reasoning method to evaluate the effect of different intervention strategies on ideological and political education and optimize the intervention scheme.

[0029] Among them, the multimodal value knowledge graph refers to the construction of value concepts, expressions and their associations in various media forms such as text, images, and videos into a structured knowledge network, which includes entity nodes, attributes and relationship edges, and can represent the mapping relationship and semantic connection between value expressions in different modalities.

[0030] Among them, implicit value expression refers to the values ​​and positions conveyed through indirect rhetoric such as euphemism, metaphor, irony, or non-explicit expressions. It is usually necessary to comprehensively judge its true tendency based on the context, cultural background, and expression method.

[0031] Among them, the value orientation index is a quantitative score obtained by calculating the similarity between the content feature vector and the preset value vector. It is used to measure the degree of fit between content and values. The value range is 0 to 1. The higher the value, the higher the fit.

[0032] Among them, the counterfactual reasoning method utilizes a counterfactual reasoning module, which is a computing unit built based on the principle of causal inference. By simulating the possible results under different intervention strategies, it evaluates the effectiveness and applicability of various intervention plans and thus selects the optimal intervention path.

[0033] Among them, the multimodal values ​​comprehension model refers to a deep neural network model that is trained with large-scale corpus and multimodal content to understand and analyze the explicit and implicit value expressions in various media. The model uses an encoder-decoder architecture to integrate cross-modal information and has a culturally sensitive adaptation mechanism.

[0034] The specific structure of the multimodal value understanding model is a deep neural network formed by stacking multi-head self-attention mechanisms and cross-modal interaction layers, which includes four core components: text encoding module, visual encoding module, audio encoding module and cross-modal fusion module. The text encoding module adopts a hierarchical bidirectional converter structure to process text information from character to sentence to paragraph level. The visual encoding module uses a visual converter to extract spatial features and temporal relationships in images and video frames. The audio encoding module extracts emotional and intonation features in sound information through a one-dimensional convolutional network and a long short-term memory network. The cross-modal fusion module adopts a sparse multi-head attention mechanism to realize dynamic weight integration of different modal information. The sparse factor parameter of the sparse attention mechanism is dynamically adjusted according to the semantic density, content complexity and cluster center contribution of each segmented text of the value knowledge graph. The number of attention heads is determined by the difficulty of recognizing implicit value expressions and the complexity of rhetoric. The final output layer of the model generates a value tendency feature vector through nonlinear activation function mapping.

[0035] The steps of establishing a training data set for the multimodal value understanding model specifically include collecting educational text, image, and video resources containing clear value labels from multi-source databases, manually annotating all collected resources to determine their value categories and strengths, dividing the data into two categories of explicit expression and implicit expression according to the content expression method and performing subdivision annotation, constructing a sample set containing positive and negative value comparisons, stratifying the data according to cultural background and contextual factors to ensure that the training set covers multiple expression scenarios, expanding rare and difficult-to-obtain value expression samples through data augmentation technology, using knowledge distillation methods to extract value judgment basis from expert evaluations to form auxiliary training signals, and finally establishing a validation set and a test set for model performance evaluation and generalization ability testing;

[0036] The steps of training the multimodal value understanding model specifically include first conducting self-supervised learning on a general large-scale corpus to acquire basic language comprehension capabilities, then conducting cross-modal alignment training on a multimodal general dataset to establish semantic mapping relationships between different modalities, then conducting supervised fine-tuning on a value annotation dataset to acquire value recognition capabilities, then adopting a contrastive learning strategy to enable the model to learn to distinguish subtle differences between similar value categories, incorporating expert knowledge into model parameters through knowledge distillation technology, utilizing adversarial training to improve the model's robustness in recognizing implicit value expressions, optimizing the model's rapid adaptability in small sample situations based on a meta-learning method, and finally improving the model's performance in actual application scenarios of ideological and political education through cyclic evaluation and parameter tuning;

[0037] Among them, the contribution of the cluster center refers to the weight of the influence of each cluster center on the overall value expression when cluster analysis is performed after text information segmentation. It is quantified by calculating the semantic relevance of each cluster center with the value standard vector and its coverage in the text. Cluster centers with high contribution will obtain more attention resource allocation in the sparse attention mechanism.

[0038] Among them, educational intervention strategies are a group of values ​​deviation types P i and the degree of deviation D i Determine the mapping function S(P i , D i )→A i , where A i Represents a specific intervention action. Each intervention strategy is defined as a triple (T j , M j , I j ), T j Represents intervention type (T1: cognitive correction, T2: emotional guidance, T3: behavioral guidance), M iRepresentative intervention methods (M1: direct explicit, M2: indirect suggestion, M3: case comparison, M4: interactive discussion, M5: practical experience), I j Represents the intervention intensity (range [0, 1]). The intervention strategy generation process is modeled using a Markov decision process (MDP). The state space S is the learner's current value state vector, the action space A is the set of optional intervention strategies, the transfer function P(s′|s, a) describes the probability of state change after the intervention, and the reward function R(s, a) measures the consistency between the intervention effect and the target value. Through the policy gradient method π*(s) = argmax a ∑ s′ P(s′|s,a)|R(s,a)+γV(s′)] learns the optimal intervention strategy, where V(s) is the state value function and γ is the discount factor (usually set to 0.85 to 0.95). The strategy evaluation uses the cumulative reward expectation E[∑ t γ t R(s t , a t )] to ensure maximum intervention effect.

[0039] Among them, optimizing the intervention plan is a multi-objective combinatorial optimization problem, which is formally defined as finding the intervention strategy sequence A*={a1,a2,...,a n} Make the objective function vector F(A) = [f1(A), f2(A), f3(A), f4(A)] reach the optimal value, where f1(A) represents the effectiveness of the intervention (the degree of correction of value bias), f2(A) represents the efficiency of the intervention (minimization of resource consumption), f3(A) represents the adaptability of the intervention (the breadth of applicable learner groups), and f4(A) represents the persistence of the intervention (how long the effect lasts). The optimization process uses a multi-objective particle swarm algorithm, where each particle represents a candidate intervention plan, and the non-dominated plans are screened through the Pareto front. Plan evaluation function where w i is the weight of each target (∑w i =1), O i (A) is the normalized objective function value. The implementation sequence of the program adopts the principle of progressive intensity, from I1≤I2≤...≤I n Or I1≥I2≥...≥I n Determine (depending on the learner acceptance assessment results). Intervention intensity dynamic adjustment function I′ j =I j ×α(R j ), where α(R j ) is based on real-time feedback R j The adjustment coefficient ranges from [0.7 to 1.3]. The program evaluation cycle is set to be conducted once after each intervention and after three cumulative interventions to ensure the optimal program execution path.

[0040] Among them, the counterfactual reasoning method is based on the structural causal model SCM=(U, V, F), where U is the set of exogenous variables (learner background characteristics), V is the set of endogenous variables (including intervention strategy variables X and value outcome variables Y), and F is the set of functions that describe the causal relationship between variables. Counterfactual reasoning calculates P(Y=y|do(X=x)) by the intervention exponent do(X=x) to evaluate the causal effect of intervention X=x on outcome Y=y. The intervention effect is measured using the average treatment effect ATE=E[Y|do(X=x1)]-E[Y|do(X=x0)] and the conditional average treatment effect CATE(z)=E[Y|do(X=x1), Z=zl-E[Y|do(X=x0), Z=zl, where Z is a covariate (learner feature grouping). The specific implementation uses a dual robust estimator. Where m1 and m0 are the result models, and e(Z) is the propensity score. The accuracy of counterfactual prediction is evaluated using the counterfactual verification error. CFE<0.15 is required to ensure the reliability of prediction. The optimal intervention path selection is based on the principle of maximizing the counterfactual prediction utility, i.e. argmax x E[U(Y)|do(X=x)], where U(Y) is the value outcome utility function.

[0041] The specific implementation of the above steps is described in detail below.

[0042] The specific implementation of step S01 involves constructing a multimodal value knowledge graph and integrating the mainstream social value system into it as a foundational semantic network. This implementation process first collects multimodal data related to values, such as text, images, and videos. Natural language processing techniques are used to extract value entities and their relationships. Convolutional neural networks are used to identify value-related visual elements from images and videos. Ontology modeling is then used to define the concept hierarchy and semantic associations of values. The value entities are then vectorized, and transfer learning is used to transfer knowledge from pre-trained language and visual models to the value domain. Graph embedding algorithms such as TransE or ComplEx are used to map the entities and relationships into a low-dimensional vector space. The concepts of the mainstream social value system are then used as core nodes and paths in the knowledge graph. Knowledge fusion techniques are used to integrate the core values ​​with the existing knowledge graph to establish a weighted system for value orientations. Finally, a graph database is used to store the constructed knowledge graph, enabling fast retrieval and reasoning. A regular update mechanism is used to maintain the timeliness and accuracy of the knowledge graph. The association strength threshold between knowledge graph nodes is set at 0.65; associations below this threshold are filtered to ensure graph quality. The purpose of this step is to provide a knowledge base for subsequent analysis, so that the identification and judgment of values ​​have a unified standard and reference system.

[0043] The specific implementation of step S02 involves using a pre-trained multimodal value understanding model to extract explicit value features from text and video, and using a recurrent neural network to capture contextual dependencies. This implementation process first inputs the text and video content to be analyzed into the multimodal value understanding model, and uses a pre-trained encoder to extract text and video features. Text features are extracted using a multi-layer bidirectional transformer network, which considers the contextual relationships between words; video features are extracted using a visual transformer and a three-dimensional convolutional neural network, which considers both spatial and temporal information. The extracted feature vectors are mapped to the value feature space using an attention pooling mechanism to form explicit value representation vectors. A recurrent neural network, such as a long short-term memory network or a gated recurrent unit, is then used to model the context sequence, capturing long-range dependencies and contextual information. The hidden layer state dimension of the recurrent network is set to 512, and the time step size is dynamically adjusted based on the content length, supporting a maximum of 1024 time steps. Finally, the output of the recurrent neural network is fused with the explicit value representation vector to generate a contextualized value feature representation. The confidence threshold for explicit feature extraction in this step is set to 0.75; features below this threshold are marked as uncertain. This step aims to accurately identify the value elements explicitly expressed in the content and understand their true meaning and orientation in context.

[0044] The specific implementation of step S03 involves applying an attention mechanism to analyze the connection between key words and visual elements in the content to identify implicit value expressions. This implementation process first segments the input multimodal content, establishing a temporal alignment between each text paragraph and its corresponding visual element. A multi-head attention mechanism is then used to calculate the attention weight matrix between key words and visual elements within the text, with the number of attention heads set to 8, each with a dimension of 64. This attention distribution identifies highly correlated regions within the text and visual content, which often contain implicit value expressions. The matching degree between key words and the value knowledge graph is then used to analyze the implicit semantics of words, using a context-sensitive word embedding model to capture the semantic variations of the same word in different contexts. For visual elements, image sentiment analysis and symbol recognition algorithms are used to detect visual cues that may contain value implications. Finally, the implicit features of the text and visual modalities are cross-fused, and an adaptive weighting method is used to integrate multimodal information to generate an implicit value representation vector. The confidence threshold for implicit value determination is set to 0.6, and it requires support from at least two modal evidence. The purpose of this step is to discover the value tendencies that are not directly stated but implied in the content, and to improve the ability to understand rhetorical devices such as euphemisms, metaphors and irony.

[0045] The specific implementation method of step S04 is to construct positive and negative sample pairs through the contrastive learning method, learn the differences in value expression, and enhance the model's recognition ability of rhetoric. The implementation process first selects content pairs that express similar values ​​but use different rhetorical methods from the training data as positive sample pairs, and selects content pairs that express opposite values ​​but use similar expressions as negative sample pairs. Each sample pair contains the original content feature vector and the corresponding value label. The contrastive loss function is used for training so that the model learns to map different expressions of the same value to similar feature spaces, while distinguishing the expressions of different values. The contrastive loss uses cosine similarity measurement, and the temperature parameter is set to 0.07. A difficult sample mining strategy is introduced during the training process, focusing on easily confused value expressions, such as the distinction between euphemistic expressions and ironic expressions. In addition, data enhancement technology is implemented to generate more diverse expression variants, including synonym replacement, syntactic reconstruction, and context transformation. Model training uses a small batch stochastic gradient descent algorithm with a batch size of 128, and the learning rate is initially set to 10 -4 , using a cosine annealing strategy for adjustment. The similarity threshold for contrastive learning is set to 0.85; pairs of samples above this threshold are considered to have the same value orientation. This step enhances the model's ability to distinguish subtle differences in values, especially for content expressed using complex rhetoric.

[0046] The specific implementation of step S05 involves calculating the cosine similarity between the content and the core value vector based on the value feature vector to generate a value orientation index. This implementation process first extracts standard vector representations of mainstream social values ​​from the value knowledge graph, which was generated during the knowledge graph construction process. For the content to be analyzed, the explicit and implicit value feature vectors extracted in the previous steps are combined and integrated using an attention-weighted method to generate a comprehensive value feature vector for the content. The cosine similarity between the content feature vector and each core value vector is calculated using the dot product of the two vectors divided by the product of their respective moduli. For multi-dimensional value assessments, the similarity between the content and each dimension is calculated to form a multi-dimensional value orientation vector. Finally, the multi-dimensional similarities are integrated into a single value orientation index using a weighted average method, with weights pre-set based on the importance of each value dimension. The value orientation index ranges from 0 to 1, with higher values ​​indicating greater alignment with core values. The threshold for determining content as positively oriented is set at 0.75, with values ​​between 0.4 and 0.75 being neutral and below 0.4 being negative. The purpose of this step is to quantify the degree of consistency between the assessment content and core values, and provide a decision-making basis for subsequent intervention strategies.

[0047] The specific implementation of step S06 is to train the content intervention agent using reinforcement learning techniques to generate education intervention strategies based on the value orientation index. The implementation process first constructs a Markov decision process framework, and the state space is defined as the current value orientation state vector of the learner, which is represented by the value orientation index calculated in the previous step. The action space is defined as the set of available intervention strategies, including different combinations of intervention types, methods and intensity. The reward function is designed to improve the degree of similarity between the value orientation state after intervention and the target value orientation state, and to consider the punishment term of intervention resource consumption. The intervention agent is trained using deep Q-network or policy gradient method, with the current state as input and the optimal intervention strategy as output. The Q-network uses a three-layer fully connected neural network with 256 and 128 hidden layer nodes, and the activation function is ReLU. To balance exploration and utilization, an ε-greedy strategy is used, with an initial ε value of 0.3, gradually decaying to 0.05. During the training process, an experience replay buffer is used to store transition samples, with a buffer size of 10 5 , and a batch size of 64 samples is randomly sampled for training each time. The target network is updated every 100 steps, and the discount factor γ is set to 0.9. The evaluation of intervention strategies uses the A / B testing method to compare the differences in effects under different strategies. The strategy selection confidence threshold is set to 0.7, and cases below this threshold will use a conservative intervention strategy. The role of this step is to automatically generate personalized education intervention strategies based on the current state of the learner, achieving intelligent and precise ideological and political education.

[0048] The specific implementation of step S07 is to evaluate the impact of different intervention strategies on the effectiveness of ideological and political education using counterfactual reasoning methods and optimize the intervention plan. The implementation process first constructs a structural causal model, defining variables including learner characteristics, intervention strategies, and education effectiveness, etc. Based on historical intervention data, a causal effect estimation model is trained, using a double robust estimator combined with propensity score and outcome regression model. The propensity score model uses logistic regression or random forest algorithm to predict the probability of learners receiving intervention strategies; the outcome regression model uses multilayer perceptron to predict the education effectiveness under given intervention. For new situations, the trained model is used to generate counterfactual predictions, i.e. to estimate "what effect will be produced if strategy B is used for learner A". By comparing the predicted effects of different intervention strategies, the optimal intervention path is selected. To handle multi-objective optimization problems, an improved Pareto optimization method is used, considering four objectives of intervention effectiveness, efficiency, adaptability and durability. The multi-objective particle swarm optimization algorithm is used to find the Pareto optimal solution set, and each particle represents a candidate intervention scheme, whose position and speed are defined by the intervention strategy parameters. The scheme selection considers the weights of each target and combines the decision maker's preferences. The implementation order of the intervention scheme adopts the principle of progressive intensity, and is dynamically adjusted according to the acceptance of learners. The accuracy threshold of counterfactual prediction is set to 0.15, and the prediction higher than this threshold is considered unreliable. The role of this step is to find the most effective combination of intervention strategies through simulation analysis, to ensure that the education intervention achieves the expected effect, and to optimize the use of resources.

[0049] The specific structure of the multi-modal value understanding model is a deep neural network system based on multi-head self-attention mechanism and cross-modal interaction layer stacking, which contains four core components. The text encoding module adopts a hierarchical bidirectional transformer structure, including three encoding levels: the character-level encoding layer uses word embedding and a one-dimensional convolutional network to capture morphological features; the sentence-level encoding layer uses a 6-layer transformer encoder to process the word order relationship within the sentence, with a hidden state dimension of 768 and 12 attention heads; the paragraph-level encoding layer uses self-attention pooling to aggregate sentence representations and generate text semantic vectors. The visual encoding module uses a visual transformer structure to divide images or video frames into 16x16 pixel image blocks, which are input into a 12-layer transformer after position encoding, with a hidden state dimension of 1024 to capture spatial relationships; for videos, a temporal attention layer is added to model inter-frame relationships. The audio encoding module extracts spectrogram features through a one-dimensional convolutional network with 64, 128, and 256 filters and a kernel size of 3, followed by a bidirectional long short-term memory network with a hidden state size of 512 to capture speech emotion and rhythm features. The cross-modal fusion module uses a sparse multi-head attention mechanism, with an attention sparsity factor dynamically adjusted based on the semantic density, content complexity, and cluster center contribution of the value knowledge graph, with a sparsity range of 0.3 to 0.7; the number of attention heads is adaptively adjusted based on the difficulty of recognizing implicit value expressions, with a range of 4 to 16. The final output layer maps the value orientation feature vector through a nonlinear activation function, with a vector dimension of 128.

[0050] The training data set establishment process of the multi-modal value understanding model includes seven key steps. First, collect resources such as teaching materials, film and television works, and public speeches with value labels from multiple sources such as education resources, news media, and social networks. The total amount of collected resources is not less than 10 6 thousands, covering multiple modalities such as text, images, audio, and video. Second, manually annotate all collected resources by not less than 5 professional annotators to determine their value categories and intensity, using a seven-point scale to score, with annotation consistency evaluated by Kappa coefficient, requiring greater than 0.8. The third step is to divide the data into two categories according to the content expression method: explicit expression refers to the content that directly and explicitly states the value point, and implicit expression refers to the content that indirectly conveys the value through rhetorical devices; further subdivide to indicate specific expression skill types such as metaphor, analogy, and irony. The fourth step is to construct a contrast sample set containing positive and negative values, with each contrast sample containing a pair of content with similar themes but different value orientations, a total of not less than 10 5Group comparison samples. The fifth step is to stratify the data sampling based on cultural background and contextual factors to ensure that the training set covers value expression scenarios in different regions, age groups and educational backgrounds, and the sample distribution matches the population distribution of the target application scenario. The sixth step is to expand rare value expression samples through data enhancement technology, including synonym replacement, syntactic reorganization and style transfer, and amplify rare category samples by 5 to 10 times. The last step is to use the knowledge distillation method to extract the basis for value judgment from expert evaluations, establish an expert explanation database, and use it as an auxiliary training signal to guide the model to learn human value judgment rules; finally, establish a validation set and a test set, accounting for 10% and 15% of the total data respectively, for model performance evaluation and generalization ability testing.

[0051] The training process of the multimodal value understanding model consists of seven stages. First, self-supervised learning is performed on a general large-scale corpus, using masked language modeling and contrastive prediction tasks for pre-training. The corpus size is at least 50GB of text data. Second, cross-modal alignment training is performed on a general multimodal dataset, using a contrastive learning objective to minimize the feature distance between matching modal pairs and maximize the feature distance between mismatching modal pairs. The temperature parameter is set to 0.1. Third, supervised fine-tuning is performed on a value annotation dataset, using a cross-entropy loss function to optimize the value classification task. The learning rate is set to 10 and the number of training epochs is 10. Fourth, a contrastive learning strategy is used to enable the model to learn to distinguish subtle differences between similar value categories. Pairs of difficult-to-separate examples are constructed for training, and the contrastive loss weight is set to 0.3. Fifth, expert knowledge is incorporated into the model parameters through knowledge distillation. The teacher model is manually annotated expert judgment results. The distillation temperature parameter is set to 2.0, and the distillation loss weight is set to 0.5. In the sixth stage, adversarial training is used to improve the model's robustness to implicit value expressions. The perturbation strength of the generated adversarial samples is 0.02, and the adversarial training ratio is 30% of the total samples. In the final stage, the model's rapid adaptability in small sample situations is optimized based on the meta-learning method. The model-independent meta-learning algorithm is used to support rapid adaptation of only 5 to 10 samples per class. The meta-batch size is set to 4 and the meta-learning rate is 0.01. The entire training process uses mixed precision training technology, uses the Adam optimizer, and the weight decay is 10 -6 , improve the actual application performance of the model through cyclic evaluation and parameter tuning.

[0052] The mathematical model or calculation process involved in the present invention is described in detail below.

[0053] In step S02, a multimodal value understanding model is used to extract explicit value features and a recursive neural network is used to capture contextual dependencies. The calculation process involved is specifically represented as follows:

[0054] h t=tanh(W x x t +W h h t-1 +b h );

[0055] Where h t is the hidden state vector at time t, with a dimension of 512; x t is the input feature vector at time t; h t-1 is the hidden state vector at time t-1; W x is the weight matrix input to the hidden layer, with a size of d x ×d h , where d x is the input feature dimension, d h is the hidden state dimension 512; W h is the weight matrix from hidden layer to hidden layer, with size d h ×d h ; b h is the hidden layer bias vector, dimension d h ; tanh is the hyperbolic tangent activation function.

[0056] This recursive computation process is based on the principle of sequence model, combining the current input with historical information through weight parameters to capture long-range dependencies in text or video sequences. x and W h Parameters are optimized using a backpropagation algorithm based on large-scale training data, with initial values ​​initialized using the Xavier method. This formula effectively models context, considers the influence of historical information on the interpretation of current features, and overcomes the limitation of simple feedforward networks that cannot handle sequential dependencies.

[0057] In step S03, when applying the attention mechanism to analyze the connection between keywords and visual elements, the multi-head attention calculation process involved is specifically represented as follows:

[0058]

[0059] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O ;

[0060] head i =Attention(QW i Q , KW i K , VW i V );

[0061] where Q is a query matrix with size n q ×d model ; K is a key matrix with size n k ×d model ; V is a value matrix with size n k ×d model ; d k is the dimension of key vector; W i Q , W i K , W i V are the projection matrices of the i-th attention head for query, key, value respectively, with size d model ×d q , d model ×d k , d model ×d v , respectively, where d q = d k = d v = d model / h = 64, h is the number of attention heads, and the value is 8; W O is an output projection matrix with size hd v ×d model ; softmax is a normalization exponential function; and Concat is a concatenation operation.

[0062] This multi-head attention mechanism realizes parallel attention on different aspects of content by projecting the input into different subspaces. The scaling factor is used to prevent the problem of gradient disappearance caused by too large inner product. The attention mechanism enables the system to dynamically calculate the correlation strength between text and visual elements, adaptively allocate attention resources, highlight key information, and improve the model's detection ability of implicit values. Compared with the traditional fixed weight method, it can better handle dynamic correlations in multi-modal information.

[0063] In step S04, when constructing positive and negative sample pairs using the contrast learning method, the calculation process of the contrast loss function is specifically represented as follows:

[0064]

[0065] where L contrastive is the contrast loss value; f(x) is a feature representation function of content x; x i is an anchor sample; is a positive sample with the same value expression as the anchor sample; x j is all samples in a batch; s() is a similarity function, and the cosine similarity τ is the temperature parameter, and its value is 0.07; N is the batch size, and its value is 128.

[0066] The contrastive loss function, based on the principle of mutual information maximization from information theory, aims to maximize the feature similarity between expressions of the same value while minimizing the feature similarity between expressions of different value. The temperature parameter τ controls the smoothness of the loss function's gradient. A smaller τ value causes the model to focus more on difficult-to-distinguish pairs of examples. By bringing semantically similar examples closer together and pushing semantically dissimilar examples further apart, this loss function enhances the model's ability to discriminate between expressions of the same value expressed in different rhetorical devices, overcoming the limitations of traditional classification methods that struggle to handle subtle semantic differences.

[0067] In step S05, the calculation process of calculating the cosine similarity between the content and the core value vector to generate the value orientation index is specifically represented as follows:

[0068]

[0069] Where, v content is the comprehensive value feature vector of the content, with a dimension of 128; is the standard vector representation of the i-th core value, and the dimension is also 128; similarity() is the cosine similarity function; w i is the weight coefficient of the i-th core value, satisfying The weight value is pre-set according to the importance of the value, ranging from 0.05 to 0.3; M is the total number of core values; Index value It is the final value orientation index, ranging from 0 to 1.

[0070] This calculation process, based on the principles of the vector space model, quantifies semantic similarity by calculating the cosine of the angle between vectors. The introduction of weighting coefficients takes into account the differential importance of different value dimensions, making the evaluation results more consistent with practical application needs. This method transforms complex value analysis into a computable similarity problem, enabling a quantitative assessment of the value orientation of content and providing an objective basis for subsequent intervention strategies.

[0071] In step S06, when using reinforcement learning to train the content intervention agent, the policy gradient calculation process is specifically expressed as follows:

[0072]

[0073] π * (s)=argmax d ∑ s′ P(s′|s,a)[R(s,a,s′)+γV(s′)];

[0074] wherein. is the policy gradient; θ is the policy network parameter; θ (a|s) is the parameterized policy function, representing the probability of choosing action a in state s; Q π (s,a) is the state-action value function, representing the expected cumulative reward of taking action a in state s following policy π; γ is the discount factor, taking value 0.9; r t is the reward obtained at time t; P(s'|s,a) is the state transition probability, representing the probability of transitioning to state s' after taking action a in state s; R(s,a,s') is the immediate reward function; V(s) is the state value function; π * (s) is the optimal policy.

[0075] The policy gradient algorithm is based on the principle of gradient ascent, adjusting the parameters to increase the probability of high reward actions and decrease the probability of low reward actions. The discount factor γ balances the importance of short-term and long-term rewards. This method can handle continuous action spaces and nonlinear policies, making it particularly suitable for intervention policy generation, which requires considering long-term effects. Compared with traditional rule-based methods, it has better adaptability and optimization ability.

[0076] In step S07, when evaluating the intervention policy using counterfactual reasoning, the calculation process of the double robust estimator is specifically represented as follows:

[0077]

[0078] where τ DR is the double robust estimate of the average treatment effect; n is the sample size; m1(X i , Z i ) and m0(X i , Z i ) are the outcome models for the treatment group and the control group, respectively, predicting the outcome when receiving intervention X i = 1 or X i = 0 under the condition of covariate Z i ; Y i is the observed actual outcome; I(X i = j) is the indicator function, taking value 1 when X i = j, otherwise 0; e(Z i ) is the propensity score, representing the probability of receiving intervention under the condition of covariate Z i .

[0079] The doubly robust estimator combines an outcome regression model with a propensity score model. As long as one of the two models is correctly specified, the estimates are consistent, hence the name "doubly robust." This approach, based on a potential outcome framework, estimates counterfactual causal effects from observed data. This estimator reduces the impact of selection bias through weighting adjustments, improving the reliability of intervention effect estimates and offering greater robustness and accuracy than single-model approaches.

[0080] The multi-dimensional evaluation calculation of the value orientation index is specifically expressed as follows:

[0081] Index multi =[s1, s2, ..., s K ];

[0082]

[0083] In the formula, Index multi is the multidimensional value orientation vector; s k is the similarity between the content and the k-th value dimension; v content is the comprehensive feature vector of the content; is the standard vector of the kth value dimension; K is the total number of dimensions of value evaluation.

[0084] This multidimensional assessment method expands a single indicator into a vector representation, enabling a more comprehensive portrayal of content across different value dimensions. Compared to single-indicator assessment, multidimensional assessment provides a richer portrait of values, supporting more refined intervention decisions and is particularly well-suited for the assessment of ideological and political education in complex scenarios.

[0085] The reward function calculation of the educational intervention strategy is specifically expressed as follows:

[0086] R(s,a)=α·sim(s′,s target )-β·C(a)+λ·Δsim(s, s′, s target );

[0087] Δsim(s, s′, s target ) = sim(s′, s target )-sim(s,s target );

[0088] Where R(s, a) is the reward function value; s is the current state, which represents the learner's current value state vector; a is the selected intervention action; s′ is the new state after the intervention; s targetis the target value state; sim(·,·) is the similarity function between states, using cosine similarity; C(a) is the resource consumption cost of action a; Δsim is the degree of state improvement, which measures the improvement in similarity with the target state before and after intervention; α, β, and λ are weight coefficients, which control the importance of target achievement, resource efficiency, and improvement, respectively, with the value ranges of α∈[0.4, 0.6], β∈[0.1, 0.3], and λ∈[0.2, 0.5], and satisfy α+β+λ=1.

[0089] This reward function comprehensively considers goal achievement, resource efficiency, and improvement magnitude, achieving a multi-objective balance through a weighted combination. The Δsim term is introduced to encourage state improvement toward the goal, providing positive feedback even when the goal has not yet been fully achieved. This reward function design overcomes the limitations of traditional binary reward mechanisms, providing a more fine-grained feedback signal and facilitating progressive optimization of the policy.

[0090] The calculation of the multi-objective evaluation function of the intervention plan is specifically expressed as follows:

[0091]

[0092] Where E(A) is the comprehensive evaluation value of intervention plan A; w i is the weight coefficient of the i-th target, satisfying The value range is between 0.1 and 0.4; i (A) is the i-th normalized objective function value; f i (A) is the original objective function value of i; f i min and f i max The objective function f i The minimum and maximum values ​​of .

[0093] This multi-objective evaluation function transforms multidimensional objectives into a single scalar through weighted combination, facilitating scenario comparison and decision-making. Normalization ensures comparability across objectives of varying dimensions, preventing excessively large values ​​in a single dimension from dominating the evaluation results. This approach balances intervention effectiveness, efficiency, adaptability, and sustainability, supporting comprehensive multi-dimensional decision-making and better meeting the complex demands of educational interventions.

[0094] The calculation of the dynamic adjustment function of intervention intensity is specifically expressed as follows:

[0095] I′ j =I j ×α(R j );

[0096]

[0097] Where I′j is the adjusted intervention intensity; I j is the initial set intervention intensity, ranging from [0, 1]; a(R j ) is the adjustment coefficient based on real-time feedback R j , ranging from [0.7, 1.3]; k is the steepness parameter of the adjustment curve, taking the value of 5; R0 is the feedback threshold, taking the value of 0.5; R j is the real-time feedback score, ranging from [0, 1], with higher values indicating better acceptance.

[0098] This adjustment function adopts the Sigmoid form, achieving smooth nonlinear adjustment. When the feedback score is higher than the threshold, the intervention intensity is increased, and when it is lower than the threshold, the intervention intensity is decreased, embodying the principle of adaptive control. Parameter k controls the sensitivity of the adjustment response, and R0 sets the dividing point of the adjustment direction. This dynamic adjustment mechanism enhances the adaptability of the intervention process, optimizes the intervention strategy in real time according to the learner's acceptance, and avoids the negative effects that may be caused by fixed intensity intervention.

[0099] The calculation of the counterfactual prediction accuracy evaluation is specifically represented as follows:

[0100]

[0101] In the formula, CFE is the counterfactual verification error; n is the sample size; X i is the observed intervention; x is the intervention scheme to be evaluated; Y i (x) is the result that individual i should have observed if he or she accepted intervention x in a counterfactual situation; is the counterfactual result predicted by the model.

[0102] The counterfactual verification error is based on the mean square error principle to calculate the deviation between the model prediction and the ideal counterfactual result. Since it is impossible to observe the results of the same individual under both intervention and non-intervention in reality, this index usually needs to be estimated through special experimental design or statistical methods. The value of this index is required to be less than 0.15 to ensure the reliability of counterfactual prediction, providing a quality assurance mechanism for intervention decision-making.

[0103] The dynamic adjustment calculation of the sparsity factor in the sparse attention mechanism is specifically represented as follows:

[0104]

[0105] In the formula, S factor is the sparsity factor parameter, ranging from [0.3, 0.7]; D semantic is the semantic density of the value knowledge graph, calculated as the average connectivity between related nodes; C complexity is the content complexity, evaluated based on syntactic analysis and semantic diversity.i is the contribution of the i-th cluster center; w i is the weight of the i-th cluster center; N is the total number of cluster centers; α, β, and γ are weight coefficients, satisfying α+β+γ=1, and the value ranges are α∈[0.2,0.4], β∈[0.3,0.5], and γ∈[0.2,0.4].

[0106] This sparsity factor adjustment formula comprehensively considers three factors: semantic density, content complexity, and cluster center contribution. Semantic density reflects the structural characteristics of the knowledge graph, content complexity indicates the difficulty of the content being analyzed, and cluster center contribution considers the core semantic distribution of the content. The dynamically adjusted sparsity factor enables the attention mechanism to adaptively allocate computing resources based on content characteristics, improving computational efficiency while maintaining expressiveness.

[0107] The calculation of cluster center contribution is specifically expressed as follows:

[0108]

[0109] Where G i is the contribution of the i-th cluster center; C i is the semantic vector of the i-th cluster center; V std is the standard vector of values; sim() is the cosine similarity function; Coverage(C i ) is the coverage of the i-th cluster; |T i | is the number of text segments belonging to the i-th cluster; |T| is the total number of text segments; N is the total number of clusters.

[0110] Cluster center contribution is calculated based on two key factors: semantic relevance to the standard values ​​and coverage. Semantic relevance reflects the consistency of the cluster center with the target values, while coverage indicates the importance of the cluster within the overall content. Normalization ensures that the sum of the contributions of all cluster centers is 1. This calculation method enables the system to identify the semantic core most relevant to the values ​​in the content, prioritize attention resources, and improve the relevance and accuracy of value analysis.

[0111] Specifically, the principle of the present invention is: the present invention has built a complete set of technical solutions to the technical problems of implicit value identification and effective intervention. Its core principles can be explained from the following aspects:

[0112] First, the multimodal values ​​knowledge graph serves as a foundational semantic network, systematically representing the mainstream social value system, forming a network structure of entities, attributes, and relationships. This knowledge representation transcends the limitations of traditional keyword lists and can capture the complex connections and semantic hierarchies between value concepts, providing a rich foundation of prior knowledge and reasoning for subsequent analysis.

[0113] Secondly, the multimodal value understanding model utilizes an encoder-decoder architecture, integrating information from three modalities: text, vision, and audio. The hierarchical bidirectional transformer within the model architecture captures text features from the character to the paragraph level, the visual transformer extracts spatial and temporal features, and the audio processing network captures sentiment and intonation. This multimodal fusion design enables the system to comprehensively perceive diverse forms of value expression, while the sparse multi-head attention mechanism dynamically integrates the weights of information from different modalities, effectively improving the ability to recognize implicit expressions.

[0114] In terms of identifying value expressions, this invention uses a recurrent neural network to capture contextual dependencies and incorporates an attention mechanism to analyze the connections between key elements in the content, effectively identifying implicit values. The introduction of contrastive learning enables the model to distinguish subtle differences in the expression of different values, enhancing its ability to identify rhetorical devices. The calculation of a value orientation index transforms qualitative analysis into a quantitative assessment, providing an objective basis for intervention decisions.

[0115] Regarding intervention strategy generation and optimization, this paper models the intervention process as a Markov decision process based on reinforcement learning techniques, and uses policy gradient methods to learn the optimal intervention strategy. The application of counterfactual reasoning enables the system to evaluate the causal effects of different intervention strategies. A structural causal model and a dual-robust estimator enable scientific evaluation of intervention effectiveness. A multi-objective particle swarm algorithm is used to optimize intervention plans, comprehensively considering multidimensional objectives such as intervention effectiveness, efficiency, adaptability, and durability.

[0116] In summary, this paper organically combines three technical approaches: multimodal fusion analysis, deep learning identification, and reinforcement learning intervention, to construct a closed-loop system from value identification to intervention strategy generation and then to effect evaluation. This system design conforms to the basic principles of cognitive science and educational psychology, enabling accurate identification of implicit values ​​and personalized intervention, thus resolving the long-standing technical problem of difficult-to-quantify methods for accurate identification and effective intervention in ideological and political education.

[0117] A specific embodiment 1 of the present invention is provided below. The specific implementation of each step in this embodiment 1 is described in detail as follows.

[0118] The specific implementation method of step S01 is to construct a multimodal value knowledge graph and integrate the mainstream social value system into it as the basic semantic network. The specific implementation process first collects multimodal data such as text, images, and videos related to values. Natural language processing technology is used to extract value entities and their relationships, named entity recognition algorithm is used to identify value concepts in text, and relationship extraction algorithm is used to identify semantic connections between concepts. Convolutional neural networks are used to identify visual elements related to values ​​from images and videos, and the recognition accuracy threshold is set to 0.85. The ontology modeling method is used to define the concept hierarchy and semantic association of values, and the resource description framework is used to construct a knowledge representation in the form of triples. The value entities are vectorized, and transfer learning is used to transfer the pre-trained language model and visual model knowledge to the value field. The transfer learning rate is set to 5×10 -5 . Use graph embedding algorithms such as TransE or ComplEx to map entity e and relation r into a d-dimensional vector space, where the vector dimension d is set to 100, and the entity relationship is represented as a triple (h, r, t). The loss function is L = ∑ (h,r,t)∈S ∑ (h′,r,t′)∈S′ [γ+d(h+r,t)-d(h′+r,t′)] + , where γ is the boundary parameter set to 1.0, d is the Euclidean distance function, S is the set of correct triples, and S′ is the set of negatively sampled triples. The concept of the mainstream social value system is used as the core node and path of the knowledge graph. Knowledge fusion technology is used to integrate core values ​​with the existing knowledge graph to establish a weight system for value orientation. A graph database is used to store the constructed knowledge graph to achieve fast retrieval and reasoning, and a regular update mechanism is used to maintain the timeliness and accuracy of the knowledge graph. The association strength threshold between knowledge graph nodes is set to 0.65, and associations below this threshold will be filtered out. The purpose of this step is to provide a knowledge basis for subsequent analysis, so that the identification and judgment of values ​​have a unified standard and reference system.

[0119] The specific implementation method of step S02 is to use a pre-trained multimodal value understanding model to extract explicit value features in text and video, and use a recursive neural network to capture contextual dependencies. The specific implementation process first inputs the text and video content to be analyzed into the multimodal value understanding model, and uses a pre-trained encoder to extract text features and video features. Text features are extracted through a multi-layer bidirectional transformer network, taking into account the contextual relationship between words; video features are extracted through a visual transformer and a three-dimensional convolutional neural network, taking into account spatial and temporal information. For the extracted feature vectors, an attention pooling mechanism is used to map them to the value feature space to form an explicit value representation vector. Then, a recursive neural network such as a long short-term memory network or a gated recurrent unit is used to model the context sequence to capture long-distance dependencies and contextual information. The recursive network calculation process is as follows:

[0120] h t =tanh(W x x t +W h h t-1 +b h );

[0121] Where h t is the hidden state vector at time t, with a dimension of 512; x t is the input feature vector at time t; h t-1 is the hidden state vector at time t-1; W x is the weight matrix input to the hidden layer, with a size of d x ×d h , where d x is the input feature dimension, d h is the hidden state dimension 512; W h is the weight matrix from hidden layer to hidden layer, with size d h ×d h ; b h is the hidden layer bias vector, dimension d h ; tanh is the hyperbolic tangent activation function. The hidden layer state dimension in the recursive network is set to 512, and the time step is dynamically adjusted according to the content length, supporting a maximum of 1024 time steps. Finally, the output of the recursive neural network is fused with the explicit value representation vector to generate a value feature representation that takes the context into account. The explicit feature extraction confidence threshold of this step is set to 0.75, and features below this threshold will be marked as uncertain. The purpose of this step is to accurately identify the value elements clearly expressed in the content and understand their true meaning and tendency in combination with the context.

[0122] The specific implementation of step S03 is to apply the attention mechanism to analyze the connection between key words and visual elements in the content to identify implicit value expressions. The specific implementation process first segments the input multimodal content and establishes a temporal alignment relationship between each text paragraph and the corresponding visual element. The multi-head attention mechanism is used to calculate the attention weight matrix between the key words and visual elements in the text. The calculation process is as follows:

[0123]

[0124] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O ;

[0125] head i =Attention(QW i Q , KW i K , VW i V );

[0126] Where Q is the query matrix, size is n q ×d model ; K is the key matrix, size n k ×d model ; V is the value matrix, size is n k ×d model ;d k is the dimension of the key vector; W i Q 、W i K 、W i V The projection matrices of query, key, and value of the i-th attention head are d model ×d q d model ×d k d model ×d v , where d q =d k =d v =d model / h=64, h is the number of attention heads, value is 8; W O Output projection matrix, size hd v ×d model; softmax is a normalized exponential function; Concat is a concatenation operation. Highly correlated regions in text and visual content are identified through attention distribution. These regions often contain implicit expressions of values. The implicit semantics of words are then analyzed using the matching degree between keywords and the value knowledge graph, and a context-sensitive word embedding model is used to capture the semantic changes of the same word in different contexts. For visual elements, image sentiment analysis and symbol recognition algorithms are used to detect visual clues that may contain value implications. Finally, the implicit features of text and visual modalities are cross-fused, and an adaptive weighting method is used to integrate multimodal information to generate an implicit value representation vector. The confidence threshold for implicit value judgment is set to 0.6, and it is required to be supported by at least two or more modal evidences. The purpose of this step is to discover value tendencies that are not directly stated but implied in the content, and to improve the ability to understand rhetorical techniques such as euphemisms, metaphors, and irony.

[0127] The specific implementation of step S04 is to construct positive and negative sample pairs through contrastive learning, learn the differences in value expression, and enhance the model's ability to identify rhetorical devices. The specific implementation process first selects content pairs expressing similar values ​​but using different rhetorical devices from the training data as positive sample pairs, and selects content pairs expressing opposite values ​​but using similar expressions as negative sample pairs. Each sample pair contains the original content feature vector and the corresponding value label. Training is performed using a contrastive loss function, which is calculated as follows:

[0128]

[0129] Where, L contrastive is the contrast loss value; f(x) is the feature representation function of content x; x i is the anchor sample; is a positive sample that expresses the same values ​​as the anchor sample; x i is all samples in the batch; s() is the similarity function, using cosine similarity τ is the temperature parameter, with a value of 0.07; N is the batch size, with a value of 128. The model learns to map different expressions of the same value to similar feature spaces, while distinguishing the expressions of different value concepts. A difficult sample mining strategy is introduced during the training process, focusing on easily confused expressions of values, such as the distinction between euphemistic expressions and ironic expressions. Difficult sample mining uses an online difficult example mining algorithm, and selects the top 30% samples with the highest loss values ​​in each batch for gradient updates. In addition, data enhancement technology is implemented to generate more diverse expression variants, including synonym replacement, syntactic reconstruction, and context transformation. Model training uses a small batch stochastic gradient descent algorithm with a batch size of 128, and the learning rate is initially set to 10 -4 , use the cosine annealing strategy to adjust, the learning rate change formula is where η t is the learning rate of step t, η min The minimum learning rate is 10 -6 , η max The maximum learning rate is 10 -4 , where T is the total number of training steps. The similarity threshold for contrastive learning is set to 0.85; pairs of samples above this threshold are considered to have the same value orientation. This step enhances the model's ability to distinguish subtle differences in values, especially for content expressed using complex rhetoric.

[0130] The specific implementation method of step S05 is to calculate the cosine similarity between the content and the core value vector based on the value feature vector to generate a value orientation index. The specific implementation process first extracts the standard vector representation of mainstream social values ​​from the value knowledge graph. These vectors have been generated during the knowledge graph construction process. For the content to be analyzed, the explicit and implicit value feature vectors extracted in the previous steps are merged and integrated using the attention weighting method to generate a comprehensive value feature vector of the content. The process of calculating the cosine similarity between the content and the core value vector and the multidimensional value orientation index is as follows:

[0131]

[0132] Index multi =[s1, s2, ..., s K ];

[0133]

[0134] Where, v content is the comprehensive value feature vector of the content, with a dimension of 128; is the standard vector representation of the i-th core value, and the dimension is also 128; similarity() is the cosine similarity function; w i is the weight coefficient of the i-th core value, satisfying The weight value is pre-set according to the importance of the value, ranging from 0.05 to 0.3; M is the total number of core values; Index value The final value orientation index ranges from 0 to 1; Index multi is the multidimensional value orientation vector; s kis the similarity between the content and the kth value dimension; K is the total number of value dimensions assessed. The value orientation index ranges from 0 to 1, with higher values ​​indicating greater alignment with core values. The threshold for determining content as positively oriented is set at 0.75, with values ​​between 0.4 and 0.75 considered neutral and below 0.4 as negative. This step quantifies the degree of consistency between content and core values, providing a basis for decision-making on subsequent intervention strategies.

[0135] The specific implementation of step S06 is to use reinforcement learning technology to train the content intervention agent and generate an educational intervention strategy based on the value orientation index. The specific implementation process first constructs a Markov decision process framework. The state space is defined as the learner's current value state vector, represented by the value orientation index calculated in the previous step. The action space is defined as the set of available intervention strategies, including different combinations of intervention types, methods, and intensities. The policy gradient and reward function calculation process of the educational intervention strategy are as follows:

[0136]

[0137] π * (s)=argmax a ∑ s′ P(s′|s,a)[R(s,a,s′)+yV(s′)];

[0138] R(s,a)=α·sim(s′,s target )-β·C(a)+λ·Δsim(s, s′, s target );

[0139] Δsim(s, s′, s target ) = sim(s′, s target )-sim(s,s target );

[0140] Wherein. is the policy gradient; θ is the policy network parameter; π θ (a|s) is a policy function with parameter θ, which represents the probability of selecting action a in state s; Qπ(s, a) is a state-action value function, which represents the expected cumulative reward of taking action a in state s according to policy π; γ is a discount factor with a value of 0.9; r t is the reward obtained at time t; P(s′|s, a) is the state transition probability, which represents the probability of transitioning to state s′ after taking action a in state s; R(s, a, s′) is the immediate reward function; V(s) is the state value function; π*(s) is the optimal strategy; R(s, a) is the reward function value; s is the current state, which represents the learner’s current value state vector; a is the selected intervention action; s′ is the new state after the intervention; starget is the target value state; sim(·,.) is the similarity function between states, using cosine similarity; C(a) is the resource consumption cost of action a; Δsim is the degree of state improvement, measuring the improvement in similarity with the target state before and after the intervention; α, β, and λ are weight coefficients, respectively controlling the importance of goal achievement, resource efficiency, and improvement magnitude, with values ​​ranging from α∈[0.4, 0.6], β∈[0.1, 0.3], and λ∈[0.2, 0.5], and satisfying α+β+λ=1. The intervention agent is trained using a deep Q-network or policy gradient method, with the current state as input and the optimal intervention policy as output. The Q-network uses a three-layer fully connected neural network with 256 and 128 hidden layer nodes, respectively, and the activation function is ReLU. To balance exploration and exploitation, a ∈-greedy strategy is adopted, with an initial ∈ value of 0.3 and gradually decaying to 0.05. The decay formula is ∈=∈en d +(∈star t -∈ end )×e -λt , where λ is the decay rate 0.001 and t is the number of training steps. During the training process, the experience replay buffer is used to store the transfer samples, and the buffer size is 10 5 , randomly sampling batches of 64 samples each time for training. The target network update frequency was set to every 100 steps, and the discount factor γ was set to 0.9. Intervention strategies were evaluated using an A / B testing method to compare the effectiveness of different strategies. The strategy selection confidence threshold was set to 0.7; below this threshold, a conservative intervention strategy was adopted. This step automatically generates personalized educational intervention strategies based on the learner's current state, achieving intelligent and precise ideological and political education.

[0141] The specific implementation method of step S07 is to use the counterfactual reasoning method to evaluate the impact of different intervention strategies on the effect of ideological and political education, and optimize the intervention plan. The specific implementation process first constructs a structural causal model, and defines variables including learner characteristics, intervention strategies, and educational effects. The causal effect estimation model is trained based on historical intervention data, and a dual robust estimator is used in combination with the propensity score and outcome regression model for training. The calculation process of the dual robust estimator, intervention plan evaluation function, intervention intensity dynamic adjustment function, and counterfactual prediction accuracy evaluation is as follows:

[0142]

[0143]

[0144] I′ j =I j ×α(R j );

[0145]

[0146] Where, T DR is the double robust estimate of the average treatment effect; n is the number of samples; m1(X i , Z i ) and m0(X i , Z i ) are the result models of the treatment group and the control group, respectively, and the covariate Z i Conditions receiving intervention X i =1 or X i =0; Y i is the actual result observed; I(X i =j) is the indicator function, when X i = j, the value is 1, otherwise it is 0; e(Z i ) is the propensity score, which indicates that the covariate Z i The probability of accepting intervention under certain conditions; E(A) is the comprehensive evaluation value of intervention plan A; w i is the weight coefficient of the i-th target, satisfying The value range is between 0.1 and 0.4; i (A) is the i-th normalized objective function value; f i (A) is the original objective function value of i; f i min and f i max The objective function f i The minimum and maximum values ​​of I′ j is the adjusted intervention intensity; I j is the initial intervention intensity, ranging from [0, 1]; α(R j ) is based on real-time feedback R j The adjustment coefficient is in the range of [0.7, 1.3]; k is the steepness parameter of the adjustment curve, which is 5; R0 is the feedback threshold, which is 0.5; R j is the real-time feedback score, with a value range of [0, 1], where a higher value indicates better acceptance; CFE is the counterfactual verification error; X i is the observed intervention; x is the intervention to be evaluated; Y i (x) is the outcome that would have been observed if individual i had received intervention x in the counterfactual scenario; is the counterfactual result predicted by the model. The propensity score model uses logistic regression or random forest algorithm, and the prediction formula is Where β0 is the intercept term and β is the regression coefficient vector; the resulting regression model uses a multi-layer perceptron to predict the educational effect under a given intervention. For new situations, the trained model is used to generate counterfactual predictions, that is, to estimate "what effect will occur if strategy B is adopted for learner A". By comparing the predicted effects of different intervention strategies, the optimal intervention path is selected. In order to deal with multi-objective optimization problems, an improved Pareto optimization method is adopted, taking into account the four goals of intervention effectiveness, efficiency, adaptability and durability. The multi-objective particle swarm algorithm is used to find the Pareto optimal solution set. Each particle represents a candidate intervention plan, and its position and speed are defined by the intervention strategy parameters. The position update formula is x i (t+1)=x i (t)+v i (t+1), the speed update formula is v i (t+1)=w×v i (t)+c1r1×(p i -x i (t))+c2r2×(p g -x i (t)), where w is the inertia weight 0.7, c1 and c2 are learning factors both 2.0, r1 and r2 are random numbers in the range [0, 1|, p i is the individual optimal position of particle i, p g The global optimal position is determined by the weight of each objective and the decision maker's preferences. The order of intervention implementation follows a principle of increasing intensity, dynamically adjusted based on learner acceptance. The counterfactual prediction accuracy assessment threshold is set at 0.15; predictions above this threshold are considered unreliable. This step aims to identify the most effective combination of intervention strategies through simulation analysis, ensuring that educational interventions achieve their intended outcomes while optimizing resource utilization.

[0147] The dynamic adjustment calculation of the sparse factor of the sparse attention mechanism in the multimodal value knowledge graph is as follows:

[0148]

[0149] Where S factor is the sparsity factor parameter, and its value range is [0.3, 0.7]; D semantic is the semantic density of the value knowledge graph, calculated as the average connectivity between related nodes; C complexity G is the content complexity, which is evaluated based on syntactic analysis and semantic diversity; i is the contribution of the i-th cluster center; w iis the weight of the i-th cluster center; N is the total number of cluster centers; α, β, and γ are weight coefficients, satisfying α+β+γ=1, and the value ranges are α∈[0.2,0.4], β∈[0.3,0.5|, and γ∈[0.2,0.4] respectively; C i is the semantic vector of the i-th cluster center; V std is the standard vector of values; sim() is the cosine similarity function; Coverage(C i ) is the coverage of the i-th cluster; |T i | is the number of text segments belonging to the i-th cluster; |T| is the total number of text segments.

[0150] Furthermore, the specific implementation of the counterfactual reasoning module is a computational unit based on the structural causal model (SCM), which is used to evaluate the impact of different intervention strategies on ideological and political education and optimize intervention plans. This module first establishes a causal graph structure containing learner characteristic variables, intervention strategy variables, and value outcome variables. It then calculates the causal effect of each strategy through intervention scenarios and selects the optimal intervention path.

[0151] The specific implementation process begins by constructing a complete structural causal model (SCM = (U, V, F), where U is the set of exogenous variables (learner background characteristics, such as age, major, and basic cognitive level), V is the set of endogenous variables (including intervention strategy variables X and value outcome variables Y), and F is the set of functions describing the causal relationships between variables. The causal relationships between variables are represented using Bayesian networks or structural equation models, and the structural parameters are determined using historical intervention data or expert knowledge.

[0152] Then, we train the causal effect estimation model based on the historical intervention data, using a dual robust estimator combined with the propensity score method and the outcome regression model. The propensity score model uses the logistic regression algorithm, and the prediction formula is Where β0 is the intercept term and β is the regression coefficient vector, which was optimized and solved by the maximum likelihood estimation method. Results The regression model used a three-layer multilayer perceptron with 64 and 32 hidden layer nodes, the activation function was ReLU, and the output layer used a linear activation function to predict the intervention effect.

[0153] Then the dual robust estimator is implemented to calculate the counterfactual effect. The specific formula is:

[0154]

[0155] Where m1 and m0 are the result models of the treatment group and the control group respectively, I(X i =j) is the indicator function.

[0156] To evaluate the effectiveness of different intervention strategy combinations, the module constructs a counterfactual prediction function that predicts the counterfactual outcome Y(x) for a new learner feature set Z and a candidate intervention strategy set X. Prediction accuracy is measured by the counterfactual verification error (CFE), which is required to be less than 0.15 to ensure reliability.

[0157] The module also integrates a multi-objective optimization component, which considers the four objectives of intervention effectiveness, efficiency, adaptability and durability, and is formalized as finding an intervention strategy sequence A * ={a1, a2, ..., a n} Make the objective function vector F(A) = [f1(A), f2(A), f3(A), f4(A)] reach the optimal value. The optimization process adopts the multi-objective particle swarm algorithm, and the update formula of the particle position is x i (t+1)=x i (t)+v i (t+1), the speed update formula is v i (t+1)=w×v i (t)+c1r1×(p i -x i (t))+c2r2×(p g -x i (t)).

[0158] Finally, the counterfactual reasoning module generates an intervention effectiveness report, including the expected effects of each intervention strategy, uncertainty estimates, and sensitivity analysis, providing a scientific basis for decision-making on ideological and political education interventions. This module continuously updates model parameters and structure through an iterative improvement mechanism to ensure the accuracy and applicability of predictions.

[0159] The implementation of the counterfactual reasoning module breaks through the limitations of traditional intervention evaluation methods. Through a rigorous causal inference framework and data-driven effect prediction, it achieves accurate evaluation and scientific optimization of the ideological and political education intervention effect, providing strong technical support for value-oriented ideological and political education.

[0160] In order to better understand and implement the present invention, Example 2 of a specific application scenario of the present invention is provided below: A research team conducted an applied study on a value-oriented content analysis and intervention method for the ideological and political education of college students. The team first constructed a multimodal value knowledge graph that includes the mainstream social value system. The knowledge graph contains 5,874 concept entity nodes and 8,521 relationship edges, covering multimodal information such as text, images, and videos. On this basis, the research team selected 320 students from a university as experimental subjects, and collected multimodal content generated by these students in classroom discussions, social media, and learning platforms, including text posts, video comments, and picture and text sharing, etc., and collected a total of about 12,500 content samples.

[0161] The research team first used the constructed value knowledge graph to analyze the value orientation of these content samples. According to the specific implementation method of S01 to S05 steps, the research team extracted the value characteristics of the content generated by students and calculated the inclination index. Through analysis, the research team divided the students into four value orientation groups, as shown in Table 1:

[0162] Table 1 Student Value Orientation Distribution Table

[0163] Value orientation type Number of students Proportion Average Value Orientation Index Highly compatible 87 27.2% 0.851 Basic fit 156 48.8% 0.683 Neutral deviation 58 18.1% 0.512 Obvious deviation 19 5.9% 0.372

[0164] Based on the above analysis results, the research team designed differentiated intervention strategies for different value orientation groups. According to the specific implementation method of S06 steps, the research team trained the content intervention agent system and used reinforcement learning technology to generate targeted intervention strategies. Table 2 shows the intervention strategy matrix for different groups:

[0165] Table 2 Differentiated Intervention Strategy Matrix

[0166] Tendency type Type of intervention Intervention methods Intervention intensity Implementation frequency Highly compatible Cognitive consolidation Case Study 0.35 Once a week Basic fit Cognitive Guidance Interactive Discussion 0.55 Twice a week Neutral deviation Emotional guidance Case Comparison 0.75 3 times a week Obvious deviation Cognitive Correction Directly express 0.90 4 times a week

[0167] The research team then conducted a 16-week intervention experiment. During the intervention process, the research team used counterfactual reasoning to evaluate the effectiveness of different intervention strategies according to the specific implementation method of S07 steps, and dynamically adjusted the intervention scheme. During the implementation process, according to the real-time feedback of students, the intervention intensity was adjusted. The intervention intensity dynamic adjustment function I' j j = I j × α (R The research team optimized part of the intervention intensity, as shown in Table 3:

[0168] Table 3 Intervention Intensity Dynamic Adjustment Results

[0169] student groups Original intervention intensity Real-time feedback scoring Adjustment factor Optimized intervention intensity Highly compatible 0.35 0.71 1.15 0.40 Basic fit 0.55 0.63 1.06 0.58 Neutral deviation 0.75 0.48 0.92 0.69 Obvious deviation 0.90 0.39 0.81 0.73

[0170] During the experiment, the research team applied the multi-head attention mechanism to analyze the connection between keywords and visual elements in the content of students in the "neutral deviation" and "obvious deviation" groups to identify implicit value expression. By calculating the attention weight matrix of text and visual content The research team successfully identified the implicit value expression characteristics in the content of these students. Table 4 shows the identification results of different types of implicit value expression:

[0171] Table 4 Implicit Value Expression Identification Results

[0172] ​

[0173]

[0174] For the identified implicit value expression, the research team adopts a contrast learning method to construct positive and negative sample pairs, learn the difference in value expression, and enhance the model's ability to recognize rhetorical devices. The contrast loss function In the optimization process of the contrast loss function, with the increase of training rounds, the model's recognition accuracy of value expression under different rhetorical devices significantly improves, as shown in Table 5:

[0175] Table 5 Contrast learning training effect

[0176] Training rounds Training loss value Validation set loss value Rhetorical device recognition accuracy Accuracy of value orientation judgment 1 0.862 0.879 68.5% 72.3% 5 0.536 0.589 76.2% 79.8% 10 0.385 0.412 82.7% 84.6% 15 0.287 0.326 87.1% 88.2% 20 0.231 0.283 89.5% 90.7%

[0177] After 16 weeks of intervention experiments, the research team re-evaluated the students' value orientation index. By calculating the cosine similarity between the content feature vector and the core value vector, the research team obtained the changes in students' value orientation before and after the intervention, as shown in Table 6:

[0178] Table 6 Changes in value orientation before and after intervention

[0179] Original tendency type Average index before intervention Average index after intervention Improvement Improvement ratio Highly compatible 0.851 0.892 0.041 4.8% Basic fit 0.683 0.752 0.069 10.1% Neutral deviation 0.512 0.647 0.135 26.4% Obvious deviation 0.372 0.536 0.164 44.1%

[0180] The research team also conducted a comprehensive evaluation of the intervention program based on the multi-objective evaluation function , considering the effectiveness, efficiency, adaptability, and durability of the intervention. After evaluation, the comprehensive effect of different intervention strategies is shown in Table 7:

[0181] Table 7 Comprehensive effect evaluation of intervention strategies

[0182] Intervention strategies Effectiveness efficiency Adaptability Persistence Overall score Cognitive consolidation + case deepening 0.83 0.91 0.78 0.85 0.843 Cognitive guidance + interactive discussion 0.87 0.76 0.89 0.81 0.832 Emotional guidance + case comparison 0.92 0.68 0.83 0.79 0.815 Cognitive correction + direct instruction 0.95 0.63 0.74 0.77 0.786

[0183] Traditional ideological and political education intervention methods mainly rely on teachers' experience and intuitive judgment, lacking precise analysis and targeted intervention of students' value orientation. Conventional methods usually use unified education content and teaching methods, which cannot meet the differentiated needs of students with different value orientations. Traditional methods have great difficulty in identifying students' implicit value expression, especially for value orientation expressed by euphemism, metaphor, irony, and other rhetorical devices, with an accuracy rate usually not exceeding 50%. At the same time, traditional intervention strategies are often static and fixed, lacking real-time adjustment mechanism, and intervention effect evaluation is mostly based on subjective judgment, lacking objective quantitative indicators.

[0184] In contrast, the value-oriented content analysis and intervention method proposed in this paper offers several technological advancements. First, by constructing a multimodal value knowledge graph, it provides a unified standard and reference system for value analysis, making value identification and judgment more objective and systematic. Second, by employing deep learning and multimodal fusion techniques, it accurately identifies both explicit and implicit value expressions, increasing the accuracy to over 80%. The recognition capability is particularly enhanced for expressions with complex rhetoric. Third, the intervention agent system based on reinforcement learning enables personalized and dynamic intervention strategy generation, adaptively adjusting intervention intensity and approach based on real-time student feedback. Finally, the intervention effectiveness is objectively evaluated using counterfactual reasoning, and a multi-objective optimization approach is employed to optimize the intervention plan, resulting in more precise and efficient intervention results. Experimental results demonstrate that this method is highly effective across groups with diverse value orientations, particularly for those with neutral and significant deviations, with increases in value alignment reaching 26.4% and 44.1%, respectively, significantly exceeding the 5%-15% improvement achieved by traditional intervention methods.

[0185] It should be noted that the variables involved in the present invention are explained in detail as shown in Tables 8, 9 and 10 below.

[0186] Table 8 Variable Explanation Table (Part 1)

[0187]

[0188] Table 9 Variable Explanation Table (Part 2)

[0189]

[0190]

[0191] Table 10 Variable Explanation Table (Part 3)

[0192]

[0193]

[0194] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.

Claims

1. A content analysis and intervention method for value orientation in ideological and political education, characterized by: include: Construct a multimodal value knowledge graph and integrate it into the mainstream social value system as the basic semantic network; use a multimodal value understanding model to extract explicit value features and use a recursive neural network to capture contextual dependencies; apply an attention mechanism to analyze the connection between key words and visual elements in the content to identify implicit value expressions; use a comparative learning method to construct positive and negative sample pairs to learn the difference in value expressions and enhance the model's ability to recognize rhetorical devices; calculate the cosine similarity between the content and the core value vector based on the value feature vector to generate a value tendency index; Reinforcement learning technology is used to train content intervention agents to generate educational intervention strategies based on the value orientation index; counterfactual reasoning methods are used to evaluate the impact of different intervention strategies on the effectiveness of ideological and political education and optimize intervention plans.

2. The content analysis and intervention method for value orientation in ideological and political education according to claim 1 is characterized by: The multimodal value knowledge graph is a structured knowledge network constructed by the concepts, expressions and associations of values ​​in various media forms. It contains entity nodes, attributes and relationship edges, and is used to represent the mapping relationship and semantic connection between value expressions in different modalities.

3. The content analysis and intervention method for value orientation in ideological and political education according to claim 2 is characterized by: Implicit expression of values ​​is the transmission of values ​​and positions through indirect rhetoric such as euphemism, metaphor, irony or non-explicit expression. It is usually necessary to comprehensively judge the true tendency based on the context, cultural background and expression method.

4. The content analysis and intervention method for value orientation in ideological and political education according to claim 3 is characterized by: The value orientation index is a quantitative score obtained by calculating the similarity between the content feature vector and the preset value vector. It is used to measure the degree of fit between content and values. The value range is 0 to 1. The higher the value, the higher the fit.

5. The content analysis and intervention method for value orientation in ideological and political education according to claim 4 is characterized by: The counterfactual reasoning module is a computing unit built based on the principle of causal inference. By simulating possible outcomes under different intervention strategies, it evaluates the effectiveness and applicability of various intervention plans and thus selects the optimal intervention path.

6. The content analysis and intervention method for value orientation in ideological and political education according to claim 5 is characterized by: The multimodal values ​​comprehension model is a deep neural network model that is trained with large-scale corpus and multimodal content and is used to understand and analyze explicit and implicit value expressions in various media. It uses an encoder-decoder architecture to integrate cross-modal information and has a culturally sensitive adaptation mechanism.

7. The content analysis and intervention method for value orientation in ideological and political education according to claim 6 is characterized by: The specific structure of the multimodal value understanding model is a deep neural network based on a multi-head self-attention mechanism and a stack of cross-modal interaction layers. It includes four core components: a text encoding module, a visual encoding module, an audio encoding module, and a cross-modal fusion module. The text encoding module uses a hierarchical bidirectional converter structure to process text information from characters to sentences to paragraphs. The visual encoding module uses a visual converter to extract spatial features and temporal relationships in images and video frames. The audio encoding module extracts emotional and intonation features from sound information through a one-dimensional convolutional network and a long short-term memory network. The cross-modal fusion module uses a sparse multi-head attention mechanism to realize the dynamic weight integration of different modal information.

8. The content analysis and intervention method for value orientation in ideological and political education according to claim 7 is characterized by: The sparse factor parameters in the sparse multi-head attention mechanism are dynamically adjusted according to the semantic density of the value knowledge graph, the complexity of the content, and the contribution of each segmented text clustering center. The number of attention heads is determined by the difficulty of identifying implicit value expressions and the complexity of rhetorical techniques. The final output layer of the model generates a value tendency feature vector through nonlinear activation function mapping.

9. The content analysis and intervention method for value orientation in ideological and political education according to claim 8 is characterized by: The contribution of a cluster center refers to the weight of the influence of each cluster center on the overall value expression when cluster analysis is performed after text information segmentation. It is quantified by calculating the semantic relevance of each cluster center to the value standard vector and its coverage in the text. Cluster centers with high contribution will obtain more attention resources allocation in the sparse attention mechanism.

10. The content analysis and intervention method for value orientation in ideological and political education according to claim 9 is characterized by: The steps for establishing a training dataset for a multimodal values ​​comprehension model include collecting educational text, image, and video resources with clear value labels from multi-source databases, manually annotating all collected resources to determine the value category and strength, dividing the data into two categories of explicit expression and implicit expression according to the content expression method and performing detailed annotation, constructing a sample set containing a comparison of positive and negative values, and stratifying the data according to cultural background and contextual factors to ensure that the training set covers multiple expression scenarios.

Citation Information

Cited By

  • Social medium anti-mental semantic recognition method and system, storage medium and electronic equipment

    CN121093970A

  • Policy field-oriented reasoning generation method, apparatus and device, and storage medium

    CN121279459A