Zero sample vertical field detection method based on cognitive mode graph enhancement
By constructing logical cognitive mode and cognitive-first-order logical rule graph CFGraph, combining large language model LLM and relational graph convolution network RGCN, the existing zero-sample position detection method in terms of labeling data dependence and generalization capabilities is solved, and a more efficient and accurate position detection effect is achieved.
Patent Information
- Application Number
- CN202510263189.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
AI Technical Summary
The existing zero-sample position detection method has shortcomings in labeling data dependence, domain adaptability, and generalization ability, which is difficult to effectively solve the accuracy and efficiency of text position detection in social media.
Using a zero-sample position detection method based on cognitive mode graph enhancement, a large language model LLM and relational graph convolution network RGCN are used to model and predict models by constructing logical cognitive mode and cognitive-first-order logic rule graph CFGraph.
It significantly reduces the dependence on labeled data, improves the generalization ability and semantic understanding ability of the model, and enhances the adaptability and accuracy of the model under zero-sample settings.
Smart Images

Figure CN120217024A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a zero-shot stance detection method enhanced based on a cognitive pattern graph, aiming to solve the problems of high dependence on labeled data and insufficient generalization ability in text stance detection on social media. Background Art
[0002] With the rapid development and popularization of Internet technology, social media platforms have been deeply integrated into people's daily lives and become an important channel for the public to exchange ideas, share emotions and express opinions. Social media platforms such as Weibo, WeChat, and Twitter gather a huge amount of content posted by users globally every day. These contents not only cover all aspects of daily life, but also involve multiple fields such as politics, economy, and culture, and become an important resource for studying public stances and public opinion trends.
[0003] Government agencies and enterprises are paying increasing attention to social media content. Government agencies can use the data on social media to monitor public opinion and grasp social trends, so as to formulate more people-oriented and targeted policies. Enterprises can, by analyzing user feedback on social media, gain insights into market demands, understand consumer preferences, and provide strong support for product development and marketing strategy formulation.
[0004] However, how to efficiently and accurately analyze the stances in social media texts has become a common challenge for the academic and industrial communities. Stance detection refers to automatically identifying the viewpoints, emotions or attitudes expressed in texts, which is of great significance for understanding public opinions, predicting market reactions and social feedbacks.
[0005] Early stance detection research mainly focused on two major tasks: single-target stance detection and cross-target stance detection. Single-target stance detection assumes that the target sets in the training set and the test set are the same, that is, the model is trained and evaluated based on data with the same target. However, in real scenarios, the targets during inference are often those that the trained stance detection model has never seen, because it is impossible to list all possible detection targets in advance during the model training stage. For this reason, cross-target stance detection partially alleviates the data annotation problem by adapting the classifier trained for a certain target to related new targets. However, cross-target stance detection relies on a strong assumption, that is, it is necessary to know the semantic relationships between targets in advance, which is difficult to meet in practical applications because it is impossible to list all potential targets and their associations in advance.
[0006] Therefore, zero-shot stance detection emerges as a promising research direction. The goal of zero-shot stance detection is to make inferences on completely unrelated targets without relying on the relevance between the training and test targets. Compared with single-target stance detection and cross-target stance detection, zero-shot stance detection is more applicable to the diverse and real-time scenarios in social media, such as stance analysis of emergencies or emerging topics.
[0007] For the zero-shot stance detection task, various methods have been proposed. Early methods mainly relied on traditional deep learning architectures, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs). With the emergence of pre-trained language models like BERT, fine-tuning-based models and prompt tuning methods have been widely used. However, these methods highly rely on a large amount of manually annotated data during training and have limitations in generalization ability.
[0008] Recently, large language models (LLMs) such as the GPT series have demonstrated excellent zero-shot performance in common sense reasoning tasks, bringing new possibilities to zero-shot stance detection. One method is called chain-of-thought prompting, which treats LLMs as effective knowledge experts and utilizes their reasoning ability for prediction through predefined prompts. However, LLMs usually lack domain-specific knowledge required for specific tasks, and the knowledge embedded in them may be outdated. In addition, stance detection methods based on LLMs generally perform worse than the latest non-LLM baseline models.
[0009] Another alternative is the LLM-enhanced fine-tuning method, which combines LLMs with traditional trainable stance detection models. The LLM-enhanced fine-tuning method consists of two stages: First, chain-of-thought prompting is used to obtain relevant background knowledge; then, the obtained knowledge is input into the trainable stance detection model together with the original text. Although this method effectively combines the reasoning ability of large models and the learning ability of labeled samples, enhancing the zero-shot learning ability of the model, its performance still depends on a large number of labeled instances and cannot effectively solve problems such as implicit stance confusion and stance label hallucination.
[0010] In summary, existing zero-shot stance detection methods have deficiencies in terms of labeled data dependence, domain adaptability, and generalization ability. In view of this, the present invention proposes a zero-shot stance detection method enhanced by a cognitive pattern graph, aiming to overcome the shortcomings of the prior art and improve the accuracy and efficiency of stance detection. Summary of the Invention
[0011] Aiming at the above-mentioned shortcomings of the prior art, the present invention provides a zero-shot stance detection method enhanced by a cognitive pattern graph, aiming to reduce the dependence on labeled data, improve the generalization ability of the model, and enhance the domain adaptability of LLMs.
[0012] The present invention achieves the above object through the following technical solutions:
[0013] A zero-shot stance detection method based on enhanced cognitive pattern graphs, comprising the following steps:
[0014] Construct a logical cognitive pattern:
[0015] Use a pre-trained large language model LLM to convert the input text into a logical expression; cluster the logical predicates in the logical expression and organize the logical predicates into different clusters; summarize the text in each cluster to generate a summary phrase for the cluster; construct a weighted multi-relationship graph as the cognitive pattern according to the logical relationships between the logical predicates and the summary phrases of the clusters.
[0016] Cognition-enhanced stance prediction:
[0017] In the prediction stage, convert the text to be predicted into a logical expression, and extract the subgraph related to the logical expression of the text to be predicted from the constructed cognitive pattern; combine the extracted subgraph with the logical expression of the text to be predicted to construct a cognitive-first-order logic rule graph CFGraph; use a relational graph convolutional network to model the constructed CFGraph for stance prediction; select the category with the highest probability as the prediction result.
[0018] According to a zero-shot stance detection method based on enhanced cognitive pattern graphs provided by the present invention, in the input stage, receive an input text containing at least one logical statement; use a large language model LLM to perform semantic analysis on the input text, extract the predicates and their logical relationships in the text, and generate a corresponding logical expression; convert the predicates in the logical expression into nodes in a first-order logic rule graph, and each node represents an independent predicate concept; establish edges between the nodes in the first-order logic rule graph according to the logical relationships between the predicates in the logical expression, and the edges represent the logical associations between the predicates, including at least logical relationships such as implication, conjunction, disjunction, and negation.
[0019] According to a zero-shot stance detection method based on enhanced cognitive pattern graphs provided by the present invention, based on the constructed first-order logic rule graph, extract the matching cognitive pattern subgraph from the weighted multi-relationship graph of the cognitive pattern, and fuse the cognitive pattern subgraph with the first-order logic rule graph to obtain a cognitive-first-order logic rule graph CFGraph.
[0020] According to a zero-shot stance detection method based on enhanced cognitive pattern graphs provided by the present invention, the following specific steps are further included to generate logical expressions using a pre-trained large language model LLM:
[0021] Construct an instruction prompt P1, which is used to guide the large language model to understand the task requirements, that is, convert the input text into the corresponding first-order logic expression;
[0022] Prepare a sample input set, where each sample includes a piece of social text and a corresponding target first-order logic expression, and the target first-order logic expression contains logical predicates and logical relationships;
[0023] Provide the constructed instruction prompt P1 and all sample inputs to the pre-trained large language model LLM for model training or fine-tuning, so that the model learns how to convert social text into first-order logic expressions;
[0024] After the model training or fine-tuning is completed, use this large language model LLM to receive new input text, and according to the learned mapping relationship, convert the new input text into the corresponding first-order logic expression, which also contains logical predicates and the logical relationships between them.
[0025] According to a zero-shot stance detection method based on enhanced cognitive pattern graph provided by the present invention, when performing model training or fine-tuning, encode the instruction prompt P1 into a format understandable by the model, and pair it with each sample in the sample input set to form training data pairs;
[0026] Input the training data pairs into the pre-trained large language model LLM and set the training parameters;
[0027] Through the supervised learning method, enable the model to gradually learn the tasks specified by the instruction prompt P1 during the training process, that is, map the semantic information in the social text to the logical predicates and logical relationships in the first-order logic expression;
[0028] During the training process, use a loss function to evaluate the difference between the logical expression generated by the model and the target logical expression, and update the parameters of the model through the backpropagation algorithm to minimize this difference;
[0029] When the performance of the model on the training set reaches a predetermined standard, stop training or perform fine-tuning to further optimize the performance of the model;
[0030] The trained or fine-tuned model is used to receive new social text inputs and automatically generate the corresponding first-order logic expressions according to the learned mapping relationship.
[0031] According to a zero-shot stance detection method based on enhanced cognitive pattern graph provided by the present invention, it also includes using the K-means clustering algorithm to organize the logical predicates in the first-order logic FOL expressions into different clusters, and the specific implementation steps are as follows:
[0032] Convert each logical predicate into a d-dimensional embedding vector e through the Sentence-BERT model to fully capture its semantic details and features;
[0033] Initialize the number of clusters K of the K-means clustering algorithm and randomly select K initial cluster center vectors;
[0034] For each d-dimensional embedding vector e corresponding to a logical predicate, calculate its distance from each cluster center vector and assign the vector to the cluster with the closest distance;
[0035] According to the attribution result, update the center vector of each cluster to the mean vector of all embedding vectors within the cluster;
[0036] Repeat the embedding vector step until the cluster center vectors no longer change significantly or reach the preset number of iterations to minimize the within-cluster sum of squared distances;
[0037] Finally, obtain K clusters, each cluster containing a set of semantically similar logical predicates to achieve effective organization and classification of logical predicates.
[0038] According to a zero-shot stance detection method based on enhanced cognitive pattern graph provided by the present invention, the K-means clustering algorithm groups semantically similar logical predicates by minimizing the within-cluster sum of squared distances, expressed as the following formula:
[0039]
[0040] where K is the number of clusters, Ci is the set of points in cluster i, μi is the center of cluster i, and the optimal number of clusters is selected by evaluating the silhouette scores of different numbers of clusters to ensure the rationality of the clusters.
[0041] According to a zero-shot stance detection method based on enhanced cognitive pattern graph provided by the present invention, it also includes using a pre-trained large language model LLM to summarize the text in each cluster to generate a concise phrase as the summary phrase of the cluster. The specific implementation steps are as follows:
[0042] For each cluster obtained by the clustering algorithm, extract all the logical predicates or corresponding text fragments contained therein;
[0043] Construct an input prompt P2, which is used to guide the large language model LLM to understand the task requirements, that is, to summarize the text in each cluster and generate a concise core description;
[0044] Take the input prompt P2 and the text fragments in each cluster as inputs and provide them to the pre-trained large language model LLM;
[0045] The large language model LLM summarizes the text in each cluster according to the input prompt P2 and the text fragments in the cluster, generating a concise phrase that reflects the core semantics of the cluster;
[0046] The generated concise phrase is used as the summary phrase of the cluster.
[0047] According to a zero-shot stance detection method based on cognitive pattern graph enhancement provided by the present invention, a weighted multi-relational graph is constructed as a cognitive pattern, including:
[0048] The cognitive pattern is defined as a weighted multi-relational graph Gs=(Vs, Es, As), where:
[0049] Vs represents the set of nodes, and each node corresponds to a summary phrase of a cluster;
[0050] Es represents the set of edges, connecting clusters based on the logical relationships between logical predicates, and using the Z3 Prover tool to detect and eliminate conflicting relationships;
[0051] As represents the weight of the edges, corresponding to the occurrence frequency of the inter-cluster logical relationships.
[0052] According to a zero-shot stance detection method based on cognitive pattern graph enhancement provided by the present invention, a relational graph convolutional network is used to model the constructed CFGraph, including:
[0053] Construct CFGraph, where CFGraph is a graph structure containing users, items, and the interaction relationships between them. Among them, users and items are nodes, and the interaction behaviors between users and items are edges;
[0054] Use the relational graph convolutional network RGCN to model CFGraph to learn the embedding representations of users and items in the graph structure. RGCN captures the complex interaction patterns and relationship features between users and items by aggregating and convolving the neighbor information of the nodes in CFGraph;
[0055] Take the user and item embedding representations learned by RGCN as inputs and input them into the feed-forward neural network layer;
[0056] In the feed-forward neural network layer, perform non-linear transformation and feature extraction on the input user and item embedding representations to learn high-order features for stance prediction;
[0057] According to the output of the feed-forward neural network layer, use a classifier to predict the stance of the user or item. The stance includes at least a positive stance, a negative stance, or a neutral stance;
[0058] Output the stance prediction result.
[0059] As can be seen, compared with the prior art, the present invention has the following beneficial effects:
[0060] 1. Significantly reduce the dependence on labeled data: The present invention ingeniously uses large language models (LLMs) to construct concept-level knowledge in an unsupervised manner, greatly reducing the dependence on a large amount of labeled data. Through unsupervised learning, the model can autonomously extract useful information and patterns from a vast amount of text, thereby achieving stance detection for unseen targets, improving the practicality and flexibility of the model. This not only reduces the costs of data collection and annotation but also enables the model to adapt to new fields and emerging topics more quickly, meeting the needs of real-time changes.
[0061] 2. Significantly enhance semantic understanding and reasoning capabilities: The present invention transforms first-order logic (FOL) expressions into cognitive patterns and combines them with relational graph convolutional networks (RGCNs) for modeling. This innovative combination enables the model to better capture semantic information and logical structures in text. By constructing a cognitive-FOL graph (CFGraph), the model can effectively integrate logical reasoning and semantic understanding, enabling it to more accurately understand the meaning of text and make reasonable inferences in complex and changing stance detection tasks, significantly enhancing the adaptability of the model in the zero-shot setting and enabling it to better handle unseen targets and situations.
[0062] 3. Enhance the generalization ability and accuracy of the model: The present invention effectively makes up for the deficiencies of LLMs in stance detection tasks by combining domain-specific knowledge and advanced prompt learning methods. The introduction of domain-specific knowledge enables the model to better adapt to the needs of specific domains and improve the accuracy of detection. Advanced prompt learning methods enable the model to more efficiently utilize limited information for reasoning and judgment, further enhancing the generalization ability of the model.
[0063] In summary, the present invention provides an efficient, accurate, and practical solution for zero-shot stance detection tasks by reducing the dependence on labeled data, enhancing semantic understanding and reasoning capabilities, and enhancing the generalization ability and accuracy of the model.
[0064] The following further elaborates on the present invention in conjunction with the accompanying drawings and specific embodiments. Description of the Drawings
[0065] Figure 1 is a flowchart of an embodiment of a zero-shot stance detection method based on cognitive pattern graph enhancement according to the present invention.
[0066] Figure 2 is a schematic diagram of an embodiment of a zero-shot stance detection method based on cognitive pattern graph enhancement according to the present invention. Detailed Embodiments
[0067] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts fall within the scope of protection of the present invention.
[0068] As used herein, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0069] See Figure 1 and Figure 2 , this embodiment provides a zero-shot stance detection method based on enhanced cognitive mode graph, and the method includes the following steps:
[0070] Step S1, constructing a logical cognitive mode:
[0071] Using a pre-trained large language model LLM to convert the input text into a logical expression; clustering the logical predicates in the logical expression and organizing the logical predicates into different clusters; summarizing the text in each cluster to generate a summary phrase for the cluster; constructing a weighted multi-relationship graph as the cognitive mode according to the logical relationships between the logical predicates and the summary phrases of the clusters; where the nodes of the graph represent the summary phrases of the clusters, the edges represent the logical relationships between the logical predicates, and the weights of the edges correspond to the occurrence frequencies of the logical relationships;
[0072] Step S2, cognitive-enhanced stance prediction:
[0073] In the prediction stage, converting the text to be predicted into a logical expression, and extracting a subgraph related to the logical expression of the text to be predicted from the constructed cognitive mode; combining the extracted subgraph with the logical expression of the text to be predicted to construct a cognitive-first-order logic rule graph CFGraph; using a relational graph convolutional network to model the constructed CFGraph for stance prediction; selecting the category with the highest probability as the prediction result.
[0074] Among them, the model constructs a cognitive mode: using a pre-trained model (such as a large language model, LLM) to convert the input text into a logical expression. Subsequently, through clustering and summarization of logical predicates, a cognitive mode is constructed. The cognitive mode is stored in the form of a graph and contains the core information of semantic and logical relationships.
[0075] Among them, the small model predicts the stance based on the cognitive mode: in the prediction stage, relevant information is extracted and a cognitive map is constructed according to the input text and the cognitive mode. The cognitive map is modeled by a graph neural network (such as a relational graph convolutional network, RGCN), and the neural network is used to predict the stance of the text towards the target. Finally, the category with the highest probability is selected as the prediction result.
[0076] Among them, in the input stage, an input text containing at least one logical statement is received, such as "Silence is golden, yes, but if there is only silence..."; the large language model LLM is used to perform semantic analysis on the input text, extract the predicates and their logical relationships in the text, and generate corresponding logical expressions; the predicates in the logical expressions are transformed into nodes in the first-order logic rule graph, and each node represents an independent predicate concept; according to the logical relationships between the predicates in the logical expressions, edges are established between the nodes in the first-order logic rule graph, and the edges represent the logical associations between the predicates, including at least logical relationships such as implication, conjunction, disjunction, and negation. The purpose of transforming the input text into a graph in this embodiment is to better combine it with the cognitive mode graph. The graph composed of the first-order logic rule text can be more perfectly integrated with the cognitive mode graph in terms of structure and semantics. This is because the first-order logic rule text is based on logical predicates and logical relationships, and can clearly express the semantics and logical structure in the text, and has a high degree of compatibility with the concept nodes and relationship edges in the cognitive mode graph.
[0077] Then, based on the constructed first-order logic rule graph, a matching cognitive mode subgraph is extracted from the weighted multi-relational graph of the cognitive mode, and the cognitive mode subgraph is fused with the first-order logic rule graph to obtain the cognitive-first-order logic rule graph CFGraph.
[0078] In the above step S1, the following specific steps are also included to generate logical expressions using the pre-trained large language model LLM:
[0079] Construct an instruction prompt P1, which is used to guide the large language model to understand the task requirements, that is, to convert the input text into a corresponding first-order logic expression;
[0080] Prepare a sample input set, where each sample includes a social text and a target first-order logic (FOL) expression corresponding to the social text, and the target first-order logic expression contains logical predicates (such as F(x), S(x), and C(x)) and logical relationships (such as "implication" → and "and" ∧).
[0081] Among them, P1: Your task is to analyze the attitude of the [given sentence] towards the [given target] using first-order logic. Formulate a response and conclude with a statement indicating the attitude (Support, Opposed, Neutral).
[0082] Provide the constructed instruction prompt P1 and all sample inputs to the pre-trained large language model LLM for model training or fine-tuning, so that the model learns how to convert social text into first-order logic expressions.
[0083] After the model training or fine-tuning is completed, use this large language model LLM to receive new input text and convert the new input text into the corresponding first-order logic expression according to the learned mapping relationship. This expression also contains logical predicates and the logical relationships between them.
[0084] When performing model training or fine-tuning, encode the instruction prompt P1 into a format understandable by the model and pair it with each sample in the sample input set to form training data pairs.
[0085] Input the training data pairs into the pre-trained large language model LLM and set training parameters such as learning rate, number of training epochs, and batch size, etc.
[0086] Through the supervised learning method, enable the model to gradually learn the task specified by the instruction prompt P1 during the training process, that is, map the semantic information in social text to the logical predicates and logical relationships in the first-order logic expression.
[0087] During the training process, use a loss function to evaluate the difference between the logical expression generated by the model and the target logical expression, and update the model's parameters through the backpropagation algorithm to minimize this difference.
[0088] When the performance of the model on the training set reaches the predetermined standard, stop training or perform fine-tuning to further optimize the model's performance.
[0089] The model after training or fine-tuning is used to receive new social text inputs and automatically generate the corresponding first-order logic expressions according to the learned mapping relationship.
[0090] In the above step S1, it also includes using the K-means clustering algorithm to organize the logical predicates in the first-order logic FOL expressions into different clusters. The specific implementation steps are as follows:
[0091] Convert each logical predicate into a d-dimensional embedding vector e through the Sentence-BERT model to fully capture its semantic details and features;
[0092] Initialize the number of clusters K of the K-means clustering algorithm and randomly select K initial cluster center vectors;
[0093] For the d-dimensional embedding vector e corresponding to each logical predicate, calculate its distance from each cluster center vector and assign the vector to the cluster with the closest distance;
[0094] According to the attribution results, update the center vector of each cluster to the mean vector of all embedding vectors within the cluster;
[0095] Repeat the embedding vector step until the cluster center vectors no longer change significantly or reach the preset number of iterations to minimize the within-cluster sum of squared distances;
[0096] Finally, obtain K clusters, each cluster containing a set of semantically similar logical predicates to achieve effective organization and classification of logical predicates.
[0097] In this embodiment, the K-means clustering algorithm groups semantically similar logical predicates by minimizing the within-cluster sum of squared distances (latex formula: \text{minimize}\sum_{i=1}^K\sum_{e\in C_i}\|e-\mu_i\|^2), which is expressed as the following formula:
[0098]
[0099] Among them, K is the number of clusters, Ci is the set of points in cluster i, μi is the center of cluster i, and the optimal number of clusters is selected by evaluating the silhouette scores of different numbers of clusters to ensure the rationality of the clusters.
[0100] In the above step S1, it also includes using a pre-trained large language model LLM to summarize the text in each cluster to generate a concise phrase as the summary phrase of the cluster. The specific implementation steps are as follows:
[0101] For each cluster obtained by the clustering algorithm, extract all the logical predicates or corresponding text fragments contained therein;
[0102] Construct an input prompt P2, which is used to guide the large language model LLM to understand the task requirements, that is, to summarize the text in each cluster and generate a concise core description to compress a large number of semantically similar text fragments;
[0103] Among them, P2: You are provided with several descriptions, each representing a predicate in first-order logic. Your task is to create new descriptions that summarize the main points of these predicates.
[0104] Use the input prompt P2 and the text fragments in each cluster as input and provide them to the pre-trained large language model LLM.
[0105] The large language model LLM summarizes the text in each cluster according to the input prompt P2 and the text fragments in the cluster, and generates concise phrases that reflect the core semantics of the cluster.
[0106] Use the generated concise phrases as the summary phrases of the cluster for subsequent text processing, information retrieval, or knowledge representation tasks to compress a large number of semantically similar text fragments and improve processing efficiency and accuracy.
[0107] In the above step S1, construct a weighted multi-relational graph as the cognitive pattern, including:
[0108] The cognitive pattern is defined as a weighted multi-relational graph Gs = (Vs, Es, As), where:
[0109] Vs represents the set of nodes, and each node corresponds to the summary phrase of a cluster.
[0110] Es represents the set of edges, connects the clusters based on the logical relationships between logical predicates, and uses the Z3 Prover tool to detect and eliminate conflicting relationships.
[0111] As represents the weights of the edges, corresponding to the occurrence frequencies of the inter-cluster logical relationships.
[0112] In the above step S2, use the relational graph convolutional network to model the constructed CFGraph, including:
[0113] Construct CFGraph, which is a graph structure containing users, items, and the interaction relationships between them. Among them, users and items are nodes, and the interaction behaviors between users and items are edges.
[0114] Use the relational graph convolutional network RGCN to model CFGraph to learn the embedding representations of users and items in the graph structure. RGCN captures the complex interaction patterns and relationship features between users and items by aggregating and convolving the neighbor information of the nodes in CFGraph.
[0115] Take the user and item embedding representations learned by the RGCN as inputs and input them into the feed-forward neural network layer;
[0116] In the feed-forward neural network layer, perform non-linear transformation and feature extraction on the input user and item embedding representations to learn high-order features for stance prediction;
[0117] According to the output of the feed-forward neural network layer, use a classifier to predict the stance of the user or item, and the stance includes at least positive stance, negative stance or neutral stance;
[0118] Output the stance prediction result.
[0119] In summary, regarding the construction and application of the logical cognitive mode: The present invention proposes a zero-shot stance detection (LCASD) framework enhanced based on the logical cognitive mode. The core innovation lies in using large language models (LLMs) to construct concept-level knowledge in an unsupervised manner and transforming it into a cognitive mode. This mode simulates the human stance detection process through first-order logic (FOL) expressions, thereby enhancing the model's ability to understand semantics and logic. This technical solution not only improves the model's reasoning ability and adaptability in the zero-shot setting but also significantly reduces the dependence on labeled data, providing a new technical path for zero-shot stance detection.
[0120] Regarding the semantic-enhanced relational graph convolutional network (SERGCN): The present invention further proposes a semantic-enhanced relational graph convolutional network (SERGCN) for integrating the cognitive mode with the input text. Specifically, by extracting subgraphs related to the input logic from the cognitive mode and combining them with FOL expressions, a cognitive-FOL graph (CFGraph) is constructed. This graph structure integrates logical nodes and concept nodes and is modeled through a relational graph convolutional network (RGCN) to finally achieve accurate prediction of the stance. The present invention combines the cognitive mode with the graph neural network, significantly improving the model's ability to model complex semantic and logical relationships.
[0121] Regarding the efficient implementation of zero-shot stance detection: The present invention proposes a complete technical solution that simulates the human cognitive reasoning process for the zero-shot stance detection task. Through the construction of the logical cognitive mode and the semantic-enhanced graph neural network, the present invention can accurately detect the stance of unseen targets without a large amount of labeled data. This technical solution not only improves the generalization ability of the model but also provides an efficient and practical solution for the field of zero-shot stance detection, with significant innovation and practicality.
[0122] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0123] The above embodiments are only the preferred embodiments of the present invention, and the scope of protection of the present invention cannot be limited thereby. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention belong to the scope of protection required by the present invention.
Claims
1. A zero-sample stance detection method based on cognitive pattern graph enhancement, characterized in that: The following steps are involved: Constructing logical cognitive models: Use the pre-trained large language model LLM to convert the input text into a logical expression; cluster the logical predicates in the logical expression and organize the logical predicates into different clusters; summarize the text in each cluster and generate a summary phrase for the cluster; According to the logical relations between logical predicates and the summary phrases of clusters, a weighted multi-relation graph is constructed as a cognitive model; Cognitive enhancement stance predictions: In the prediction stage, the text to be predicted is converted into a logical expression, and the subgraph related to the logical expression of the text to be predicted is extracted from the constructed cognitive model; the extracted subgraph is combined with the logical expression of the text to be predicted to construct a cognitive-first-order logic rule graph CFGraph; the constructed CFGraph is modeled using a relational graph convolutional network to perform stance prediction; Select the category with the highest probability as the prediction result.
2. The method according to claim 1, characterized in that: In the input stage, an input text containing at least one logical statement is received; a large language model LLM is used to perform semantic analysis on the input text, extract predicates and their logical relations in the text, and generate corresponding logical expressions; the predicates in the logical expressions are converted into nodes in a first-order logic rule graph, each node representing an independent predicate concept; according to the logical relations between the predicates in the logical expressions, edges between nodes are established in the first-order logic rule graph, and the edges represent the logical associations between the predicates, including at least logical relations such as implication, conjunction, disjunction, and negation.
3. The method according to claim 2, characterized in that: Based on the constructed first-order logic rule graph, the matching cognitive pattern subgraph is extracted from the weighted multi-relation graph of cognitive pattern, and the cognitive pattern subgraph is fused with the first-order logic rule graph to obtain the cognitive-first-order logic rule graph CFGraph.
4. The method according to claim 1, characterized in that: The following specific steps are also included to use the pre-trained large language model LLM for logical expression generation: Construct instruction prompt P1, which is used to guide the large language model to understand the task requirements, that is, to convert the input text into the corresponding first-order logic expression; Prepare a sample input set, where each sample includes a social text and a target first-order logic expression corresponding to the social text, and the target first-order logic expression includes a logical predicate and a logical relationship; Provide the constructed instruction prompt P1 and all sample inputs to the pre-trained large language model LLM to train or fine-tune the model so that the model learns how to convert social text into first-order logic expressions; After the model training or fine-tuning is completed, the large language model LLM is used to receive new input text, and according to the learned mapping relationship, the new input text is converted into a corresponding first-order logic expression, which also contains logical predicates and the logical relations between them.
5. The method according to claim 4, characterized in that: When training or fine-tuning the model, the instruction prompt P1 is encoded into a format that the model can understand and paired with each sample in the sample input set to form a training data pair; Input the training data pairs into the pre-trained large language model LLM and set the training parameters; Through supervised learning, the model gradually learns the task specified by the instruction prompt P1 during the training process, that is, mapping the semantic information in the social text to the logical predicates and logical relations in the first-order logic expression; During the training process, a loss function is used to evaluate the difference between the logical expression generated by the model and the target logical expression, and the parameters of the model are updated through the back-propagation algorithm to minimize the difference; When the performance of the model on the training set reaches the predetermined standard, stop training or perform fine-tuning to further optimize the performance of the model; The trained or fine-tuned model is used to receive new social text input and automatically generate the corresponding first-order logic expression based on the learned mapping relationship.
6. The method according to claim 1, characterized in that It also includes using the K-means clustering algorithm to organize the logical predicates in the first-order logic FOL expression into different clusters. The specific implementation steps are as follows: Each logical predicate is converted into a d-dimensional embedding vector e through the Sentence-BERT model to fully capture its semantic details and characteristics; Initialize the number of clusters K of the K-means clustering algorithm and randomly select K initial cluster center vectors; For each d-dimensional embedding vector e corresponding to a logical predicate, calculate its distance from the center vector of each cluster and assign the vector to the cluster with the closest distance; According to the attribution results, the center vector of each cluster is updated to the mean vector of all embedded vectors in the cluster; Repeat the embedding step until the cluster center vector no longer changes significantly or the preset number of iterations is reached to minimize the intra-cluster sum of squared distances. Finally, K clusters are obtained, each of which contains a group of logical predicates with similar semantics, so as to achieve effective organization and classification of logical predicates.
7. The method according to claim 6, characterized in that: The K-means clustering algorithm groups semantically similar logical predicates by minimizing the intra-cluster sum of squared distances, expressed as the following formula: Among them, K is the number of clusters, Ci is the set of points in cluster i, μi is the center of cluster i, and the optimal number of clusters is selected by evaluating the silhouette scores of different numbers of clusters to ensure the rationality of the clusters.
8. The method according to claim 1, characterized in that It also includes summarizing the text in each cluster using a pre-trained large language model LLM to generate a concise phrase as the summary phrase of the cluster. The specific implementation steps are as follows: For each cluster obtained by the clustering algorithm, all logical predicates or corresponding text fragments contained therein are extracted; Construct an input prompt P2, which is used to guide the large language model LLM to understand the task requirements, that is, to summarize the text in each cluster and generate a concise core description; Provide the input prompt P2 and the text fragments in each cluster as input to the pre-trained large language model LLM; The large language model LLM summarizes the text in each cluster based on the input prompt P2 and the text fragments in the cluster, and generates a concise phrase that reflects the core semantics of the cluster; The generated concise phrase is used as the summary phrase of the cluster.
9. The method according to any one of claims 1 to 8, characterized in that: The constructing of a weighted multi-relationship graph as a cognitive model includes: The cognitive model is defined as a weighted multi-relation graph Gs = (Vs, Es, As), where: Vs represents a set of nodes, each node corresponds to a summary phrase of a cluster; Es represents the edge set, which connects clusters based on the logical relationships between logical predicates, and uses the Z3Prover tool to detect and eliminate conflicting relationships; As represents the weight of the edge, which corresponds to the frequency of occurrence of the logical relationship between clusters.
10. The method according to any one of claims 1 to 8, characterized in that: The method of modeling the constructed CFGraph using a relational graph convolutional network includes: Construct CFGraph, which is a graph structure containing users, projects, and the interactions between them, where users and projects are nodes and the interactions between users and projects are edges; CFGraph is modeled using the relational graph convolutional network (RGCN) to learn the embedded representation of users and items in the graph structure. RGCN captures the complex interaction patterns and relational features between users and items by aggregating and convolving the neighbor information of nodes in CFGraph. The user and item embedding representations learned by RGCN are used as input to the feed-forward neural network layer; In the feedforward neural network layer, nonlinear transformation and feature extraction are performed on the input user and item embedding representations to learn high-order features for stance prediction; According to the output of the feed-forward neural network layer, a classifier is used to predict the position of the user or item, and the position includes at least a positive position, a negative position or a neutral position; Output the stance prediction result.