Personalized cognitive path extraction method based on large language model

By combining similarity search and prompt engineering methods, targeted prompt strategies are designed to guide large language models to extract cognitive paths, solving the problem of lack of targeted guidance and insufficient interpretability in cognitive path extraction of existing models, and achieving more accurate and interpretable cognitive path generation.

CN120067299AActive Publication Date: 2025-05-30BEIJING UNIV OF TECH

Patent Information

Application Number
CN202510212807.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

Existing large language models lack targeted guidance prompt design in cognitive path extraction tasks, making it difficult to capture complex cognitive logic chains and key nodes, and lack of interpretability, which limits its application in the field of psychology.

Method used

A personalized cognitive path extraction method based on large language models is proposed. By combining similarity search and prompt engineering, basic prompts, few sample prompts and thinking chain prompts are designed to gradually guide the model to extract cognitive paths, and optimize the generated paths through automated inspection and large model self-evaluation feedback.

Benefits of technology

It significantly improves the completeness and interpretability of cognitive path extraction, and the generated cognitive path has high accuracy and psychological consistency, and is suitable for practical application scenarios such as clinical psychology and educational intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067299A_ABST
    Figure CN120067299A_ABST
Patent Text Reader

Abstract

The invention provides a personalized cognitive path extraction method based on a large language model, and relates to natural language processing. The method comprises the following steps of: firstly, preprocessing and coding an individual statement text; then, by calculating the similarity, similar cognitive path extraction examples are retrieved from a cognitive path database, and reference is provided for few-sample prompt construction; thirdly, designing a comprehensive prompt strategy, and guiding the large language model to gradually and accurately extract cognitive paths of the individuals in combination with basic prompt, few-sample prompt and thinking chain prompt; besides, automatic inspection and large-scale language model self-evaluation feedback modes are introduced, and consistency verification and optimization are carried out on the generated cognitive path, so that the accuracy and reliability of an extraction result are ensured. According to the method, the explanatability of a model reasoning process is enhanced and the efficiency and accuracy of personalized cognitive path extraction are improved by combining example retrieval and prompt engineering and particularly introducing a thinking chain prompt strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for extracting personalized cognitive paths based on large language models, belonging to the field of natural language processing. Background Art

[0002] In recent years, large language models (LLMs) have made remarkable progress in the field of natural language processing. They have demonstrated excellent performance in tasks such as text understanding and generation, and sentiment analysis, providing new technical means for text analysis in the field of psychology. However, although LLMs have been preliminarily applied in aspects such as emotion recognition and psychological state analysis, there are still obvious gaps in the specific field of cognitive path extraction, which are specifically reflected in the following aspects:

[0003] 1. Lack of targeted guidance in prompt design: Existing large language models mostly rely on prompts for task-driven operations. However, in the field of psychology, general prompts may not fully reflect the dynamics and contextuality of individual cognitive processes. This results in the model's difficulty in effectively capturing complex cognitive logic chains and key nodes.

[0004] 2. Insufficient interpretability: The task of cognitive path extraction essentially involves a layer-by-layer analysis of complex psychological processes, which poses a relatively high requirement for the interpretability of the model. However, existing language models often exhibit "black box" characteristics when generating results, that is, it is difficult to clearly show the reasoning logic from input to output. This not only hinders researchers' understanding and optimization of the model's behavior but also limits its credibility and adoption in practical application scenarios such as clinical psychology and educational intervention. Summary of the Invention

[0005] In response to the above challenges, the present invention proposes a method for extracting personalized cognitive paths based on large language models. This method combines similarity retrieval and prompt engineering to guide the large language model to deeply understand and accurately extract information related to cognitive paths, thereby generating more reasonable and accurate results. Specifically, the basic prompt clarifies the task objective, helping the model establish a clear task direction; the few-shot prompt combines examples of cognitive path extraction to improve the model's understanding and adaptation ability for the task; and the chain-of-thought prompt gradually decomposes the task structure, sequentially guiding the model to identify the core components of the cognitive path, including triggering events, irrational beliefs, emotional reactions, and refutation methods. Through this systematic prompt strategy design, the integrity and interpretability of cognitive path extraction are significantly improved. In addition, the present invention also introduces an automated inspection and self-evaluation feedback of the large model to further improve the accuracy and psychological consistency of the generated paths, making the model output more in line with the actual application requirements of the field of psychology.

[0006] The technical solution of the present invention includes the following steps: First, preprocess and encode the individual statement data; then, use semantic similarity technology to retrieve similar examples from the cognitive path database to provide reference and support for subsequent few-shot prompting; subsequently, design a comprehensive prompting strategy to organically combine basic prompting, few-shot prompting, and chain-of-thought prompting to guide the large language model to gradually extract the cognitive path of the patient; finally, automatically check the generated cognitive path and further optimize the path from three dimensions: fluency, accuracy, and psychological professionalism in combination with the self-evaluation feedback of the large model, so as to ensure the quality and application value of the cognitive path.

[0007] The specific solution of the present invention is as shown in the appendix Figure 1 .

[0008] Step 1: Receive input text and encoding

[0009] This step aims to preprocess and encode the statement text provided by the individual, extract useful features, and lay a foundation for subsequent steps.

[0010] Step 1.1: Text preprocessing. Receive the statement text X. To protect the privacy of the individual and improve data quality, preprocess the text, remove its privacy information, and clean meaningless characters.

[0011] Step 1.2: Text semantic encoding. Use the BERT model to generate semantic embeddings for each segmented sentence s i . The BERT can capture the context semantic features of the text and generate high-dimensional vector representations to express the semantic characteristics of the sentence. The embedding vectors of all sentences constitute the semantic feature set E semantic .

[0012] Step 1.3: Emotional feature extraction and fusion. Use an emotion dictionary to match the emotion words in the text and extract the emotional features of each sentence. Integrate the emotional features with the semantic feature vectors generated by BERT to form emotion-enhanced semantic feature vectors E enhanced .

[0013] Step 1.4: Normalization processing: Perform L2 norm normalization on the emotion-enhanced semantic feature vectors E enhanced to eliminate the scale difference of the features and improve the consistency of the model's processing of the features. The normalized feature set is denoted as E final .

[0014] Step 1.5: Mean pooling aggregation: Generate the overall feature embedding E of the statement text through the mean pooling operation on the sentence feature vectors query . This global feature will be used as the query vector in the retrieval stage and provide input for the next similar example retrieval.

[0015] Step 2: Data Preprocessing

[0016] In this step, through vector retrieval technology, the most similar examples to the individual statement text are obtained from the cognitive path knowledge base to build a reference for few-shot prompting.

[0017] Step 2.1: Knowledge Base Construction. Build the cognitive path database D, and all data are annotated by psychology experts to ensure the authority and accuracy of the annotation results. The cognitive path of each text X i is annotated as L i . Use BERT to encode each text in the knowledge base to generate the embedding vector E i , and use FAISS to build an index library.

[0018] Step 2.2: Retrieval and Sorting. Based on the overall feature vector E query of the statement text, perform vector retrieval operations through the FAISS index library, and the retrieved examples are sorted according to the cosine similarity to ensure the maximization of the similarity of the filtered statement text.

[0019] Step 2.3: Example Screening. Since the cognitive path structure is relatively complex, the extraction results often contain long text sequences, and there is a length limit for the input tokens of the large model. Therefore, in this step, 2 cognitive path examples with the highest similarity are screened based on the cosine similarity, that is Among them, represents the 2 similar example texts retrieved, is the corresponding cognitive path extraction result.

[0020] Step 3: Prompt Engineering Construction

[0021] Prompt engineering is the key step to guide the large language model to complete cognitive path extraction. In this step, a comprehensive strategy that combines basic prompts, few-shot prompts, and chain-of-thought prompts is designed, as shown in the example Figure 3 as follows:

[0022] Step 3.1: Basic Prompt Prompt basic : Clearly define the role and task objectives of the model.

[0023] Step 3.2: Few-shot Prompt Prompt shot : Combine the example texts retrieved in Step 2 to provide specific references for the model. These example texts and their corresponding cognitive path annotations help the model understand the task details and generation requirements.

[0024] Step 3.3: Chain-of-thought Prompt Prompt CoT: According to the structural characteristics of the cognitive path, the step-by-step guidance model sequentially extracts four core parts: triggering events, irrational beliefs, emotional responses, and refutation methods. Each step of the prompt clarifies specific tasks, helping the model shift from a global understanding to a gradually refined task output.

[0025] Step 3.3: Prompt combination: Integrate the basic prompt, few-shot prompt, and chain-of-thought prompt into a complete prompt Prompt, providing multi-level and step-by-step task guidance.

[0026] Step 4: Cognitive path extraction

[0027] Input the designed comprehensive prompt and the statement text into the large language model to generate a preliminary cognitive path.

[0028] Step 5: Cognitive path optimization

[0029] To improve the quality of the generated path, it is adjusted and improved by combining two stages: automatic optimization and self-evaluation feedback optimization of the large model.

[0030] Step 5.1: Automatic optimization: Through consistency checking, verify whether the subcategory logically matches its parent category, as shown in the cognitive path parent and child nodes Figure 2 as shown.

[0031] Step 5.2: Self-evaluation feedback generation and optimization: Further submit the cognitive path that has passed the automated check to the LLM for scoring based on three dimensions: fluency, accuracy, and psychological professionalism. The full score for each dimension is 10 points, and a score of ≥8 is considered qualified. If all dimensions are qualified, the cognitive path can be directly output; otherwise, the LLM generates self-evaluation feedback F self , and adjusts and optimizes the generated cognitive path according to the feedback to generate the final version of the cognitive path:

[0032] The above technical solution has the following advantages or beneficial effects:

[0033] First, the personalized cognitive path extraction method based on the large language model proposed by the present invention makes full use of the prompt design template, including the basic prompt, few-shot prompt, and chain-of-thought prompt, and gradually guides the large language model to complete the complex extraction task of the cognitive path, effectively improving the model's understanding ability and generation quality of the task; in addition, the prompt template can be adjusted according to specific needs, and the method has good scalability.

[0034] Second, the personalized cognitive path extraction method based on the large language model proposed by the present invention is based on the theoretical framework of cognitive behavioral therapy, and combines the chain-of-thought prompt to gradually extract the core components of the cognitive path. In this way, the generated cognitive path has high accuracy and interpretability. Description of the drawings

[0035] Figure 1 This is the overall architecture diagram of the method proposed by the present invention.

[0036] Figure 2 This is the ABCD model diagram of the cognitive path.

[0037] Figure 3 This is the detailed process schematic diagram of the method proposed by the present invention Specific implementation manner

[0038] The following combines the specification drawings to elaborate on the implementation examples of the present invention:

[0039] The present invention is a personalized cognitive path extraction method based on a large language model. This method combines semantic similarity retrieval and prompt engineering to guide the large language model from multiple perspectives to deeply understand and accurately extract information related to an individual's cognitive path.

[0040] The present invention adopts the following technical solutions: First, preprocess and encode the individual's statement data; then, use semantic similarity technology to retrieve examples from the cognitive path database annotated by experts to support subsequent few-shot prompting; subsequently, design a comprehensive prompting strategy that combines basic prompting, few-shot prompting, and chain-of-thought prompting to guide the large language model to gradually extract personalized cognitive paths; finally, use an automatic optimization mechanism to perform consistency checks on the generated cognitive paths, and combine self-evaluation feedback to further optimize the paths from dimensions such as fluency, accuracy, and psychological professionalism, thereby ensuring the quality and application value of the cognitive paths.

[0041] Specifically, the method includes the following steps:

[0042] Step 1: Receive the input text X and encoding: Receive the individual's statement text, preprocess it, perform text semantic encoding using BERT, and then extract emotional features using an emotion dictionary and fuse the two.

[0043] Step 1.1: Text preprocessing

[0044] Obtain the statement text X, delete the privacy information in the text, and remove meaningless characters (HTML tags, special characters, and URL links), handle spelling mistakes, and use the toolkit NLTK to split the preprocessed statement text into sentences, that is, X = {s 1 , s 2 , …, s n};

[0045] Step 1.2: Text semantic encoding

[0046] For each sentence s iSemantic embedding generation is performed using the BERT model to extract semantic features, that is The sequence formed by the embedding vectors of all sentences is

[0047] Step 1.3: Emotional feature extraction and fusion

[0048] Use the Harbin Institute of Technology emotional dictionary to match the emotional words in each sentence, and then calculate the emotional features of each sentence And fuse it with the semantic embedding to form an emotion-enhanced semantic feature vector Finally, the feature vectors of all sentences are

[0049] Step 1.4: Normalization processing

[0050] For the emotion-enhanced semantic feature vector of each sentence Perform L 2 norm normalization to ensure feature consistency in subsequent processing, that is The set of normalized sentence features:

[0051] Step 1.5: Mean pooling aggregation

[0052] Use mean pooling to generate the overall feature embedding of the patient text for retrieval

[0053] Step 2: Data preprocessing

[0054] In this step, through vector retrieval technology, the most similar examples to the statement text are obtained from the labeled cognitive path knowledge base to build a reference for few-shot prompting.

[0055] Step 2.1: Knowledge base construction: For each statement text, 3 professional psychology experts with professional qualifications are organized to independently annotate according to the defined cognitive path labels. If the label annotation consistency among experts is less than 80%, the review mechanism is triggered to review and re-annotate the controversial data to ensure annotation consistency. Through the above operations, a cognitive path dataset is constructed Where X i represents an individual statement text, and L i is the labeled cognitive path. According to cognitive behavioral therapy, the defined cognitive path labels consist of four parts:

[0056] A: Triggering event (A1: Disease symptoms, A2: Social relations, A3: Life, A4: Learning and work, A5: Emotions)

[0057] B: Irrational Beliefs (B1: All-or-Nothing, B2: Overgeneralization, B3: Mental Filtering, B4: Disqualifying the Positive, B5: Jumping to Conclusions, B6: Magnification and Minimization, B7: Emotional Reasoning, B8: Should Statements, B9: Labeling, B10: Self / Other Blame)

[0058] C: Emotional Reactions (C1: Affective Effect, C2: Behavioral Effect)

[0059] D: Disputation (D1: Habitual Disputation, D2: Effective Disputation).

[0060] Finally, an index library Index is constructed based on FAISS. Specifically, during the FAISS index construction process, first, each text data X in the dataset is encoded using the BERT model i to obtain the feature vector representation E i , and these E i vectors are used as storage items in the index library Index for subsequent similar example retrieval, defined as follows:

[0061]

[0062] where BuildIndex represents the index construction function, implemented based on the vector retrieval algorithm FAISS.

[0063] S22 Retrieval and Sorting: In the retrieval stage, according to the query vector E obtained in step S1 query the similar index vectors E are retrieved from the index library Index i , and sorted according to the cosine similarity:

[0064]

[0065] where Retrieve(Index, E query , k) represents the vector retrieval operation, Sim(E query , E i ) is the similarity metric function, ||E query || and ||E i || represent the L2 norms of vectors E query and E i respectively, and are used for normalized dot product calculation.

[0066] S23 Screening: Since the cognitive path structure is relatively complex, the extraction results often contain long text sequences, and there is a length limit for the input tokens of the large model. Therefore, in this step, 2 cognitive path examples with the highest similarity are selected based on the cosine similarity, that is where represents the 2 similar example texts retrieved, It is the corresponding cognitive path extraction result.

[0067] Step 3: Prompt Engineering Construction

[0068] In this step, a comprehensive strategy integrating basic prompt, few-shot prompt, and chain-of-thought prompt is designed. An example is shown as Figure 3 follows:

[0069] Step 3.1: Basic Prompt basic : It is used to clarify the model's role (psychology expert) and task objective, so as to better guide it to extract the cognitive path based on psychological knowledge. The basic prompt template is:

[0070] As an experienced psychology expert, you focus on researching cognitive behavioral therapy and are good at analyzing the cognitive structure and thinking mode of individuals. Please extract the personalized cognitive path according to the given individual's statement text.

[0071] Step 3.2: Few-shot Prompt shot : Provide specific examples of cognitive path extraction to enable the model to better understand the task requirements and output format. The few-shot prompt template is:

[0072] Example 1:

[0073] {Individual statement text 1}

[0074] {Cognitive path extraction result 1}

[0075] Example 2:

[0076] {Individual statement text 2}

[0077] {Cognitive path extraction result 2}

[0078] Step 3.3: Chain-of-thought Prompt CoT : According to the structural characteristics of the cognitive path, construct a chain-of-thought prompt to gradually guide the large model to extract the four parts of triggering event, irrational belief, emotional reaction, and refutation in sequence. Specifically, the chain-of-thought prompt template is as follows:

[0079] The cognitive path consists of four parts: triggering event, irrational belief, emotional reaction, and refutation. Please extract these four parts from the statement text in sequence:

[0080] The first part: Identification of triggering event, that is, identify the key event that leads to the change of their emotion or cognition from the statement text, including five aspects: disease symptoms, social relations, life, study and work, and emotion;

[0081] Part 2: Identification of Irrational Beliefs. Identify the cognitive distortions present in the text, specifically referring to the ten cognitive distortions mentioned in "The New Mood Therapy" by David D. Burns (all-or-nothing thinking, overgeneralization, mental filtering, discounting the positive, jumping to conclusions, magnification and minimization, emotional reasoning, should statements, labeling, personalization / blame).

[0082] Part 3: Identification of Emotional Reactions, including two aspects: emotional effects and behavioral effects;

[0083] Part 4: Identification of Refutations. Extract the ways in which the individual refutes their own cognitive distortions from the text, including habitual refutations and effective refutations.

[0084] Finally, combine the content obtained from the four parts to obtain the final cognitive path.

[0085] Step 3.4 Prompt Engineering Construction: Combine the basic prompt, sample prompt, and chain-of-thought prompt into a complete Prompt:

[0086] Prompt = Prompt basic + Prompt shot + Prompt CoT

[0087] The complete Prompt is as follows:

[0088] As an experienced psychology expert specializing in cognitive-behavioral therapy and proficient in analyzing an individual's cognitive structure and thinking patterns, please extract the personalized cognitive path based on the given statement text of the individual.

[0089] Example 1:

[0090] {Individual statement text 1}

[0091] {Cognitive path extraction result 1}

[0092] Example 2:

[0093] {Individual statement text 2}

[0094] {Cognitive path extraction result 2}

[0095] The cognitive path consists of four parts: triggering event, irrational belief, emotional reaction, and refutation. Please extract these four parts from the statement text in sequence:

[0096] Part 1: Identification of Triggering Events, that is, identify the key events that lead to changes in their emotions or cognitions from the statement text, including five aspects: disease symptoms, social relationships, life, study / work, and emotions;

[0097] Part 2: Identification of Irrational Beliefs. Identify the cognitive distortions existing in the text, which refer to the ten cognitive distortions mentioned in "The New Mood Therapy" by Burns (all-or-nothing thinking, overgeneralization, mental filter, discounting the positive, jumping to conclusions, magnification and minimization, emotional reasoning, should statements, labeling, personalization / blame).

[0098] Part 3: Identification of Emotional Reactions, including two aspects: emotional effects and behavioral effects;

[0099] Part 4: Identification of Refutations. Extract the ways in which individuals refute their own cognitive distortions from the text, including habitual refutations and effective refutations.

[0100] Finally, combine the content obtained from the four parts to get the final cognitive path.

[0101] Step 4: Extraction of Cognitive Path

[0102] Input the above constructed Prompt and the statement text into the large language model to generate the corresponding cognitive path:

[0103]

[0104] Among them, LLM represents the large language model, Cognitive_pathway represents the extracted cognitive path, and p i represents the parent label in the cognitive path (i.e., the triggering event, irrational belief, emotional reaction, refutation mentioned in step S21), is the set of all sub-labels corresponding to the parent label p i (for example, the triggering event corresponds to five sub-labels: study / work, social relations, life, study / work, emotion; the sub-labels corresponding to the three parent labels of irrational belief, emotional reaction, and refutation can be specifically viewed in step S21). The detailed definitions of the parent and sub-labels of the cognitive path are as Figure 2 shown.

[0105] Step 5: Optimization of Cognitive Path

[0106] To improve the quality of the generated path, it is adjusted and improved by combining two stages: automated inspection and LLM feedback generation and optimization.

[0107] Step 5.1: Automated Inspection: In this stage, an automated consistency check is performed on the generated cognitive path to ensure the logical matching between sub-labels and parent labels. For example, the "study / work" label extracted from the text should belong to the parent label of triggering event, rather than refutation. The specific operations are as follows:

[0108] By traversing each parent category p i and sub-category Check subcategory Whether it belongs to the set of legal subcategories of the parent category p i

[0109]

[0110] For all items, mark them as inconsistent subcategories for improvement in terms of accuracy in the next step:

[0111]

[0112] where represents a subcategory node under the parent category p i

[0113] After the above processing, the cognitive pathway consists of two parts:

[0114] Cognitive pathway = Cp consistent + Cp inconsistent

[0115] Step 5.2: Feedback generation and optimization: In step S4, the LLM acts as the generator of the cognitive pathway. In this stage, the LLM acts as the evaluator and optimizer. The cognitive pathway that has undergone automated inspection is further submitted to the LLM for scoring based on three dimensions: fluency, accuracy, and psychological professionalism. The full score for each dimension is 10 points, and a score of ≥8 is considered qualified. If all dimensions are qualified, the cognitive pathway can be directly output; otherwise, the LLM generates a self-evaluation feedback F self , and adjusts and optimizes the generated cognitive pathway according to the feedback to generate the final version of the cognitive pathway:

[0116] F self = {f 1 , f 2 , f 3}

[0117] Cognitive pathway optimized = Optimize(Cognitive pathway, F self )

[0118] where f 1 , f 2 and f 3 ​​They respectively represent the feedback suggestions generated by the LLM from three scoring perspectives, including error descriptions and adjustment suggestions. Optimize represents the process in which the LLM automatically adjusts and optimizes the cognitive path based on the feedback. If there are still dimensions with scores lower than 8 points after optimization, the LLM will perform secondary optimization until all evaluation criteria are met. Finally, the qualified cognitive path is output as the final extraction result.

[0119] The detailed scoring rules are as follows:

[0120]

[0121] References:

[0122] [1] Wei J, Wang X, Schuurmans D, et al. Chain-of-thought prompting elicits reasoning in large language models[J]. Advances in neural information processing systems, 2022, 35: 24824-24837.

[0123] [2] Ji B, Liu H, Du M, et al. Chain-of-Thought Improves Text Generation with Citations in Large Language Models[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2024, 38(16): 18345-18353.

[0124] [3] Gu X, Chen X, Lu P, et al. AGCVT-prompt for sentiment classification: Automatically generating chain of thought and verbalizer in prompt learning[J]. Engineering Applications of Artificial Intelligence, 2024, 132: 107907.

Claims

1. A personalized cognitive path extraction method based on a large language model, characterized in that: The following steps are involved: S1: Input text reception and encoding: First, the statement text provided by the user is preprocessed, including removing personal identity information and redundant information; the statement text is encoded using the BERT model, and the sentiment analysis of the text is performed in combination with the sentiment dictionary, and the final vector representation E is obtained by fusing the BERT semantic embedding and the sentiment feature vector. query ; S2: Similar example retrieval: Retrieve cognitive path examples similar to the statement text from the constructed cognitive path knowledge base; S3: Prompt Engineering Construction: Design a comprehensive prompt strategy that combines basic prompts, few-sample prompts, and thought chain prompts to improve the accuracy and rationality of cognitive path extraction by gradually guiding large language models; S4: Cognitive path extraction: Input the constructed comprehensive prompt and statement text into the large language model to extract the corresponding cognitive path; S5: Cognitive path optimization: Combine automated checking with self-evaluation feedback from large models to perform consistency checks, error corrections, and optimization adjustments on the generated cognitive paths to improve their logical rationality, content accuracy, and consistency with psychological theories.

2. According to claim 1, a personalized cognitive path extraction method based on a large language model is characterized in that S2, as follows: S21 Knowledge Base Construction: For each statement text, three psychology experts with professional qualifications are organized to independently annotate it according to the defined cognitive path labels. If the labeling consistency between experts is less than 80%, the review mechanism is triggered to review and re-annotate the disputed data to ensure the consistency of the annotations. Through the above operations, the cognitive path dataset is constructed. Where X i Indicates the statement text, L i The cognitive pathways are labeled; according to cognitive behavioral therapy, the cognitive pathway labels are defined as consisting of four parts: composition: A: Triggering events (A1: Disease symptoms, A2: Social relationships, A3: Life, A4: Study and work, A5: Emotions) B: Irrational beliefs (B1: Either this or that, B2: Generalizing, B3: Psychological filtering, B4: Denying positive thinking, B5: Jumping to conclusions, B6: Magnifying and minimizing, B7: Emotional reasoning, B8: Should sentences, B9: Random labeling, B10: Blaming oneself / others) C: Emotional response (C1: emotional effect, C2: behavioral effect) D: refutation (D1: habitual refutation, D2: effective refutation); Finally, we build an index based on FAISS. Specifically, in the process of building the FAISS index, we first use the BERT model to analyze each text data X in the dataset. i Encode and get the feature vector representation E i , these E i The vector is used as a storage item in the index library Index for subsequent similar example retrieval and is defined as follows: Among them, BuildIndex represents the index building function, which is implemented based on the vector retrieval algorithm FAISS; S22 Retrieval and sorting: In the retrieval stage, according to the query vector E obtained in step S1, query Retrieve the index vector E similar to it from the index library Index i , and sorted by cosine similarity: Among them, Retrieve(Index,E query ,k) represents the vector retrieval operation, Sim(E query ,E i ) is the similarity measurement function, ||E query || and ||E i || respectively represent vector E query and E i The L2 norm of is used for normalized dot product calculation; S23 screening: Since the cognitive path structure is complex, the extraction results often contain long text sequences, and the input token of the large model has a length limit. Therefore, in this step, the two cognitive path examples with the highest similarity are screened based on cosine similarity, namely in, Indicates the two similar sample texts retrieved. Extract the results for the corresponding cognitive path.

3. According to claim 1, a personalized cognitive path extraction method based on a large language model is characterized in that S3, specifically as follows: S31: Basic Prompt basic : It is used to clarify the role (psychology expert) and task objectives of the large model, so as to better guide it to extract cognitive paths based on psychological knowledge. The basic prompt template is: As an experienced psychology expert, you focus on cognitive behavioral therapy and are good at analyzing individual cognitive structures and thinking patterns. Please extract the individual's personalized cognitive path based on the individual's statement text. S32: Prompt for few samples shot : Provide specific cognitive path extraction examples to enable the model to better understand task requirements and output formats; the few-sample prompt template is: Example 1: {Individual statement text 1} {Cognitive path extraction results 1} Example 2: {Individual statement text 2} {Cognitive path extraction results 2} S33: Thought Chain Prompt CoT : Based on the structural characteristics of cognitive paths, a thinking chain prompt is constructed to gradually guide the large model to extract the four parts of triggering events, irrational beliefs, emotional reactions, and refutations in turn; specifically, the thinking chain prompt template is as follows: The cognitive path consists of four parts: triggering events, irrational beliefs, emotional reactions, and refutations. Please extract these four parts from the statement text in turn: Part I: Trigger event identification, that is, identifying the key events that lead to emotional or cognitive changes from the statement text, including disease symptoms, social relationships, life, study and work, and emotions; Part II: Identification of irrational beliefs, identifying cognitive distortions in the text, which refers to the ten cognitive distortions mentioned in Burns' New Emotional Therapy (either-or, generalizing, psychological filtering, jumping to conclusions, magnification and reduction, emotional reasoning, should sentences, random labeling, and blaming oneself / others); Part III: Emotional response recognition, identifying the emotional response reflected in the text, including both emotional effects and behavioral effects; Part 4: Refutation identification: extracting from the text the individual's refutation of his or her own cognitive distortions, including habitual refutation and effective refutation; Finally, the contents obtained from the four parts are combined to obtain the final cognitive path; S34 prompt engineering construction: basic prompts, sample prompts and thought chain prompts are combined into a complete prompt: Prompt=Prompt basic +Prompt shot +Prompt CoT The complete prompt is as follows: As an experienced psychology expert, you focus on cognitive behavioral therapy and are good at analyzing individual cognitive structures and thinking patterns. Please extract the individual's personalized cognitive path based on the individual's statement text. Example 1: {Individual statement text 1} {Cognitive path extraction results 1} Example 2: {Individual statement text 2} {Cognitive path extraction results 2} The cognitive path consists of four parts: triggering events, irrational beliefs, emotional reactions, and refutations. Please extract these four parts from the statement text in turn: Part I: Trigger event identification, that is, identifying the key events that lead to emotional or cognitive changes from the statement text, including disease symptoms, social relationships, life, study and work, and emotions; Part II: Identification of irrational beliefs, identifying cognitive distortions in the text, which refers to the ten cognitive distortions mentioned in Burns' New Emotional Therapy (either-or, generalizing, psychological filtering, denying positive thinking, jumping to conclusions, magnification and reduction, emotional reasoning, should sentences, random labeling, and blaming oneself / others); Part III: Emotional response identification, including both emotional effects and behavioral effects; Part 4: Refutation identification, extracting from the text the individual's refutation of his or her own cognitive distortions, including habitual refutation and effective refutation; Finally, the contents of the four parts are combined to obtain the final cognitive path.

4. The personalized cognitive path extraction method based on a large language model according to claim 1 is characterized in that S4, specifically as follows: S41 prompt input and cognitive path generation: The above-constructed prompt and statement text are input into the large language model to generate the corresponding cognitive path: Among them, LLM represents the large language model, cognitive_pathway represents the extracted cognitive path, and p i Represents the parent label in the cognitive path, For the parent tag p i The corresponding set of all sub-tags.

5. According to claim 1, a personalized cognitive path extraction method based on a large language model is characterized in that S5, specifically as follows: S51: Automated Check: In this stage, the generated cognitive paths are automatically checked for consistency to ensure the logical matching between the sub-labels and the parent labels; for example, the "learning work" label extracted from the text should be attributed to the parent label of the triggering event, rather than refutation; the specific operations are as follows: By traversing each parent category p in the path i and subcategories Check Subcategory Whether it belongs to the parent category p i The set of legal subcategories of For all , mark them as inconsistent subcategory to improve on the accuracy dimension in the next step: in, Represents the parent category p i A subcategory node under After the above processing, the cognitive pathway consists of two parts: Cognitive pathway=Cp consistent +Cp inconsistent S52: Self-evaluation feedback generation and optimization: In step S4, the big model acts as the generator of cognitive paths. In this stage, the big model acts as an evaluator and optimizer. The cognitive paths that have undergone automated inspection are further submitted to the big model for scoring based on three dimensions: fluency, accuracy, and psychological professionalism. The full score of each dimension is 10 points, and scores ≥ 8 are considered qualified. If all dimensions are qualified, the cognitive path can be directly output, otherwise the big model generates self-evaluation feedback F according to the evaluation results. self , and adjust and optimize the generated cognitive path based on the feedback to generate the final version of the cognitive path: <h2 style=";text-align:left;direction:ltr">F<h2 style=";text-align:left;direction:ltr"> self <h2 style=";text-align:left;direction:ltr"> (f1,f2,f3) Cognitive pathway optimized =Optimize(Cognitive pathway,F self ) Among them, f1, f2 and f3 represent the feedback generated by the large model according to the three scoring dimensions, including error descriptions and adjustment suggestions; Optimize represents the process of the large model automatically adjusting and optimizing the cognitive path according to the feedback; if there are still dimensions with a score lower than 8 points after optimization, the large model will perform secondary optimization until all evaluation criteria are met; finally, the qualified cognitive path is output as the final extraction result.

Citation Information

Patent Citations

  • Small-sample chapter-level event extraction method based on large language model thinking chain

    CN118394941A

  • Large language model joint reasoning method for enhancing thinking chain prompt based on knowledge graph

    CN118940840A

  • System and methods for upsampling of decompressed data after lossy compression using a neural network

    US12058333B1

  • Deep Embedding for Natural Language Content Based on Semantic Dependencies

    US20180336183A1

  • Systems and methods for contextualized and quantized soft prompts for natural language understanding

    US20230342552A1

Cited By

  • Big language model-based ecological system culture service intelligent evaluation method

    CN121706939A