Personal thinking simulation method and system
Through multi-source data collection and privacy preprocessing, combined with graph attention network and five-dimensional knowledge graph, dynamic prompt words are generated, which solves the limitations and privacy protection issues of user behavior and thinking mode simulation in the existing technology, and achieves accurate personal thinking imitation and decision-making support.
Patent Information
- Application Number
- CN202510256287.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The existing technology has the limitations of single-dimensional and static analysis in user behavior and thinking mode simulation, and it is difficult to fully and accurately characterize user characteristics. It lacks privacy protection in the process of data collection and processing, and cannot effectively support the in-depth imitation and decision-making support of personal thinking.
By collecting multi-source data in the public domain, private domain and dynamic questionnaires for target users and performing privacy preprocessing, using a three-level segmentation strategy to extract features, using a graph attention network for cross-modal association, building a five-dimensional knowledge graph and vector storage database, combining the five-dimensional weight injection template and real-time environmental data to generate dynamic prompt words, and finally generating decision suggestions.
It realizes accurate imitation of personal thinking and efficient decision-making assistance, takes into account privacy protection, and can generate decision-making suggestions that meet user characteristics and actual needs.
Smart Images

Figure CN120256682A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field, and specifically, to a personal thinking imitation method and system. Background Art
[0002] With the development of artificial intelligence technology, the demand for simulating and understanding user behavior and thinking patterns is increasing day by day. Traditional user portraits have limitations such as single dimension and static analysis, and it is difficult to comprehensively and accurately depict user characteristics. At the same time, in the process of data collection and processing, privacy protection issues have also attracted much attention. Existing methods have deficiencies in integrating multi-dimensional information, dynamic modeling, and privacy enhancement processing, and cannot well meet the requirements of deeply imitating personal thinking and providing effective decision-making support. Therefore, a more advanced personal thinking imitation method is needed to break through the above technical bottlenecks. Summary of the Invention
[0003] The purpose of the present invention is to provide a personal thinking imitation method and system.
[0004] In a first aspect, an embodiment of the present invention provides a personal thinking imitation method, including:
[0005] Collecting multi-source data of a target user in the public domain, private domain, and dynamic questionnaire and performing privacy preprocessing to obtain desensitized behavior data;
[0006] Extracting features from the desensitized behavior data by using a three-level segmentation strategy, and performing cross-modal association through a graph attention network to obtain a multi-dimensional joint semantic vector;
[0007] Constructing a graph database and a vector storage database based on the multi-dimensional joint semantic vector, where the graph database is used to store a five-dimensional knowledge graph, and the five-dimensional knowledge graph includes a cognitive dimension, a behavior dimension, an emotional dimension, a social dimension, and an evolutionary dimension;
[0008] Constructing dynamic prompt words based on the five-dimensional knowledge graph, a preset five-dimensional weight injection template, and real-time environmental data;
[0009] Performing multi-level retrieval enhancement based on the dynamic prompt words, the graph database, and the vector storage database, and combining a pre-trained large model to generate a final decision-making suggestion as the personal thinking imitation result for the target user.
[0010] In a second aspect, an embodiment of the present invention provides a personal thinking imitation system, including:
[0011] An acquisition module, which is used to collect multi-source data of a target user in the public domain, private domain, and dynamic questionnaire and perform privacy preprocessing to obtain desensitized behavior data; adopt a three-level segmentation strategy to extract features from the desensitized behavior data, and perform cross-modal association through a graph attention network to obtain a multi-dimensional joint semantic vector; construct a graph database and a vector storage database based on the multi-dimensional joint semantic vector, where the graph database is used to store a five-dimensional knowledge graph in the graph database, and the five-dimensional knowledge graph includes a cognitive dimension, a behavior dimension, an emotional dimension, a social dimension, and an evolutionary dimension; construct dynamic prompt words based on the five-dimensional knowledge graph, a preset five-dimensional weight injection template, and real-time environmental data;
[0012] An execution module, which is used to perform multi-level retrieval enhancement based on the dynamic prompt words, the graph database, and the vector storage database, and combine a pre-trained large model to generate a final decision recommendation as the personal thinking imitation result for the target user.
[0013] Compared with the prior art, the beneficial effects provided by the present invention include: adopting a personal thinking imitation method and system disclosed by the present invention, including: first, collecting multi-source data of a target user in the public domain, private domain, and dynamic questionnaire and performing privacy preprocessing to obtain desensitized behavior data. Then, extracting features through a three-level segmentation strategy, and obtaining a multi-dimensional joint semantic vector through cross-modal association of a graph attention network. Accordingly, a graph database and a vector storage database including a five-dimensional knowledge graph such as cognition are constructed. Combining the five-dimensional knowledge graph, a preset template, and real-time environmental data to construct dynamic prompt words, and after multi-level retrieval enhancement, combining a large model to generate a final decision recommendation for the target user, realizing accurate personal thinking imitation and efficient decision-making assistance, and taking into account privacy protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 It is a schematic flow chart of the steps of the personal thinking imitation method provided by the embodiment of the present invention;
[0016] Figure 2 It is a hierarchical framework diagram of the personal thinking imitation system provided by the embodiment of the present invention;
[0017] Figure 3 It is an interaction schematic diagram of the personal thinking imitation system provided by the embodiment of the present invention;
[0018] Figure 4Structural schematic block diagram of the personal thinking imitation system provided by an embodiment of the present invention;
[0019] Figure 5 Structural schematic block diagram of the computer device provided by an embodiment of the present invention. Detailed implementation manners
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0021] The following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings.
[0022] To solve the technical problems in the foregoing background art, Figure 1 Flow schematic diagram of a personal thinking imitation method provided by an embodiment of the present disclosure. The following will introduce this personal thinking imitation method in detail.
[0023] Step S201: Collect multi-source data of the target user in the public domain, private domain, and dynamic questionnaire and perform privacy preprocessing to obtain desensitized behavior data;
[0024] Step S202: Extract features from the desensitized behavior data by using a three-level segmentation strategy and perform cross-modal association through a graph attention network to obtain a multi-dimensional joint semantic vector;
[0025] Step S203: Construct a graph database and a vector storage database based on the multi-dimensional joint semantic vector. The graph database is used to store a five-dimensional knowledge graph, and the five-dimensional knowledge graph includes a cognitive dimension, a behavior dimension, an emotional dimension, a social dimension, and an evolution dimension;
[0026] Step S204: Construct dynamic prompt words based on the five-dimensional knowledge graph, a preset five-dimensional weight injection template, and real-time environmental data;
[0027] Step S205: Perform multi-level retrieval enhancement based on the dynamic prompt words, the graph database, and the vector storage database, and combine a pre-trained large model to generate a final decision recommendation as the personal thinking imitation result for the target user.
[0028] In an embodiment of the present invention, exemplarily, the local end, as one of the starting points of data collection, undertakes the access work of the user's private domain data and interaction data. Assume that the target user is a scientific researcher, Xiao Li.
[0029] Private domain data collection: Xiao Li uses Obsidian to record scientific research notes on a daily basis. The local end connects to the Obsidian note-taking tool and obtains Xiao Li's note content in real time through relevant technologies. For example, Xiao Li recorded his ideas about a cross-disciplinary study of medicine and engineering in Obsidian, including inductive reasoning on different research methods, deductive reasoning on research directions, and other content. These notes contain rich cognitive dimension information, and the local end will collect these notes as private domain data. At the same time, Xiao Li will also record his project progress and daily thoughts in Notion. The local end also supports the import of Notion data and includes it in the category of private domain data.
[0030] Interaction data collection: The system provides Xiao Li with an embedded dynamic questionnaire system (Typeform customized), which supports voice and gesture input. During an AI training process, the system popped up a dynamic situational choice question in real time: "When conducting scientific research experiments, do you want the system to pay more attention to logical rigor or emotional resonance?" Xiao Li chose to pay more attention to logical rigor through voice, and the local end collected Xiao Li's choice as interaction data. In addition, Xiao Li can also interact with the system through voice and gestures. For example, when explaining his scientific research ideas, he can use gestures to assist in expressing the key points, and the local end will collect these interaction information.
[0031] The edge plays a certain role in transfer and preliminary processing during the data collection process, and is also responsible for collecting part of the public domain data.
[0032] Public domain data collection: Taking Weibo as an example, the edge obtains Xiao Li's public interaction data through the Weibo open platform API. Xiao Li often reposts some public welfare research-related content on Weibo, such as public welfare projects on rare disease research, which reflects his tendency to social responsibility; he also publishes some original science and technology answers to explain his rational decision-making preferences for certain scientific research issues. The edge collects data such as these reposts, comments, likes, and original content. In addition, the edge also integrates public domain data such as GitHub code submission mode and Zhihu answer style. Assuming that Xiao Li often submits code for scientific research algorithm optimization on GitHub at night, this reflects his "deep focus" characteristics in the cognitive dimension. The edge will integrate and collect these cross-platform public domain data.
[0033] Whether it is private domain data collected locally or public domain data collected at the edge, it needs to go through the privacy preprocessing module.
[0034] Local preprocessing: At the local end, for the private data of Xiao Li, such as sensitive information in notes (e.g., specific research experiment details may involve the confidentiality of personal research results), gradient desensitization (ε = 0.4) will be implemented, that is, differential privacy noise will be added to the data so that single data cannot be traced. For sensitive information such as geographical location, generalization processing will be carried out. For example, if Xiao Li occasionally records the location of an academic conference he attended in the note, the system will generalize the specific street address to the city level.
[0035] Preprocessing before cloud synchronization: Before synchronizing the data to the cloud, k-anonymization (k = 50) will be implemented to obfuscate the group characteristics of users. For example, the relevant data of Xiao Li and 49 other users will be mixed so that Xiao Li's individual characteristics cannot be accurately identified from the cloud data, thus protecting Xiao Li's privacy.
[0036] After the above multi-source data collection and privacy preprocessing, desensitized behavior data is obtained, laying a foundation for subsequent steps such as feature extraction.
[0037] A three-level segmentation strategy is adopted to extract features from the desensitized behavior data, and this process is mainly completed by the fractal analysis engine at the feature processing layer.
[0038] Extraction of backbone concepts: Taking a draft of Xiao Li's scientific research paper as a data example, the LayoutLMv3 model is used to analyze the document. The LayoutLMv3 model can identify information such as titles and paragraph structures in the document and extract the backbone concepts from them. For example, the title "Research on a New Algorithm for Medical Image Processing" and the related chapter main ideas in the paper draft are identified as backbone concepts, which reflect Xiao Li's core concerns in this research and belong to important information in the cognitive dimension.
[0039] Identification of supporting evidence: Use SpaCy_NER to process the document and identify the supporting evidence in it. In Xiao Li's paper draft, for the content such as the principle explanation of the new algorithm and the comparison of experimental data, SpaCy_NER can identify it as evidence supporting the backbone concept. These evidences further enrich the understanding of Xiao Li's scientific research thinking and help to comprehensively grasp the characteristics of his cognitive dimension.
[0040] Splitting of meta-knowledge atoms: Use Tesseract_OCR to process the document and split the content in it into meta-knowledge atoms. For example, professional terms, formulas, etc. in the paper draft are all split into the most basic knowledge units. These meta-knowledge atoms are of great significance for constructing a knowledge graph and in-depth analysis of Xiao Li's knowledge system and cognitive structure.
[0041] Cross-modal association is performed on the features extracted by the three-level segmentation strategy through a Graph Attention Network (GAT) to generate a 768-dimensional joint semantic vector.
[0042] Suppose Xiao Li has not only text data such as drafts of scientific research papers, but also his speech videos at academic conferences and related presentation image data.
[0043] Text-image association: In the speech video, Xiao Li showed some example images of medical images, which are related to the medical image processing algorithms mentioned in his paper draft. The Graph Attention Network can identify the association between the description of the algorithm in the text and the image features in the images, such as the correspondence between a certain image feature extraction method mentioned in the text and the specific image manifestation in the image, thus associating the data of these two modalities of text and image.
[0044] Text-speech association: The speech of Xiao Li's explanation of the algorithm during the speech also contains important information. The Graph Attention Network can align the key content conveyed in the speech with the information in the paper draft text, such as matching the advantages of the algorithm emphasized in the speech with the semantics of relevant paragraphs in the text, to achieve the association of the text and speech modalities.
[0045] Through this cross-modal association, a multi-dimensional joint semantic vector is obtained, which more comprehensively reflects Xiao Li's thinking and behavioral characteristics, providing a rich data basis for subsequent steps such as constructing a knowledge graph.
[0046] Construct a Neo4j graph database based on the multi-dimensional joint semantic vector to store a five-dimensional knowledge graph.
[0047] Cognitive dimension storage: Taking Xiao Li as an example, the thinking patterns reflected in his scientific research notes and paper drafts, such as inductive reasoning (summarizing various medical image processing methods), deductive reasoning (deriving the processing method for specific images from general algorithm principles), and cross-domain transfer ability (applying the optimization algorithm of engineering to medical image processing), etc., are stored as nodes in the graph database in the cognitive dimension. The nodes also contain timestamps to record the time when these thinking patterns are reflected, and confidence labels to reflect the reliability of the system's judgment of these features. The nodes are connected by edges, and the edges define the association relationships between different thinking patterns within the cognitive dimension, such as the logical relationship between a certain inductive reasoning method and the subsequent deductive reasoning based on it.
[0048] Behavior Dimension Storage: Public domain behaviors of Xiao Li on Weibo, such as reposting and commenting, as well as behavioral data such as selection preferences in dynamic questionnaires, are stored as nodes in the behavior dimension. For example, the behavior node of his frequent reposting of public welfare scientific research content is connected by an edge to the related node of his social responsibility tendency shown in the dynamic questionnaire, reflecting the internal connection between behaviors. At the same time, the behavior dimension also stores information related to innovative measurement indicators such as value selection entropy and social influence radius.
[0049] Emotion Dimension Storage: Emotion expressions of Xiao Li during academic exchanges, such as anxiety when discussing scientific research problems and joy when achieving scientific research progress, and the emotional polarity information obtained through multimodal semantic parsing, are stored as nodes in the emotion dimension. Information about emotions obtained through methods such as indirect inference of biological signals (such as inferring his possible emotional state of stress release tendency by his late-night high-frequency liking of entertainment content in public domain data) and the relevant content of constructing a three-dimensional mapping model of "emotion - value - physiological response" are also stored in this dimension. Information about measurement indicators such as the entropy value of the emotion network is also recorded in the emotion dimension nodes.
[0050] Social Dimension Storage: The role migration situation of Xiao Li in social networks (such as WeChat Moments, LinkedIn, etc.), for example, his transformation from an information receiver in an academic communication group to a disseminator who can spread valuable scientific research information, and this role migration coefficient is stored as an important indicator in the social dimension node. Information such as the coupling strength between his work collaboration network and professional network is also stored in this dimension. For example, his cooperation relationship with colleagues and connection with industry experts on LinkedIn, and a composite social graph is constructed through cross-platform relationship extraction and stored in a graph database.
[0051] Evolution Dimension Storage: Xiao Li's long-term digital footprint, such as the change in blog writing style over 5 years (from initially simply recording the scientific research process to later deeply analyzing scientific research problems and sharing research insights) and the trajectory of knowledge payment course selection (from basic scientific research method courses to cutting-edge interdisciplinary research courses), are stored as important data in the evolution dimension. Information such as key event markers, such as his career turning points in the scientific research career (promotion from an ordinary scientific researcher to a project leader) and learning breakthrough nodes (mastering a certain key scientific research technology), is also stored in this dimension. Information about breakthrough indicators such as knowledge phase transition threshold and fitness surface is also recorded in the evolution dimension nodes.
[0052] Use the Milvus vector database to store multimodal feature vectors, supporting millisecond-level similarity retrieval (optimized by the Faiss engine).
[0053] Multimodal feature vectors such as the text feature vectors of Xiao Li's scientific research papers, and the image and speech feature vectors in his speech videos are all stored in the Milvus vector database. When the system needs to find similar cases related to Xiao Li's current scientific research problems, it retrieves the vector database through the Faiss engine. For example, when Xiao Li is researching a new medical image denoising algorithm, the system can quickly retrieve other cases in the database regarding image processing algorithms with relatively high similarity of feature vectors, providing him with references. When the similarity > 0.7, it triggers an association suggestion to help Xiao Li make better scientific research decisions.
[0054] Construct dynamic prompt words based on a five-dimensional knowledge graph, a preset five-dimensional weight injection template, and real-time environmental data. Taking Xiao Li as an example, based on the information in his five-dimensional knowledge graph, the system learns that he has strong logical reasoning and cross-domain transfer abilities in the cognitive dimension, a high sense of social responsibility in the behavioral dimension, is relatively rational in the emotional dimension, has a certain social influence in the social dimension, and is in the transition stage from single-discipline research to interdisciplinary research in the evolutionary dimension.
[0055] Suppose the current preset five-dimensional weight injection template is "cognitive weight 0.8 & emotional weight 0.6". According to this template, the system will focus more on the cognitive and emotional dimension features of Xiao Li when generating prompt words.
[0056] Taking the calendar schedule as an example of real-time environmental data. Xiao Li's calendar shows that he has an important scientific research project report in the next week, and the schedule tension is relatively high (schedule_tension > 0.8). Based on this real-time environmental data, the system will make adjustments when constructing dynamic prompt words.
[0057] Combining the above information, the generated dynamic prompt words are as follows:
[0058] ```markdown
[0059] # Cognitive weight 0.8 & emotional weight 0.6
[0060] Generation rules:
[0061] Provide 4 logical comparison schemes to highlight the cost-benefit analysis of different methods in the medical image denoising algorithm, and utilize Xiao Li's logical reasoning ability in the cognitive dimension to help him clearly elaborate on the algorithm advantages in the project report.
[0062] Incorporate 3 historical similar cases, retrieve cases from the vector storage database that are similar to the current scientific research problem and show rational decision-making in the emotional dimension, providing references for Xiao Li's report.
[0063] Limit the emotional intensity to Level 2. Considering Xiao Li's relatively rational characteristics in the emotional dimension and the formal occasion of the project report, avoid overly emotional expressions.
[0064] Such dynamic prompt words can provide targeted guidance for generating decision-making suggestions in the follow-up based on Xiao Li's personal characteristics and real-time environment.
[0065] Multi-level retrieval enhancement and decision-making suggestion generation. The first level: Vector similarity retrieval based on Faiss: According to the requirements in the dynamic prompt words, the system first conducts vector similarity retrieval based on Faiss in the Milvus vector database. Taking Xiao Li's research on medical image denoising algorithms as an example, the system retrieves the top 50 candidate cases with high similarity in feature vectors to the current research problem. These cases may include other similar image processing algorithm studies, related technology applications, etc., providing a preliminary case pool for subsequent decision-making suggestions. The second level: Calculation of path credibility of graph neural networks: For the candidate cases obtained from the first-level retrieval, the system calculates the path credibility in the Neo4j graph database through graph neural networks (optimized with the HNSW algorithm). For example, whether the association paths between the candidate cases and Xiao Li's cognition, behavior, and other dimensions in the five-dimensional knowledge graph are reasonable and can truly provide valuable references for Xiao Li's current research. By calculating the path credibility, more relevant and reliable cases are screened out. The third level: Weighting of user portrait correlation: According to Xiao Li's user portrait, weight the cases screened in the first two levels according to their correlation. Xiao Li has a higher weight in the cognitive dimension (according to the five-dimensional weight injection template and his own characteristics), and the dimensions with a historical adoption rate > 60% are prioritized. The system will give priority to cases that are closely related to Xiao Li's cognitive dimension and have a high adoption rate in the past, further optimizing the retrieval results and providing a more accurate basis for generating decision-making suggestions.
[0066] After multi-level retrieval enhancement, the system combines a pre-trained large model (such as GPT-4) to generate the final decision-making suggestions.
[0067] The system inputs the screened and weighted case information, dynamic prompt words, and other relevant content into the large model. Based on this information and combined with its own language understanding and generation capabilities, the large model generates decision-making suggestions for Xiao Li's research on medical image denoising algorithms. For example, the large model may suggest that Xiao Li adopt a specific algorithm comparison display method in the project report, combined with the successful experience of historical similar cases, to highlight the innovation and practicality of the algorithm; or according to Xiao Li's emotional dimension characteristics, provide suggestions on how to appropriately express emotions in the report to enhance persuasion. These decision-making suggestions, as the results of personal thinking imitation for Xiao Li, can assist him in scientific research decision-making, project reporting, and other work.
[0068] Through the above description, the personal thinking imitation method can start from multi-source data collection, go through a series of complex and delicate processing and analysis, and finally generate decision-making suggestions that fit the personal characteristics and actual needs of the target user, achieving accurate imitation and effective assistance for personal thinking.
[0069] In the embodiment of the present invention, for multi-source data collection of the target user in the public domain, private domain, and dynamic questionnaire and privacy preprocessing to obtain desensitized behavior data, the following examples can be used for implementation.
[0070] Through the open platform API in the public domain, obtain the public interaction data of the target user;
[0071] Through the mind map structure parsing strategy in the private domain, obtain the private data of the target user;
[0072] By outputting a dynamic questionnaire embedded with dynamic scenario multiple-choice questions, obtain the questionnaire feedback of the target user;
[0073] Take the public interaction data, the private data, and the questionnaire feedback as the multi-source data, and perform gradient desensitization and anonymization processing to obtain the desensitized behavior data.
[0074] In the embodiment of the present invention, by way of example, assume that the target user is a workplace person, Xiao Zhang.
[0075] In terms of public domain data collection, the system obtains Xiao Zhang's public interaction data through the open platform API of Weibo. Xiao Zhang often participates in industry topic discussions on Weibo, forwards some articles on workplace skill improvement, and also comments on some hot events. For example, he forwarded an article on "how to improve team collaboration efficiency" and attached his own insights, believing that clear division of labor and effective communication are the keys; he also commented on a major industry cooperation event, expressing his optimism about the cooperation prospects. These public interaction data such as forwards and comments are all obtained by the system and become part of the multi-source data.
[0076] In terms of private domain data collection, Xiao Zhang uses Obsidian to record work notes, and his notes have a mind map structure. The system parses Xiao Zhang's Obsidian notes based on the mind map structure parsing strategy. Xiao Zhang sorted out the planning ideas of a project in the form of a mind map in the notes, including goal setting, task assignment, time node arrangement, etc. The system obtains private data such as Xiao Zhang's thinking mode and decision-making logic during the project planning process by parsing the nodes and connection relationships in the mind map. For example, it can be seen from the mind map that Xiao Zhang considered the advantages and expertise of team members when assigning tasks, reflecting his cognition and behavior tendencies in personnel management.
[0077] For the acquisition of questionnaire feedback, the system outputs a dynamic questionnaire embedded with dynamic scenario multiple-choice questions to Zhang. During a training session on work decisions, the dynamic questionnaire presented the question: "When faced with an urgent and important task, and there is a shortage of resources among team members, would you choose to prioritize allocating external resources of the company or try internal coordination?" After thinking, Zhang chose to try internal coordination, and the system recorded this questionnaire feedback.
[0078] The system uses the publicly available Weibo interaction data, Obsidian note private data, and dynamic questionnaire feedback obtained from Zhang as multi-source data. Subsequently, privacy preprocessing is performed on these data. In the gradient desensitization process, for the content in Zhang's Weibo comments that may involve personal sensitive views, differential privacy noise is added, making it difficult to trace a single comment back to Zhang himself. During the anonymization process, Zhang's data is mixed with the data of a certain number of other users to implement k-anonymization (assuming k = 50), confusing the user group characteristics, and finally obtaining desensitized behavior data. These desensitized behavior data will be used for subsequent operations such as feature extraction to achieve the imitation and analysis of Zhang's personal thinking.
[0079] In the embodiment of the present invention, the desensitized behavior data is subjected to feature extraction using a three-level segmentation strategy and cross-modal association is performed through a graph attention network to obtain a multi-dimensional joint semantic vector, which can be implemented through the following examples.
[0080] Successively perform backbone concept extraction on the desensitized behavior data based on LayoutLMv3, identify supporting arguments for the desensitized behavior data based on SpaCy_NER, and perform meta-knowledge atom splitting on the desensitized behavior data based on Tesseract_OCR to obtain multiple features to be processed obtained based on the three-level segmentation strategy;
[0081] Adjust and align the multiple features to be processed through a cross-modal attention mechanism;
[0082] Use the graph attention network to perform cross-modal association on the aligned multiple features to be processed to obtain the multi-dimensional joint semantic vector. The cross-modal association includes vocabulary-level association based on FastText cross-training, entity-level association based on the TransEdge algorithm, attribute-level association based on the graph attention network, instance-level association based on adversarial training, and rule-level association based on SWRL reasoning.
[0083] In the embodiment of the present invention, by way of example, continue to take the workplace person Zhang as an example.
[0084] When extracting features for the three - level segmentation strategy, first, based on LayoutLMv3, the desensitized behavior data of Xiao Zhang is used for backbone concept extraction. The desensitized behavior data of Xiao Zhang includes a project summary report document he wrote. The LayoutLMv3 model analyzes this document and identifies the title "Quarterly Project Summary and Future Plan" and the main ideas of each chapter, such as "Overview of Project Results", "Problems Encountered and Solutions", "Future Work Plan", etc., which are determined as backbone concepts.
[0085] Next, based on SpaCy_NER, the supporting evidence in Xiao Zhang's desensitized behavior data is identified. In the "Overview of Project Results" chapter, Xiao Zhang mentioned that "by optimizing the process, the project delivery time was shortened by 20% and the cost was reduced by 15%". SpaCy_NER can identify contents such as "optimizing the process", "delivery time shortened by 20%", and "cost reduced by 15%" as evidence to support the backbone concept of "remarkable project results".
[0086] Then, based on Tesseract_OCR, the meta - knowledge atoms in the desensitized behavior data are split. There are some professional terms in Xiao Zhang's project summary report, such as "agile development", "KPI assessment", etc. Tesseract_OCR splits these terms, as well as the formulas and specific abbreviations in the report, into meta - knowledge atoms. Thus, multiple features to be processed obtained based on the three - level segmentation strategy are obtained.
[0087] After that, these multiple features to be processed are adjusted and aligned through a cross - modal attention mechanism. In addition to the text form of Xiao Zhang's project summary report, he also made relevant presentation slides containing image information and gave an oral report at the team meeting with voice information. The cross - modal attention mechanism aligns the backbone concepts, supporting evidence, etc. in the text with the chart data in the image and the key points emphasized in the voice to ensure semantic consistency of information in different modalities.
[0088] Finally, a graph attention network is used for cross-modal association. At the lexical level, based on cross-training with FastText, professional terms in the text are associated with the same words mentioned in the speech. For example, "agile development" appears in both the text and the speech, and a connection is established between them. At the entity level, based on the TransEdge algorithm, the representations of specific entities in the project, such as team member names and partner company names, in different modalities are associated. At the attribute level, a graph attention network is used to analyze the relationship between the project attributes described in the text (such as delivery time, cost, etc.) and the attributes shown in the charts in the images. At the instance level, through adversarial training, the information about specific project instances in different modalities is associated. For example, the correspondence between the execution of a specific task in the text description and the image display. At the rule level, based on SWRL reasoning, the manifestations of some rules in project management in different modality data are deduced, such as the consistency of performance appraisal rules in the KPI assessment description in the text and the speech explanation. Through this series of operations, multi-dimensional joint semantic vectors are obtained, providing a rich and closely related data basis for subsequent construction of knowledge graphs, etc.
[0089] In the embodiments provided by the present invention, the following implementation manners are also provided.
[0090] When the data deviation of at least three dimensions of data in the five-dimensional knowledge graph, including the cognitive dimension, the behavioral dimension, the emotional dimension, the social dimension, and the evolutionary dimension, exceeds a preset deviation threshold, a multi-modal verification process is triggered, and the five-dimensional knowledge graph is adjusted according to the verification result;
[0091] The cognitive dimension uses an LSTM-GNN hybrid model to capture the long-term cognitive transition trajectory; the behavioral dimension develops a cognitive-behavior consistency verification algorithm, and triggers a double-verification process when the deviation between the questionnaire statement and the public domain behavior exceeds a preset deviation threshold; the emotional dimension develops an emotional state transition matrix and combines a seasonal decomposition algorithm to identify periodic patterns; the social dimension introduces a dynamic community discovery algorithm and applies the preferential attachment model in complex network theory to predict the change trend of social influence within a preset time range; the evolutionary dimension introduces complex system theory, establishes a cognitive attractor model to explain the non-linear growth law, and simulates the cognitive evolution path in different environments through the Markov chain Monte Carlo method.
[0092] In the embodiments of the present invention, by way of example, continue to take the workplace person Xiao Zhang as an example for detailed description.
[0093] When monitoring Zhang's five-dimensional knowledge graph, if it is found that the data deviation in at least three of the cognitive dimension, behavioral dimension, emotional dimension, social dimension, and evolutionary dimension exceeds the preset deviation threshold (assuming the preset deviation threshold is 20%), a multi-modal verification process will be triggered. For example, the recent thinking mode change data in Zhang's cognitive dimension, the selection preference data in the behavioral dimension, and the social relationship change data in the social dimension, after calculation, the deviation exceeds 20%. Immediately start the multi-modal verification process, comprehensively verify and analyze multi-modal information such as Zhang's text data and voice interaction data, and then adjust the five-dimensional knowledge graph according to the verification results. In the cognitive dimension, an LSTM-GNN hybrid model is used to capture Zhang's long-term cognitive transition trajectory. When Zhang first entered the workplace, he relied more on experience-driven in project planning. As the working years increased and he participated in various trainings and projects, he gradually changed to a data-driven thinking mode, and later began to have a systematic thinking. The LSTM-GNN hybrid model can analyze data such as Zhang's work documents and decision records at different stages, capture his phased leap trajectory from experience-driven to data-driven to systematic thinking, and record it in the cognitive dimension of the five-dimensional knowledge graph. For the behavioral dimension, a cognitive-behavioral consistency verification algorithm has been developed. Zhang stated in a dynamic questionnaire that he would choose efficiency first in case of time conflicts. However, by analyzing his public domain behavior data (such as in the work routine shared on Weibo, sacrificing efficiency many times to ensure work quality), it is found that the deviation between the questionnaire statement and the public domain behavior exceeds the preset deviation threshold (assuming 25%). At this time, a double verification process is triggered, and further combined with Zhang's other behavior data and interaction records for re-judgment and verification to ensure the accuracy of the behavioral dimension data. In the emotional dimension, an emotional state transition matrix has been developed, and combined with the seasonal decomposition algorithm to identify periodic patterns. At the end of each quarter of the year, through the behavior patterns in the public domain data (such as frequently posting content related to anxiety emotions on social media and giving high-frequency likes to entertainment content late at night) and multi-modal semantic parsing (analyzing the emotional polarity of the dynamics he posted), Zhang uses the emotional state transition matrix and the seasonal decomposition algorithm to identify the periodic pattern of the increase in his anxiety index at the end of the quarter and record it in the emotional dimension. In terms of the social dimension, a dynamic community discovery algorithm is introduced, and the preferential attachment model in complex network theory is applied to predict the change trend of social influence within a preset time range (such as the next 6 months). In professional social networks such as LinkedIn, through the dynamic community discovery algorithm, it is analyzed that the social circle Zhang is in is changing, and it is found that he is gradually integrating from a smaller technical communication community into a larger industry elite communication community. At the same time, using the preferential attachment model, combined with his social behavior data (such as the number of contacts added and the interaction volume of the content posted), it is predicted that his social influence may gradually increase in the next 6 months, and relevant prediction information is recorded in the social dimension. In the evolutionary dimension, the complex system theory is introduced, and a cognitive attractor model is established to explain Zhang's non-linear growth law.After Zhang encountered a major industry transformation event (such as the emergence of new technologies), his thinking mode and behavior habits changed significantly. By analyzing indicators such as the change rate of the standard deviation of his behavior before and after, and combining with the cognitive attractor model to explain this non-linear growth. At the same time, the cognitive evolution path of Zhang in different industry environments is simulated by the Markov chain Monte Carlo method, providing a basis for more accurately depicting his growth trajectory and recording it in the evolution dimension.
[0094] In the embodiment provided by the present invention, the following embodiments are also provided.
[0095] The five-dimensional knowledge graph dynamically adjusts the weights corresponding to the cognitive dimension, the behavior dimension, the emotion dimension, the social dimension, and the evolution dimension through a five-dimensional weight matrix, and dynamically adjusts the emphasis of the recommendation dimension in combination with the real-time environmental data;
[0096] The five-dimensional weight matrix passes the formula: W t+1 =αW t +β(ΔF + γΔC) to achieve dynamic adjustment, where W t+1 is the adjusted weight, W t is the weight before adjustment, α is the historical decay coefficient, β is the feedback gain, γ is the cross-dimensional coupling factor, ΔF is the feedback weight adjustment factor, and ΔC is the dimension weight adjustment factor.
[0097] In the embodiment of the present invention, by way of example, still taking Zhang, a workplace person, as an example.
[0098] Zhang will have various interactions with the system in his daily work. The system constructs a five-dimensional knowledge graph about Zhang based on these interaction data, and dynamically adjusts the weights corresponding to the five dimensions of cognition, behavior, emotion, society, and evolution through a five-dimensional weight matrix, and dynamically adjusts the emphasis of the recommendation dimension in combination with the real-time environmental data. First, the five-dimensional weight matrix realizes dynamic adjustment according to the formula W t+1 =αW t +β(ΔF + γΔC). Suppose at a certain moment t, the weight Wt(cognition) of the cognitive dimension in Zhang's five-dimensional weight matrix is 0.7, the historical decay coefficient α is set to 0.85, the feedback gain β (explicit feedback) is 1.5, β (implicit feedback) is 0.8, and the cross-dimensional coupling factor γ takes dynamic values according to the correlation between dimensions. In terms of explicit feedback, when Zhang uses the system to generate decision-making suggestions, he finds that the solution given by the system is not perfect enough in logical reasoning. He actively adjusts the content generated by the AI and emphasizes more on logical rigor in the rewritten prompt words, which belongs to explicit feedback. At this time, the feedback weight adjustment factor ΔF(explicit) is 1. Since the cognitive dimension is closely related to logical reasoning, the system calculates the adjusted weight W t+1 (cognition): W t+1(Cognition) = 0.85 × 0.7 + 1.5 × 1 = 0.595 + 1.5 = 2.095. Then, it is restricted between 0.1 and 1.0 through the np.clip function to finally obtain a suitable adjusted weight. In terms of implicit feedback, the system previously predicted that Zhang would consult friends when dealing with a work project (prediction based on the social dimension), but in fact, Zhang chose to make an independent decision, continuously deviating from the predicted path, which triggered implicit feedback. At this time, the feedback weight adjustment factor ΔF (implicit) is 0.6, and the system also adjusts the weight of the social dimension according to the above formula. At the same time, the dimension weight adjustment factor ΔC will consider the coupling relationship between dimensions. For example, the server analyzes and finds that there is a certain antagonistic effect between Zhang's cognitive dimension and emotional dimension (the correlation coefficient is 0.32). When the weight of the cognitive dimension is adjusted, ΔC will take this cross-dimensional influence into account to comprehensively adjust the weights of each dimension. When dynamically adjusting the dimension focus of the recommendation in combination with real-time environmental data, assume that Zhang's calendar shows that he will have an important project report in a week. This real-time environmental data indicates that the current is a relatively tense and important work stage. The server dynamically adjusts the dimension focus of the decision recommendation according to the adjusted weights of each dimension in the five-dimensional weight matrix and the current real-time environmental data. Since Zhang has a relatively high weight in the cognitive dimension, and the project report requires strong logical thinking and clear expression, the server will focus more on the cognitive dimension when generating decision recommendations, such as providing more suggestions on the construction of the logical framework of the project report and data argument analysis; at the same time, it will also appropriately consider other dimensions. For example, in the social dimension, it will give suggestions on how to communicate better with team members and superiors during the report to assist Zhang in better completing the project report.
[0099] In the embodiments provided by the present invention, the following implementation manners are also provided.
[0100] Obtain the desensitized behavior data on the local side based on the AES-256 encryption algorithm;
[0101] Deploy the graph attention network on the local side for cross-modal association, and combine differential privacy noise to obtain a multi-dimensional joint semantic vector;
[0102] Synchronize the multi-dimensional joint semantic vector to the cloud based on anonymization processing; wherein, when synchronizing the multi-dimensional joint semantic vector between the local side and the cloud, only upload the dimension adjustment parameters;
[0103] Implement parameter aggregation and update for the multi-dimensional joint semantic vector through the FedAvg algorithm on the cloud;
[0104] Deploy a parameter mapping gateway on the cloud, convert the multi-dimensional joint semantic vector into a large model adaptation format, and access the pre-trained large model.
[0105] In an embodiment of the present invention, by way of example, continue to take Xiao Zhang, a working professional, as an example.
[0106] Xiao Zhang uses the system on his personal device (local end). The system obtains his public Weibo interaction data through the open platform API in the public domain, obtains his private Obsidian note data through the mind map structure analysis strategy in the private domain, and obtains his questionnaire feedback through a dynamic questionnaire that outputs embedded dynamic scenario multiple-choice questions. After obtaining this multi-source data, the desensitized behavior data is encrypted on the local end based on the AES256 encryption algorithm. For example, Xiao Zhang's Weibo interaction data contains some work-related discussion content of his, and the private notes record project planning details, etc. These data are encrypted and stored on the local end to ensure data security and prevent the data from being illegally obtained and read during local storage.
[0107] A graph attention network is deployed on the local end to perform cross-modal association on the desensitized behavior data. Xiao Zhang's desensitized behavior data includes multi-modal data such as work report documents in text form, project charts in image form, and meeting speech records in voice form. The graph attention network processes this multi-modal data, and performs correlation analysis on the key information in the text, the data features in the image, and the key content in the voice. At the same time, combined with differential privacy noise (for example, setting ε = 0.4), data privacy is further protected to obtain a multi-dimensional joint semantic vector. For example, when analyzing Xiao Zhang's project report, the graph attention network correlates the description of the project objectives in the document, the project progress data shown in the chart, and the key tasks emphasized in the meeting speech to generate a multi-dimensional joint semantic vector that can comprehensively reflect Xiao Zhang's thinking and behavior characteristics in this project.
[0108] Based on anonymization processing (such as implementing k-anonymization, k = 50), the multi-dimensional joint semantic vector is synchronized to the cloud. When synchronizing the multi-dimensional joint semantic vector between the local end and the cloud, only the dimension adjustment parameters are uploaded. After Xiao Zhang's multi-dimensional joint semantic vector is anonymized on the local end, the information that can directly identify Xiao Zhang's identity is removed. Then, the system only uploads the dimension adjustment parameters to the cloud, such as the adjustment values of the weights of each dimension in the five-dimensional weight matrix, etc., without transmitting the original multi-dimensional joint semantic vector, further protecting privacy while reducing the data transmission volume.
[0109] Implement parameter aggregation and update for multi-dimensional joint semantic vectors through the FedAvg algorithm in the cloud. The cloud will receive dimension adjustment parameters from numerous users like Zhang. The FedAvg algorithm aggregates and calculates these parameters. For example, the cloud collects dimension adjustment parameters of multiple professionals in the same industry, which reflect their different characteristics and changing trends at work. Through the FedAvg algorithm, these parameters are integrated to update the model parameters in the cloud, enabling the model to better adapt to different users' situations and improving the generality and accuracy of the model.
[0110] Deploy a parameter mapping gateway in the cloud to convert multi-dimensional joint semantic vectors into a format suitable for large models and connect to a pre-trained large model. After the multi-dimensional joint semantic vectors of Zhang are processed by the cloud, the parameter mapping gateway converts them into a format that large models such as GPT4 can understand and process. For example, convert the multi-dimensional joint semantic vectors related to Zhang's thinking and behavior characteristics at work into a vector representation form that meets the input requirements of the large model, and then connect to the large model. Based on this input information, the large model combines its own knowledge and algorithms to generate more accurate and personalized decision-making suggestions for Zhang to assist Zhang in making better decisions in the workplace.
[0111] In the embodiment of the present invention, the multi-level retrieval enhancement is performed based on the dynamic prompt words, the graph database, and the vector storage database, and the final decision-making suggestion is generated in combination with the pre-trained large model as the personal thinking imitation result for the target user, which can be implemented through the following examples.
[0112] Based on the dynamic prompt words, retrieve a preset number of first-level multi-dimensional joint vectors from the vector storage database through Faiss vector retrieval;
[0113] Calculate the HNSW path credibility based on the five-dimensional knowledge graph stored in the graph database, and screen out second-level multi-dimensional joint vectors from the first-level multi-dimensional joint vectors;
[0114] Weight the second-level multi-dimensional joint vectors based on the user portrait correlation degree determined by the historical adoption rate to determine the third-level multi-dimensional joint vectors;
[0115] Execute the pre-trained large model based on the dynamic prompt words and the third-level multi-dimensional joint vectors to generate the final decision-making suggestion as the personal thinking imitation result for the target user.
[0116] In the embodiment of the present invention, by way of example, still taking the professional Zhang as an example.
[0117] The server obtains the dynamically generated prompt for Xiao Zhang, such as "Cognitive weight 0.8 & Emotional weight 0.6, generation rule: Provide 4 logical comparison schemes, incorporate 3 historical similar cases, and limit the emotional intensity to Level 2". Based on this dynamic prompt, the server conducts Faiss vector retrieval from the vector storage database (such as the Milvus vector database) to find multi-dimensional joint vectors related to Xiao Zhang's current work scenario or problem. Assuming the preset quantity is 50, the server retrieves 50 first-level multi-dimensional joint vectors with relatively high similarity in dimensions such as cognition and emotion involved in the dynamic prompt. For example, Xiao Zhang is conducting a market promotion plan for a new product. These first-level multi-dimensional joint vectors may include multi-modal feature vectors of previous similar product promotion cases, such as text feature vectors of relevant documents, image feature vectors during the planning process, and voice feature vectors of discussion meetings, etc.
[0118] Next, the server calculates the HNSW path credibility based on the five-dimensional knowledge graph of Xiao Zhang stored in the graph database (Neo4j graph database). The five-dimensional knowledge graph records various characteristics and their interrelationships of Xiao Zhang in the dimensions of cognition, behavior, emotion, society, and evolution. For the above 50 first-level multi-dimensional joint vectors, the server analyzes their associated paths with the nodes and edges in Xiao Zhang's five-dimensional knowledge graph. For example, in a case corresponding to a first-level multi-dimensional joint vector, whether the thinking mode involved in its decision-making process is similar to Xiao Zhang's thinking chain in the cognitive dimension, and whether the behavior pattern conforms to the value selection entropy and social influence radius characteristics of Xiao Zhang in the behavior dimension, etc. Through the HNSW path credibility calculation, second-level multi-dimensional joint vectors with higher credibility are selected. Assuming 20 remain after screening.
[0119] The server weights the second-level multi-dimensional joint vectors based on the user portrait correlation degree determined by the historical adoption rate. When Xiao Zhang used the decision-making suggestions generated by the system in the past, the adoption situations of suggestions in different dimensions were different. The system constructs the user portrait correlation degree based on these historical adoption rates. For example, Xiao Zhang has a relatively high adoption rate of logical analysis type suggestions based on the cognitive dimension, reaching 70%, while the adoption rate of some emotional expression type suggestions in the emotional dimension is relatively low, at 30%. Then when weighting the second-level multi-dimensional joint vectors, vectors closely associated with the cognitive dimension will be given higher weights, and vice versa. After the weighting process, third-level multi-dimensional joint vectors are determined.
[0120] Finally, the server executes a pre-trained large model (such as GPT4) based on the dynamic prompt words and the three-level multi-dimensional joint vector. The large model combines the specific requirements in the dynamic prompt words, such as providing a logical comparison scheme, integrating historical similar cases, etc., and the relevant feature information of Xiaozhang and the similar case information contained in the three-level multi-dimensional joint vector to generate final decision-making suggestions. For example, the large model may provide a detailed logical comparison scheme for Xiaozhang's new product market promotion plan, including the cost-benefit analysis of different promotion channels, the precise positioning strategy of the target customer group, etc., and integrate the experience of 3 historical similar and successful promotion cases screened from the three-level multi-dimensional joint vector. At the same time, the emotional intensity is controlled at Level2 to meet the requirements of the dynamic prompt words. These decision-making suggestions, as the result of personal thinking imitation for Xiaozhang, assist Xiaozhang in making more scientific and reasonable market promotion decisions.
[0121] In the embodiment of the present invention, the construction of the five-dimensional knowledge graph adopts an incremental learning mechanism, starts sub-graph update at a preset cycle node, processes data changes within the preset data valid range, and triggers full-graph rebalancing when the change impact degree exceeds the preset impact degree threshold;
[0122] The five-dimensional knowledge graph performs exponential decay on nodes that have not been activated for a preset duration, and archives them to the historical database when the confidence level is lower than the preset decay confidence level.
[0123] In an embodiment of the present invention, by way of example, take the workplace person Zhang as an example. The five-dimensional knowledge graph constructed by the system for Zhang adopts an incremental learning mechanism, and the preset cycle node is set to every early morning. Every day, Zhang generates new work-related data, such as his chat records in the work group, newly completed project documents, records of online training courses he participated in, etc. These data will be reflected in different dimensions of the five-dimensional knowledge graph. In the early morning of each day, the server starts the sub-graph update program to process data changes within the preset data valid range. Assuming the preset data valid range is the most recent 24 hours, the server will filter out the new data generated by Zhang in the past 24 hours. For example, Zhang discussed a new project plan with team members yesterday. His speech content and viewpoints during the discussion involve the cognitive dimension (showing thinking patterns and decision-making logics) and the social dimension (interaction relationships with team members). The server adds the corresponding nodes and edges of these new data to the corresponding sub-graph of the five-dimensional knowledge graph. At the same time, the server calculates the impact degree of the data change. If Zhang proposed a completely new project execution idea during yesterday's discussion, which is quite different from the previous thinking patterns, this may have a greater impact on the relevant nodes and edges in the cognitive dimension. The system will evaluate the impact degree of this change. When the change impact degree exceeds the preset impact degree threshold (assumed to be 0.3), a full-graph rebalancing is triggered. At this time, the server recalculates the association relationships between the nodes in each dimension of the five-dimensional knowledge graph, adjusts the weights of the edges, etc., to ensure that the knowledge graph can accurately reflect Zhang's latest status and characteristics. The five-dimensional knowledge graph also processes nodes that have not been activated for a preset duration. The preset duration is set to 6 months, that is, if a certain node has not been accessed or updated within 6 months, it is regarded as an unactivated node. For example, Zhang participated in a small project half a year ago. At that time, some behavior characteristic nodes of his in this project were recorded in the behavior dimension of the five-dimensional knowledge graph. But later, due to the shift of his work focus, he no longer involved relevant content, and this node was in an unactivated state. The system performs exponential decay on such unactivated nodes. The calculation formula is confidence = initial value × 0.9^t (t is the time, in months). As time goes by, the confidence of this node continuously decreases. When the confidence is lower than the preset decay confidence (assumed to be 0.2), the system archives this node to the historical database. This can avoid accumulating too much invalid or outdated information in the five-dimensional knowledge graph, improve the storage and query efficiency of the knowledge graph, and at the same time ensure that the current knowledge graph focuses on Zhang's recent main behaviors and characteristics, so as to provide him with more accurate decision-making suggestions and personal thinking imitation.
[0124] In an embodiment of the present invention, the graph database uses Neo4j to store the five-dimensional knowledge graph. The nodes include decision-making cases, psychological characteristics, environmental context, timestamps, and confidence labels. The edges define the cross-dimensional influence coefficients;
[0125] The vector storage database uses Milvus to store the multi-dimensional joint semantic vectors and is optimized by the Faiss engine. When the node similarity exceeds the preset similarity threshold, an association recommendation is triggered.
[0126] In an embodiment of the present invention, by way of example, take Xiao Zhang, a working professional, as an example.
[0127] The system uses Neo4j to store Xiao Zhang's five-dimensional knowledge graph. In terms of nodes: Decision cases: Xiao Zhang has participated in many project decisions at work. For example, in a product optimization project, he proposed a decision-making plan to understand user needs through market research and then determine the direction of product function improvement. This decision case will be stored as a node in the knowledge graph, and details such as the background, process, and results of the decision will be recorded. Psychological characteristics: Through the analysis of multi-source data such as Xiao Zhang's dynamic questionnaire feedback and daily communication, the system determines that Xiao Zhang has strong stress resistance when facing pressure and is relatively rational when making decisions. These psychological characteristic information will also be stored as nodes to reflect Xiao Zhang's inner traits. Environmental context: The company where Xiao Zhang works is currently in a stage of business expansion, and the market competition is fierce. This kind of company environment and market environment context information will be recorded as nodes in the knowledge graph because environmental factors will affect his decisions and behaviors. Timestamp and confidence level label: Each of the above nodes is tagged with a timestamp. For example, the timestamp of the product optimization project decision case records the specific time when the decision occurred; at the same time, the system will assign a confidence level label to the node according to the data source and the reliability of the analysis. For example, for the psychological characteristic node that Xiao Zhang has strong stress resistance obtained based on a large amount of reliable data, a relatively high confidence level label is assigned.
[0128] In terms of the definition of edges, edges are used to define the cross-dimensional influence coefficient. For example, Xiao Zhang's rational thinking mode in the cognitive dimension (cognitive dimension node) will affect his decision-making behavior in the behavioral dimension (behavioral dimension node). The two are connected by an edge, and the edge will define a cross-dimensional influence coefficient to represent the degree and direction of this influence. If Xiao Zhang's rational thinking makes him more inclined to analyze data and risks when making decisions, then this influence coefficient will reflect the tightness of this association.
[0129] Milvus is used to store the multi-dimensional joint semantic vectors of Zhang. After the feature extraction and cross-modal association of Zhang's work-related data, such as project report texts, meeting speech, presentation images, etc., multi-dimensional joint semantic vectors are generated and stored in Milvus. The Faiss engine optimizes Milvus to improve the efficiency and accuracy of vector retrieval. When the node similarity exceeds the preset similarity threshold (assumed to be 0.7), an association recommendation is triggered. For example, when Zhang is preparing a new marketing plan, the system will calculate the similarity between the multi-dimensional joint semantic vectors related to marketing and the vectors in the database. If it is found that the similarity between a stored vector (corresponding to a previous successful marketing case) and the current vector reaches 0.75, exceeding the preset threshold, the system will trigger an association recommendation and display relevant information about the previous successful case, such as marketing strategies, market feedback, etc., for his reference. This helps Zhang draw on past experience and make more reasonable decisions, and also reflects the important role of the vector storage database in assisting personal thinking imitation and decision support.
[0130] In order to more clearly describe the solution provided by the embodiments of the present invention, a relatively complete implementation manner is provided below.
[0131] 1. Dimension Reconstruction and Theoretical Deepening of the Five-Dimensional Knowledge Graph Analysis Method
[0132] Proposed background: The limitations of traditional user portraits (single dimension, static analysis), combined with the integrated innovation of psychology, behavioral science, and knowledge graph technology.
[0133] 1.1 Cognitive Dimension: Dynamic Modeling of Thinking Patterns
[0134] 1.1.1 Core Definition:
[0135] Reveal the core decision-making logic and information processing paradigm of users, including inductive reasoning, deductive reasoning, cross-domain transfer ability, etc. Different from traditional cognitive models, this dimension emphasizes capturing the depth and temporal evolution law of the thinking chain.
[0136] 1.1.2 Measurement System:
[0137] Logical Chain Density: Calculated by the number of consecutive follow-up questions and the number of decision path branches in the log (for example, if the number of consecutive follow-up questions ≥ 3 and the number of path bifurcations ≥ 2, it is regarded as deep cognition);
[0138] Cross-Domain Association Strength: Based on the co-occurrence frequency and semantic similarity of cross-disciplinary nodes in the knowledge graph (for example, if the co-occurrence rate of medical and engineering concepts > 30% is regarded as a strong association).
[0139] 1.1.3 Innovation in Data Sources:
[0140] Private domain data: Mind map structure analysis (such as the bidirectional link density in Obsidian notes), e-book annotation mode (highlighted paragraph topic clustering);
[0141] Public domain data: Academic paper citation network analysis (obtaining interdisciplinary features through the SemanticScholar API).
[0142] 1.1.4 Dynamic evolution mechanism:
[0143] Adopt an LSTM-GNN hybrid model to capture the long-term cognitive leap trajectory (such as the phased leap from experience-driven to data-driven to systematic thinking).
[0144] 1.2 Behavioral dimension: Privacy-friendly behavior modeling
[0145] In response to privacy protection requirements, establish a dual-track analysis system based on choice preferences and public domain footprints, breaking through the traditional dependence on device sensors.
[0146] 1.2.1 Innovative measurement indicators:
[0147] Value selection entropy: Quantify behavioral inertia through virtual scenario questionnaires (such as "When there is a time conflict, do you choose efficiency first / quality first?");
[0148] Social influence radius: Calculate based on the forwarding level of the social network (such as WeChat Moments) and the connection strength of the professional network (such as LinkedIn) (a three-level forwarding depth + 50% strong connection ratio is considered high influence).
[0149] 1.2.2 Data collection strategy:
[0150] Enhanced questionnaire design: Embed dynamic scenario multiple-choice questions (such as "Do you want the system to pay more attention to logical rigor / empathic resonance?" when training AI);
[0151] Analysis of public domain footprints: Obtain the topic participation matrix through public domain APIs (such as Weibo, Xiaohongshu, Twitter) (original content ratio > 40% is considered an active output behavior).
[0152] 1.2.3 Contradiction handling mechanism:
[0153] Develop a cognitive-behavioral consistency verification algorithm to trigger a double-verification process when the deviation between the questionnaire statement and public domain behavior > 25%.
[0154] 1.3 Emotional dimension: Multimodal emotion map
[0155] 1.3.1 Theoretical breakthrough:
[0156] Integrate knowledge graph emotion computing technology to construct a three-dimensional mapping model of "emotion-value-physiological response".
[0157] 1.3.2 Measurement Innovation:
[0158] Emotional Network Entropy Value: Calculate the distribution dispersion of text emotional word vectors in the knowledge graph (a standard deviation > 0.5 is regarded as high volatility);
[0159] Value Tendency Index: Quantify egoistic / altruistic preferences based on the choice experiment method (the conflict intensity between donation behavior and self - interest options > level 3).
[0160] 1.3.3 Data Source Expansion:
[0161] Indirect Inference of Biological Signals: Infer the emotional state through the behavior patterns in public domain data (such as high - frequency liking of entertainment content late at night - tendency of stress release);
[0162] Multi - modal Semantic Parsing: Construct compound emotion labels by analyzing the emotional polarity of user - released dynamics.
[0163] 1.3.4 Dynamic Modeling:
[0164] Develop an emotional state transition matrix and combine it with the seasonal decomposition algorithm to identify periodic patterns (such as the upward trend of anxiety index at the end of the quarter).
[0165] 1.4 Social Dimension: Reconstruction Analysis of the Relationship Network
[0166] 1.4.1 Methodology Innovation:
[0167] Introduce a dynamic community discovery algorithm to break through the limitations of traditional static social graphs.
[0168] 1.4.2 Core Indicators:
[0169] Role Migration Coefficient: Calculate the centrality change rate in the social network within half a year (such as the transformation speed from information receiver to disseminator);
[0170] Cross - platform Collaboration Degree: For example, evaluate the coupling intensity between the work collaboration network and the professional network (a common node ratio > 15% is regarded as high collaboration).
[0171] 1.4.3 Data Fusion Strategy:
[0172] Cross - platform Relationship Extraction: Construct a compound social graph through the social network data provided by users;
[0173] Latent Relationship Mining: Identify unstated master - apprentice relationships, potential cooperation intentions, etc. through knowledge graph reasoning technology.
[0174] 1.4.4 Evolution Prediction:
[0175] Apply the preferential attachment model in complex network theory to predict the changing trend of social influence in the next 6 months.
[0176] 1.5 Evolution Dimension: Phase Transition Analysis of Cognitive Growth
[0177] Introduce complex system theory and establish a cognitive attractor model to explain the non-linear growth law.
[0178] 1.5.1 Breakthrough Indicators:
[0179] Knowledge Phase Transition Threshold: Identify the key stimulus intensity for the mutation of thinking patterns (e.g., the change rate of behavioral standard deviation before and after major events > 40%);
[0180] Fitness Landscape: Simulate the cognitive evolution path in different environments through the Markov chain Monte Carlo method.
[0181] 1.5.2 Data Source Innovation:
[0182] Long-term Digital Footprint: Integrate the changes in blog writing styles and the selection trajectories of knowledge payment courses over 5 years;
[0183] Key Event Marking: Identify career turning points and learning breakthrough nodes through calendar data.
[0184] 1.5.3 Dynamic Modeling:
[0185] Construct a cognitive fitness landscape model to predict the evolution direction of thinking patterns in the next 12 - 18 months.
[0186] 1.6 Five-Dimension Synergy Mechanism
[0187] 1.6.1 Cross-Dimension Coupling Matrix:
[0188] Construct a five-dimensional correlation intensity map to identify the synergy / antagonistic effects between dimensions (e.g., a high social dimension may inhibit the intensity of emotional expression, and the correlation coefficient reaches -0.32).
[0189] 1.6.2 Contradiction Early Warning System:
[0190] When the data deviation of three or more dimensions > 20%, trigger a multi-modal verification process (e.g., combine voice emotion analysis and questionnaire results for cross-verification).
[0191] 1.6.3 Dynamic Weight Allocation:
[0192] Design a reinforcement learning framework to automatically adjust the dimension weights according to the user's adoption rate of AI suggestions. Update formula:
[0193] ```python
[0194] W t+1 =αWt +β(ΔF + γΔC), where α = historical decay coefficient, β = feedback gain, γ = cross - dimensional coupling factor
[0195] ```
[0196] 1.7 Theoretical innovation value of the five - dimensional knowledge graph analysis method:
[0197] Initiate the "cognitive attractor" model, breaking through the static analysis framework of traditional psychology;
[0198] Establish a privacy - friendly behavior analysis paradigm to solve the ethical dilemma of device data collection;
[0199] Achieve the deep coupling of knowledge graph technology and complex system theory.
[0200] 2. Data - driven five - dimensional dynamic modeling and knowledge base construction
[0201] 2.1 Data collection and feature engineering optimization
[0202] 2.1.1 Multi - source data fusion strategy
[0203] 2.1.1.1 Public domain behavior analysis:
[0204] Public domain social network data collection: Obtain users' public interaction data (reposts, comments, likes) through API interfaces (such as Weibo Open Platform) to construct a "behavior - value" association model. For example, high - frequency reposting of public welfare content reflects the tendency of social responsibility, and original science and technology - related answers reflect rational decision - making preferences.
[0205] Cross - platform data integration: Integrate public domain data such as GitHub code submission patterns and Zhihu answer styles to form a composite behavior portrait (such as continuous late - night code submissions - cognitive dimension "deep focus" feature + weight).
[0206] 2.1.1.2 Enhanced questionnaire design: Develop a dynamic scenario multiple - choice question bank, embed five - dimensional guiding questions (such as "Do you rely more on intuition / logic when making decisions"), and use contrastive learning technology to align questionnaire self - reports with public domain behavior data. Trigger a double - verification process when the deviation > 25%.
[0207] 2.1.2 Privacy - friendly feature extraction
[0208] Text semantic parsing: Use the BERT model for sentiment word vector analysis, and combine with the LDA topic model to mine implicit behavior features in users' original content (such as systematic thinking tendency in Zhihu answers).
[0209] Synthetic data generation: For sensitive data (such as medical records), use generative adversarial networks (GANs) to generate synthetic data to ensure that the distribution characteristics are consistent with the real data.
[0210] 2.2 Five - dimensional Weight Dynamic Adjustment Algorithm
[0211] 2.2.1 Dual - feedback Driving Mechanism
[0212] ```python
[0213] class DimensionOptimizer:
[0214] def __init__(self, dimensions):
[0215] self.weights = dimensions # Initial five - dimensional weights
[0216] self.history_decay = 0.85 # Historical influence decay coefficient
[0217] self.feedback_gain = {
[0218] 'explicit': 1.5, # User directly modifies
[0219] 'implicit': 0.8 # Behavioral deviation prediction
[0220] }
[0221] def update(self, dimension, feedback_type):
[0222] delta = self.feedback_gain[feedback_type] * (1 if feedback_type == 'explicit' else -0.6)
[0223] new_weight = self.weights[dimension] * self.history_decay + delta
[0224] self.weights[dimension] = np.clip(new_weight, 0.1, 1.0)
[0225] return self.weights
[0226] ```
[0227] Explicit feedback: The user actively adjusts the AI - generated content (such as rewriting the sentiment expression in the prompt), corresponding to a 15% weight gain in the corresponding dimension.
[0228] Implicit feedback: Continuous deviation from the predicted path (e.g., social dimension prediction is "consult friends" but actually makes an independent decision) - triggers weight decay.
[0229] 2.2.2 Cross-dimensional coupling analysis
[0230] Construct a five-dimensional correlation matrix to identify synergistic / antagonistic effects (e.g., a high emotional dimension may inhibit the cognitive dimension, with a correlation coefficient of -0.32).
[0231] Adopt Granger causality test to verify the direction of time series influence (e.g., whether the change in the behavior dimension precedes the evolution dimension).
[0232] 2.3 Knowledge base construction and evolution mechanism
[0233] 2.3.1 Multimodal RAG architecture
[0234] Graph database storage: Use Neo4j to store the five-dimensional relationship network, where nodes include (decision-making cases, psychological characteristics, environmental context), and edges define the cross-dimensional influence coefficients.
[0235] Vectorized retrieval enhancement: Encode the user's historical decisions into 768-dimensional vectors and achieve context-aware matching through the Faiss engine (association suggestions are triggered when the similarity > 0.7).
[0236] 2.3.2 Dynamic prompt word generation engine
[0237] Five-dimensional label injection template:
[0238] ```markdown
[0239] # Cognitive weight 0.7 & Social weight 0.5
[0240] The generated suggestions need to meet:
[0241] Provide 3 rational comparison schemes (highlighting cost-effectiveness)
[0242] Integrate 2 community decision-making reference cases
[0243] Limit the emotional expression intensity below Level 3
[0244] ```
[0245] Context adaptive optimization: Dynamically adjust the dimension focus of suggestions by combining real-time environmental data (such as calendar schedule tightness).
[0246] 2.3.3 Knowledge evolution strategy
[0247] Incremental learning mechanism: Start subgraph update at midnight every day, only process data changes in the last 24 hours (full graph rebalancing is triggered when the change impact > 0.3).
[0248] Confidence Elimination Algorithm: Exponentially decay the nodes that have not been activated for 6 months (Confidence = Initial Value × 0.9^t), and archive them to the historical database when it is lower than 0.2.
[0249] 2.4 Privacy Protection and System Efficiency
[0250] 2.4.1 Gradient Desensitization Framework
[0251] Add differential privacy noise (ε = 0.4) during the feature extraction stage to ensure that single data cannot be traced.
[0252] Implement k-anonymization (k = 50) during cloud synchronization to obfuscate the characteristics of user groups.
[0253] 2.4.2 Edge-Cloud Collaborative Computing
[0254] Deploy a lightweight GNN model locally (parameter quantity < 80MB), and only upload the dimension adjustment parameters.
[0255] Adopt the federated learning framework to update the group model and avoid the out-of-domain of raw data.
[0256] 2.5 Innovative Technology Integration
[0257] 2.5.1 Cross-Dimensional Coupling Analysis: Break through the limitations of traditional single-dimensional portraits, and integrate the five-dimensional dynamic associations of cognition, behavior, emotion, society, and evolution.
[0258] 2.5.2 Synthetic Data and Federated Learning: Generate synthetic data through GANs to solve the cold start problem, and combine federated learning to achieve privacy protection.
[0259] 2.5.3 Dynamic Knowledge Base Architecture: Support millisecond-level retrieval of hundreds of millions of decision cases, and combine incremental learning and confidence elimination to achieve knowledge self-evolution.
[0260] 3. System Architecture and Data Flow Design
[0261] 3.1 Overall System Architecture
[0262] Core Architecture: Adopt a four-layer and three-ring architecture of local-cloud collaboration (data access layer, feature processing layer, knowledge storage layer, application service layer), integrate the edge computing module and five-dimensional dynamic modeling capabilities, and form a closed-loop data flow. Combine multi-modal fusion, dynamic knowledge evolution, and privacy protection mechanisms to achieve three-dimensional modeling of user behavior and cognition. Please refer to Figure 2 , Figure 2 which is the system framework diagram of a personal thinking imitation system provided by an embodiment of the present invention.
[0263] 3.2 Detailed Explanation of the Hierarchical Architecture
[0264] 3.2.1 Data Access Layer
[0265] 3.2.1.1 Multi-source data acquisition module: Supports voice, gesture interaction, and dynamic questionnaire systems, and captures five-dimensional guiding behavior data in real time (such as preference selection for logical rigor). Integrates multi-scenario interaction interfaces such as chat record import, psychological tests, and literature upload.
[0266] Public domain data: Obtains public interaction data (retweets / comments / original content) of users through the open platform API of social platforms such as Weibo.
[0267] Private domain data: Connects to note-taking tools such as Obsidian / Notion, and uses the BERT model to extract entity relationship triples in the notes (such as mind map nodes - cognitive dimension features).
[0268] Interaction data: Embedded dynamic questionnaire system (Typeform customization), which captures the preference adjustment behavior of users during AI training in real time.
[0269] 3.2.1.2 Privacy preprocessing module:
[0270] Implement gradient desensitization (ε = 0.4) and k-anonymization (k = 50), and generalize sensitive information (such as geographical location).
[0271] 3.2.2 Feature processing layer
[0272] 3.2.2.1 Fractal analysis engine:
[0273] Three-level segmentation strategy:
[0274] ```python
[0275] # Example of document parsing process
[0276] def fractal_parse(doc):
[0277] main_concept = LayoutLMv3(doc).extract_headings() # Main concept extraction
[0278] supporting_args = SpaCy_NER(doc).get_arguments() # Identification of supporting arguments
[0279] atomic_knowledge = Tesseract_OCR(doc).segment() # Atomic splitting of meta-knowledge return construct_subgraph(main_concept, supporting_args, atomic_knowledge)
[0280] ```
[0281] Multimodal alignment: A spatio-temporal alignment technology that aligns text, image, and speech features through a cross-modal attention mechanism).
[0282] Dynamic feature fusion: Using a graph attention network (GAT) to model cross-modal associations and generate 768-dimensional joint semantic vectors.
[0283] 3.2.3 Knowledge storage layer
[0284] 3.2.3.1 Dual-engine storage architecture:
[0285] Graph database (Neo4j): Stores a five-dimensional relational network (cognitive / behavioral / affective / social / evolutionary dimensions), and nodes contain timestamps and confidence labels.
[0286] Vector database (Milvus): Stores multimodal feature vectors and supports millisecond-level similarity retrieval (optimized by the Faiss engine).
[0287] 3.2.3.2 Incremental update mechanism:
[0288] The subgraph update is started at 0:00 every day, and the temporal graph convolutional network for full-graph rebalancing is triggered when the change impact degree > 0.3.
[0289] 3.2.4 Application service layer
[0290] 3.2.4.1 Dynamic prompt generation engine:
[0291] Five-dimensional weight injection template (RAG enhancement strategy):
[0292] ```markdown
[0293] # Cognitive weight 0.8 & Emotional weight 0.6
[0294] Generation rules:
[0295] Provide 4 logical comparison schemes;
[0296] Integrate 3 historical similar cases;
[0297] Emotional intensity is limited to Level 2;
[0298] ```
[0299] Federated learning service: Deploy a lightweight GNN model (80MB) at the edge, only upload dimensional parameters, and the cloud aggregates and updates through the FedAvg algorithm.
[0300] 3.3 Core data flow:
[0301] Input stage: Multi-source data aggregation - De-identification processing - Local encrypted storage.
[0302] Processing stage: Feature extraction (Temporal pattern analysis, Sentiment analysis) - Knowledge fusion (Dynamic update of the knowledge graph).
[0303] Output and application: Generate decision-making suggestions on demand - User feedback - Modify weights - Iterate the knowledge base.
[0304] 3.3.1 Data input stage
[0305] 3.3.1.1 Multimodal data stream: Weibo API - Text cleaning; User notes - OCR parsing; Dynamic questionnaire - Behavior coding; The text cleaning results, OCR parsing results, and behavior coding results are feature-aligned by the feature alignment module.
[0306] 3.3.1.2 Key processing:
[0307] Text data extracts entity relationships through BERT+CRF
[0308] Image / video data uses the CLIP model to extract cross-modal features
[0309] 3.3.2 Knowledge construction stage
[0310] 3.3.2.1 Cross-dimensional association modeling:
[0311] Construct a five-dimensional coupling matrix to identify the synergistic / antagonistic effects between dimensions (e.g., high social dimension inhibits emotional expression, r=-0.32)
[0312] Adopt Granger causality test to verify the leading influence of the behavior dimension on the evolution dimension (p<0.01)
[0313] 3.3.3 Service output stage
[0314] 3.3.3.1 Multi-level retrieval enhancement:
[0315] First level: Vector similarity retrieval based on Faiss (Top 50 candidates)
[0316] Second level: Calculation of the path credibility of the graph neural network (Optimized by the HNSW algorithm)
[0317] Third level: Weighting of the user profile correlation (Dimensions with a historical adoption rate>60% are prioritized)
[0318] 3.4 Key technology implementation
[0319] 3.4.1 Fractal analysis optimization
[0320] The LayoutLMv3 model is used to implement document layout analysis (F1-score 0.89), and the parsing speed is three times faster than that of the traditional PyMuPDF solution.
[0321] 3.4.2 Dynamic Vector Space Management
[0322] Design a decay mechanism with a forgetting coefficient λ = 0.9:
[0323] ```python
[0324] def update_vector(old_vec, new_vec, λ = 0.9):
[0325] return λ * old_vec + (1 - λ) * new_vec
[0326] ```
[0327] Support the timeliness maintenance of historical decision-making cases (half-life of 6 months).
[0328] 3.4.3 Cross-Domain Mapping Gateway
[0329] Five-Stage Alignment Strategy:
[0330] 1. Lexical level: Cross-training with FastText;
[0331] 2. Entity level: TransEdge algorithm;
[0332] 3. Attribute level: Graph attention network;
[0333] 4. Instance level: Adversarial training;
[0334] 5. Rule level: SWRL reasoning;
[0335] 4. User Interaction Process Design: Based on Multimodal Data Input and Dynamic Feedback Mechanism
[0336] User local data - Encryption and cleaning - Five-dimensional analysis - Generate RAG knowledge base - Cloud synchronization - Inject exclusive prompt words when calling large models (such as GPT-4)
[0337] 4.1 Interaction Process Framework
[0338] This system adopts a "five-stage closed-loop interaction model" to cover the entire process from user data input to obtaining decision-making suggestions. Please refer to Figure 3 , Figure 3 the interaction schematic diagram of the personal thinking imitation system provided by the embodiments of the present invention, and optimize the operation path in combination with the characteristics of the dynamic knowledge base.
[0339] 4.2 Detailed Explanation of Core Interaction Stages
[0340] 4.2.1 Data Authorization and Input
[0341] Multi-modal Access Channels:
[0342] Public Domain Data: Automatically scraped from Weibo / Zhihu APIs (user authorization for OAuth2.0 protocol required);
[0343] Private Domain Data: Support for importing Notion / Obsidian notes (using document parsing technology);
[0344] Interaction Data: Embedded dynamic questionnaires (Typeform customization, support for voice / gesture input);
[0345] Privacy Control Panel:
[0346] Provide 23 fine-grained permission options (such as "only analyze forwarded content", "do not parse like records");
[0347] The default authorization validity period is 7 days, and a renewal reminder is triggered 3 days before expiration.
[0348] 4.2.2 Dynamic Calibration and Training
[0349] 4.2.2.1 Deviation Detection Mechanism:
[0350] ```python
[0351] # Questionnaire and Behavior Data Alignment Algorithm
[0352] def alignment_check(questionnaire, behavior_data):
[0353] cosine_sim = cosine_similarity(
[0354] bert_embed(questionnaire),
[0355] lda_embed(behavior_data) )
[0357] if cosine_sim < 0.75: # Deviation > 25%
[0358] trigger_double_validation()
[0359] ```
[0360] 4.2.2.2 AI Training Sandbox:
[0361] Provide a visual parameter adjustment interface (cognitive / affective / social dimension sliders)
[0362] Support the playback and annotation of historical decision-making cases (feedback enhancement strategy)
[0363] 4.2.3 Decision Generation and Optimization
[0364] 4.2.3.1 Multi-level suggestion generation, please refer to Table 1 for reference
[0365] Table 1
[0366]
[0367] 4.2.3.2 Real-time optimization interface
[0368] Support one-key optimizations such as "logic strengthening" and "emotion weakening". Provide visual traceability of the decision-making path (Neo4j relationship graph rendering).
[0369] 4.3 Exception Handling and Feedback Mechanism
[0370] 4.3.1 Fault Tolerance Design
[0371] 4.3.1.1 Input Exception
[0372] The failure of unstructured document parsing triggers the OCR enhancement mode (Tesseract + LayoutLMv3); the conflict of sensor data starts the multi-modal voting mechanism.
[0373] 4.3.1.2 Logic Exception
[0374] The contradiction of five-dimensional weights activates the Granger causality test (reconstruct the correlation matrix when p < 0.01).
[0375] 4.3.2 Progressive Feedback
[0376] 4.3.2.1 Immediate Feedback
[0377] Operation confirmation (micro-vibration + visual highlighting).
[0378] Processing progress (circular progress bar + estimated remaining time).
[0379] 4.3.2.2 Delayed Feedback
[0380] Daily behavior analysis report (PDF + interactive dashboard).
[0381] Monthly cognitive evolution graph (dynamic GNN visualization).
[0382] 4.4 Key Technical Innovations
[0383] 4.4.1 Context-Aware Process Jumps
[0384] Identifying decision-making urgency through calendar events:
[0385] ```python
[0386] ifschedule_tension>0.8:# Three days before the deadline
[0387] bypass_calibration()# Skip the questionnaire calibration phase
[0388] activate_efficiency_mode()# Enable the efficiency-first weight
[0389] ```
[0390] Optimize the interaction path based on the time pressure model.
[0391] 4.4.2 Interaction Synchronization in the Federated Learning Environment
[0392] Edge-cloud collaborative update mechanism:
[0393] Locally save the interaction records of the last 7 days (lightweight SQLite)
[0394] Implement gradient obfuscation during cloud synchronization (differential privacy scheme)
[0395] Through the above technical solutions, the present invention realizes the accurate imitation of personal thinking and efficient auxiliary decision-making, while taking into account the local fast response and the powerful computing power of the cloud. By integrating edge computing, blockchain, distributed machine learning, encryption technology, as well as AI and large AI models, the system not only improves performance and security, but also enhances scalability and user-friendliness.
[0396] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a personal thinking imitation system 110 provided by an embodiment of the present invention, including:
[0397] An acquisition module 1101, configured to collect multi-source data of a target user in the public domain, private domain, and dynamic questionnaire and perform privacy preprocessing to obtain desensitized behavior data; extract features from the desensitized behavior data by using a three-level segmentation strategy, and perform cross-modal association through a graph attention network to obtain a multi-dimensional joint semantic vector; construct a graph database and a vector storage database based on the multi-dimensional joint semantic vector, where the graph database is used to store a five-dimensional knowledge graph in the graph database, and the five-dimensional knowledge graph includes a cognitive dimension, a behavior dimension, an emotional dimension, a social dimension, and an evolutionary dimension; construct a dynamic prompt word based on the five-dimensional knowledge graph, a preset five-dimensional weight injection template, and real-time environmental data;
[0398] The execution module 1102 is used to perform multi-level retrieval enhancement based on the dynamic prompt words, the graph database and the vector storage database, and generate a final decision suggestion as a personal thinking simulation result for the target user in combination with a pre-trained large model.
[0399] It should be noted that the implementation principle of the aforementioned personal thinking simulation system 110 can refer to the implementation principle of the aforementioned personal thinking simulation method, which will not be repeated here. It should be understood that the division of the various modules of the above device is only a division of logical functions, and in actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated.
[0400] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned personal thinking simulation system 110. Figure 5 As shown, Figure 5 The computer device 100 provided in the embodiment of the present invention is a structural block diagram. The computer device 100 includes a personal thinking simulation system 110, a memory 111, a processor 112 and a communication unit 113.
[0401] In order to realize data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are electrically connected to each other directly or indirectly. For example, the electrical connection between these elements can be realized through one or more communication buses or signal lines. The personal thinking simulation system 110 includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the personal thinking simulation system 110 stored in the memory 111, such as the software function modules and computer programs included in the personal thinking simulation system 110.
[0402] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is running, it controls the computer device where the readable storage medium is located to execute the aforementioned personal thinking simulation system 110.
[0403] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.
Claims
1. A method for imitating personal thinking, characterized in that, Including: Performing multi-source data collection on the target user in the public domain, private domain, and dynamic questionnaire, and performing privacy preprocessing to obtain desensitized behavior data; Adopting a three-level segmentation strategy to extract features from the desensitized behavior data, and performing cross-modal association through a graph attention network to obtain multi-dimensional joint semantic vectors; Constructing a graph database and a vector storage database based on the multi-dimensional joint semantic vectors, where the graph database is used to store a five-dimensional knowledge graph in the graph database, and the five-dimensional knowledge graph includes a cognitive dimension, a behavior dimension, an emotion dimension, a social dimension, and an evolution dimension; Constructing dynamic prompt words based on the five-dimensional knowledge graph, a preset five-dimensional weight injection template, and real-time environmental data; Performing multi-level retrieval enhancement based on the dynamic prompt words, the graph database, and the vector storage database, and combining a pre-trained large model to generate a final decision recommendation as the personal thinking imitation result for the target user.
2. The method according to claim 1, characterized in that, The performing multi-source data collection on the target user in the public domain, private domain, and dynamic questionnaire, and performing privacy preprocessing to obtain desensitized behavior data includes: Obtaining the public interaction data of the target user through the open platform API in the public domain; Obtaining the private data of the target user through a mind map structure parsing strategy in the private domain; Obtaining the questionnaire feedback of the target user by outputting a dynamic questionnaire embedded with dynamic scenario multiple-choice questions; Taking the public interaction data, the private data, and the questionnaire feedback as the multi-source data, and performing gradient desensitization and anonymization processing to obtain the desensitized behavior data.
3. The method according to claim 1, wherein The adopting a three-level segmentation strategy to extract features from the desensitized behavior data, and performing cross-modal association through a graph attention network to obtain multi-dimensional joint semantic vectors includes: Successively performing backbone concept extraction on the desensitized behavior data based on LayoutLMv3, identifying supporting arguments for the desensitized behavior data based on SpaCy_NER, and splitting meta-knowledge atoms for the desensitized behavior data based on Tesseract_OCR to obtain multiple features to be processed obtained based on the three-level segmentation strategy; Adjusting and aligning the multiple features to be processed through a cross-modal attention mechanism; Using the graph attention network to perform cross-modal association on the aligned multiple features to be processed to obtain the multi-dimensional joint semantic vectors, where the cross-modal association includes vocabulary-level association based on FastText cross-training, entity-level association based on the TransEdge algorithm, attribute-level association based on the graph attention network, instance-level association based on adversarial training, and rule-level association based on SWRL reasoning.
4. The method according to claim 1, wherein The method further includes: When the data deviation of at least three dimensions of data in the cognitive dimension, behavior dimension, emotion dimension, social dimension, and evolution dimension included in the five-dimensional knowledge graph exceeds a preset deviation threshold, triggering a multi-modal verification process, and adjusting the five-dimensional knowledge graph according to the verification result. The cognitive dimension uses an LSTM-GNN hybrid model to capture long-term cognitive transition trajectories; the behavioral dimension develops a cognitive-behavior consistency verification algorithm that triggers a dual-verification process when the deviation between the questionnaire statement and public domain behavior exceeds a preset deviation threshold; the emotional dimension develops an emotional state transition matrix and combines seasonal decomposition algorithms to identify periodic patterns; the social dimension introduces a dynamic community discovery algorithm and applies the preferential attachment model in complex network theory to predict the changing trend of social influence within a preset time range; the evolutionary dimension introduces complex system theory, establishes a cognitive attractor model to explain non-linear growth patterns, and simulates cognitive evolution paths in different environments through the Markov chain Monte Carlo method.
5. The method according to claim 1, characterized in that The method further includes: The five-dimensional knowledge graph dynamically adjusts the weights corresponding to the cognitive dimension, the behavioral dimension, the emotional dimension, the social dimension, and the evolutionary dimension through a five-dimensional weight matrix, and dynamically adjusts the focus of the recommended dimension in combination with the real-time environmental data; The five-dimensional weight matrix is adjusted dynamically through the formula: W t+1 = αW t + β(ΔF + γΔC), where W t+1 is the adjusted weight, W t is the weight before adjustment, α is the historical decay coefficient, β is the feedback gain, γ is the cross-dimensional coupling factor, ΔF is the feedback weight adjustment factor, and ΔC is the dimensional weight adjustment factor.
6. The method according to claim 1, characterized in that The method further includes: Obtain the desensitized behavior data on the local side based on the AES-256 encryption algorithm; Deploy the graph attention network on the local side for cross-modal association, and combine differential privacy noise to obtain a multi-dimensional joint semantic vector; Synchronize the multi-dimensional joint semantic vector to the cloud based on anonymization processing; among them, when synchronizing the multi-dimensional joint semantic vector between the local side and the cloud, only the dimension adjustment parameters are uploaded; Implement parameter aggregation and update for the multi-dimensional joint semantic vector through the FedAvg algorithm on the cloud; Deploy a parameter mapping gateway on the cloud to convert the multi-dimensional joint semantic vector into a format adapted to the large model and connect it to the pre-trained large model.
7. The method according to claim 1, characterized in that, The multi-level retrieval enhancement based on the dynamic prompt words, the graph database, and the vector storage database, and the generation of the final decision recommendation in combination with the pre-trained large model as the personal thinking imitation result for the target user includes: Based on the dynamic prompt words, obtain a preset number of first-level multi-dimensional joint vectors from the vector storage database through Faiss vector retrieval; Calculate the HNSW path credibility based on the five-dimensional knowledge graph stored in the graph database, and screen out the second-level multi-dimensional joint vectors from the first-level multi-dimensional joint vectors; Weight the second-level multi-dimensional joint vectors based on the user portrait correlation degree determined by the historical adoption rate to determine the third-level multi-dimensional joint vectors; Execute the pre-trained large model based on the dynamic prompt words and the third-level multi-dimensional joint vectors to generate the final decision recommendation as the personal thinking imitation result for the target user.
8. The method according to claim 1, wherein The construction of the five-dimensional knowledge graph adopts an incremental learning mechanism, starts subgraph updates at preset cycle nodes, processes data changes within a preset data valid range, and triggers a full-graph rebalancing when the change impact degree exceeds a preset impact degree threshold; The five-dimensional knowledge graph exponentially decays nodes that have not been activated for a preset duration, and archives them to the historical database when the confidence level is lower than the preset decay confidence level.
9. The method according to claim 1, wherein The graph database uses Neo4j to store the five-dimensional knowledge graph. The nodes include decision-making cases, psychological characteristics, environmental context, timestamps, and confidence labels, and the edges define the cross-dimensional influence coefficients. The vector storage database uses Milvus to store the multi-dimensional joint semantic vectors and is optimized by the Faiss engine. When the node similarity exceeds the preset similarity threshold, an association recommendation is triggered.
10. A personal thinking imitation system, characterized in that, It includes: An acquisition module, configured to perform multi-source data collection on the target user in the public domain, private domain, and dynamic questionnaire and perform privacy preprocessing to obtain desensitized behavior data. Adopt a three-level segmentation strategy to extract features from the desensitized behavior data, and perform cross-modal association through a graph attention network to obtain multi-dimensional joint semantic vectors; construct a graph database and a vector storage database based on the multi-dimensional joint semantic vectors. The graph database is used to store the five-dimensional knowledge graph in the graph database. The five-dimensional knowledge graph includes a cognitive dimension, a behavior dimension, an emotional dimension, a social dimension, and an evolutionary dimension; construct dynamic prompt words based on the five-dimensional knowledge graph, a preset five-dimensional weight injection template, and real-time environmental data. An execution module, configured to perform multi-level retrieval enhancement based on the dynamic prompt words, the graph database, and the vector storage database, and combine a pre-trained large model to generate a final decision recommendation as the personal thinking imitation result for the target user.
Citation Information
Patent Citations
Deep reinforcement learning interactive recommendation system and method based on knowledge enhancement
CN114117220A
Government affair service field multi-strategy fusion dialogue method based on knowledge graph
CN116628172A
Knowledge graph construction method, device and equipment and readable storage medium
CN118152591A
User-driven knowledge graph construction method and device and medium
CN118469004A
Literature systematic retrieval enhancement and knowledge mining method based on citation network
CN118861130A
Cited By
Multi-modal homework correction system based on image-text interlaced thinking chain
CN121388999A
Multi-agent collaborative product closed-loop optimization system and method based on user feedback data
CN121615880A
AI interaction-based personalized knowledge graph dialogue generation method for old people
CN121658617A
Figure digital twinning duplicating method, interaction method and system
CN122453992A