Interaction task generation method and device, and storage medium

By introducing long-term preference graphs and short-term emotional state memories, combined with fine-tuning of large language models and reinforcement learning, this approach addresses the shortcomings of existing task generation methods. It achieves spatial consistency, semantic rationality, and personalized adaptive recommendation in the task generation process, thus meeting the comprehensive needs for cognitive authenticity, execution accessibility, and user experience in multimodal interaction scenarios.

CN120822625BActive Publication Date: 2025-12-26雅安市人民医院
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511333344.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-12-26
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

Existing interactive task generation methods lack effective modeling of user spatial state and geographical relationships, making it difficult to introduce spatial semantic constraints and dynamic scene understanding into interactive tasks. This results in deviations between the generated results and actual needs, affecting the naturalness of the interaction and execution efficiency.

Method used

By constructing long-term preference maps and short-term emotional state memories, combined with fine-tuning of large language models and reinforcement learning, personalized interactive tasks are generated. The task content is adjusted in real time to adapt to changes in user behavior and state. Spatial semantic maps are introduced to enhance the system's modeling capabilities.

Benefits of technology

It achieves spatial consistency, semantic rationality, and personalized adaptive recommendation in the task generation process, meeting the comprehensive needs for cognitive authenticity, execution accessibility, and user experience in multimodal interaction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822625B_ABST
    Figure CN120822625B_ABST
Patent Text Reader

Abstract

The application discloses an interactive task generation method and device and a storage medium, relates to the field of human-computer interaction, and comprises the following steps: acquiring context information, task knowledge and a task target; constructing a long-term preference graph and a heterogeneous task knowledge base; fine-tuning a large language model; acquiring interactive state information; constructing a short-term emotional state memory, analyzing preference results; generating a query vector; generating a retrieval result; generating a memory-enhanced template; generating a current interactive task; collecting behavior data during execution of the current interactive task; analyzing the behavior data, constructing a multi-dimensional reward; and updating the long-term preference graph and the short-term emotional state memory. By constructing a long-term preference graph and a short-term emotion recognition mechanism, the modeling capability for individual state and environmental context of a user is improved, spatial consistency, semantic rationality and personalized adaptive recommendation in the task generation process are realized, and the comprehensive needs for cognitive authenticity, execution accessibility and user experience continuity in a multi-modal interactive scene are met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of human-computer interaction, and in particular to an interactive task generation method and device and storage medium. BACKGROUND

[0002] In recent years, with the rapid development of artificial intelligence technology, personalized interactive task generation methods have been widely used in cognitive rehabilitation, human-computer dialogue, intelligent companionship and education guidance, etc. The degree of personalization and dynamic adaptability of the generated content has become a key factor affecting the quality of system interaction and the efficiency of user task completion.

[0003] Conventional interactive tasks often occur in physical environments with clear spatial constraints and situational characteristics, such as multi-story rehabilitation training sites, interactive learning spaces, etc. Different areas have functional differentiation, equipment differences and task adaptability, and users' behavior, emotional response and task acceptance vary significantly in different spatial locations. However, existing methods generally lack effective modeling of user spatial state and geographical relationship, making it difficult to introduce spatial semantic constraints and dynamic scene understanding capabilities into interactive tasks. In addition, existing task generation methods mainly rely on static templates or preset rules, lacking perception and modeling of users' long-term interest preferences and short-term interaction states, making it difficult to dynamically adjust and personalize task content based on user behavior and state changes, resulting in deviations between generated results and actual needs, and making it difficult to meet the requirements of continuity, diversity and semantic consistency. In particular, in scenarios involving spatial movement or multi-task coherent recommendation, traditional systems cannot identify the spatial logical relationship between tasks, making it prone to problems such as task jumping, unreasonable paths or situational fragmentation, affecting the naturalness of interaction and execution efficiency. SUMMARY

[0004] The purpose of the present application is to design an interactive task generation method, device and storage medium to solve the above problems.

[0005] The present application achieves the above-mentioned purpose through the following technical solutions:

[0006] The interactive task generation method comprises:

[0007] S1, obtaining the context information C of the user u, the task knowledge and the task target, the context information C including the historical dialogue, the user u preference and the current knowledge demand;

[0008] S2, constructing a long-term preference graph according to the context information C, and constructing a heterogeneous task knowledge base according to the task knowledge;

[0009] S3, fine-tuning the large language model;

[0010] S4, obtaining the interaction state information S t ;

[0011] S5、According to the current interaction state information S t Construct a short-term emotional state memory, and determine a preference result according to a task target and a long-term preference graph;

[0012] S6, generating a query vector according to the long-term preference graph and the short-term emotional state memory;

[0013] S7, retrieving in a heterogeneous task knowledge base according to the query vector, and obtaining a knowledge fragment set as a retrieval result;

[0014] S8, embedding the preference result, the short-term emotional state memory and the retrieval result into a prompt word template to generate a memory-enhanced template;

[0015] S9, inputting the memory-enhanced template into a fine-tuned large language model according to the task target to generate a current interaction task;

[0016] S10, collecting behavior data of a user u in a process of performing the current interaction task;

[0017] S11, evaluating and analyzing the behavior data to construct a multi-dimensional reward model R;

[0018] S12, using a proximal policy optimization algorithm and a REINFORCE algorithm to optimize and update the long-term preference graph and the short-term emotional state memory according to the multi-dimensional reward model R;

[0019] S13, judging whether the task target is completed, if yes, ending; otherwise, taking the behavior data as the interaction state information S t , and returning to S4.

[0020] An interaction task generation device, comprising:

[0021] a storage; the storage is used for storing a computer program;

[0022] an executor; the executor is used for executing the computer program stored in the storage, and when the computer program is executed, the interaction task generation method as described above is realized.

[0023] A computer readable storage medium, the computer readable storage medium stores a computer program, and when the computer program is executed, the interaction task generation method as described above is realized.

[0024] The method of the present application has the beneficial effect that the method constructs a long-term preference graph and a short-term emotion recognition mechanism by introducing a spatial semantic map, improves the modeling capability of the system for the individual state of the user and the environmental context, combines the dynamically constructed prompt input and the strategy adjustment mechanism based on user feedback, realizes the spatial consistency, semantic rationality and personalized adaptive recommendation of the task generation process, and thus meets the comprehensive needs of cognitive authenticity, execution accessibility and user experience continuity in the multi-modal interaction scene. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 is the interactive task generation process of the method;

[0026] Figure 2 is a schematic diagram of an embodiment. DETAILED DESCRIPTION

[0027] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings can be arranged and designed in various different configurations.

[0028] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative labor are within the scope of protection of the present application.

[0029] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0030] In the description of the present application, it should be understood that the terms "upper", "lower", "inner", "outer", "left", "right", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship commonly placed when the product of the present application is used, or the orientation or positional relationship commonly understood by those skilled in the art, and are only for the convenience of describing the present application and simplifying the description, and therefore cannot be understood as indicating or implying that the indicated device or element must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application.

[0031] In addition, the terms "first", "second", etc. are only used for differentiation in description, and cannot be understood as indicating or implying relative importance.

[0032] In the description of the present application, it should be noted that, unless otherwise explicitly specified and limited, the terms such as "arrangement", "connection" should be understood in a broad sense, for example, "connection" can be fixed connection, or detachable connection, or integral connection; can be mechanical connection, or electrical connection; can be direct connection, or indirect connection through intermediate medium, or internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0033] The specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0034] As Figure 1 shown, the interactive task generation method comprises:

[0035] S1, obtaining the context information C of the user u, the task knowledge and the task target, the context information C comprising the historical dialogue, the user u preference and the current knowledge demand.

[0036] S2, constructing a long-term preference graph according to the context information C, and constructing a heterogeneous task knowledge base according to the task knowledge; the long-term preference graph is constructed in particular as follows: a user u spatial behavior graph is constructed based on a graph neural network, and is represented as: , wherein the node set provides entity basis for the triple. Among them, V1 is a floor node, V2 is a region node, and V3 is an object node; the edge set E includes a 3D space connection edge and a user u behavior edge, and is represented as: , E topo is used to connect the accessible area, E act is used to connect the user u and the interactive object; the user u spatial behavior graph is taken as the long-term preference graph, and a graph attention network is used to encode the graph, and the embedding of the node v at the n th layer is represented as , wherein is the attention weight between nodes, is a linear transformation matrix, and sigma is an activation function; the heterogeneous task knowledge base K is represented as: , wherein each segment d c represents the c th task related semantic unit, and is encoded into a knowledge vector by inputting a double tower embedding model, and is represented as: ;

[0037] The long-term preference graph provides a robust context basis for personalized rehabilitation task generation, generates customized main tasks and subtasks that meet medical standards and have higher adaptability. By introducing a structured spatial graph and a semantic modeling mechanism, the matching degree and stability between the generated task content and the user u cognitive habits are improved, and long-term memory support is provided for subsequent semantic retrieval and prompt word construction;

[0038] S3, fine-tuning the large language model; the fine-tuning process of the large language model uses the instruction response data set D train with the goal of maximizing the generation probability of the target output sequence y = {y1, y2, …, y Q} is to optimize the cross-entropy loss function represented as: , wherein, is a large language model containing integrated low-rank adaptive adjustment LoRA insertion parameters; x is an input memory-enhanced template containing task objectives, preference results, and retrieval results; y is the current interactive task; is the probability of the large language model predicting the qth unit token under the memory-enhanced template x.

[0039] S4, real-time acquisition of current interaction state information S t Before the first task execution, before the first round of task generation, an initial interaction question is initiated to the user, and the state of the user answering the interaction question is obtained as the interaction state information S t When the next round of tasks is generated, the interaction state information S t of the previous round of interactive tasks is used. The interaction question includes but is not limited to the current emotional state, physical energy condition, space location, recent activity, and expected task objective, etc. The user's answer is encoded into an initial short-term memory vector by a multi-modal embedding model to fill the starting position of the sliding window, ensuring that the first task generation has a basic situational personalized task.

[0040] S5, constructing a short-term emotional state memory according to the current interaction state information S t , specifically: aggregating the continuous multiple interaction state information S t into a short-term emotional state memory M t through a sliding window, represented as: , , wherein the interaction state information includes input images I t obtained from the interaction interface or perception device, current task specified text x t , state information o t extracted from the task-related objects in the input image, and three-dimensional coordinates p t returned by the perception device or simulator. The preference result is determined according to the task objective and the long-term preference graph, specifically: according to the long-term preference graph, first encode the nodes using a graph attention network, and calculate the node importance weight by combining the user interaction frequency, feedback score, etc. Then, the task objective is combined with the node vector The similarity matching is performed, and the relevance score is obtained by integrating the node importance weight. The node and its relationship most matched with the task target are screened according to the score, and then entity recognition, relationship extraction and triple organization are performed to generate a structured knowledge expression, and finally a triple set most relevant to the current task target, i.e. the preference result, is formed; B is the number of graph attention network layers.

[0041] S6, generating a query vector according to the long-term preference graph and the short-term emotional state memory; specifically comprising:

[0042] S601, obtaining a preference vector of the user u according to the long-term preference graph , which is expressed as: , wherein V u is a set of historical interaction nodes of the user u, γ v is a node importance weight, and B is the number of graph attention network layers.

[0043] S602, obtaining a short-term memory vector according to the short-term emotional state memory , which is expressed as: , wherein, e is a multi-modal memory vector of the e-th round task, is a learnable attention weight parameter, is an importance weight of the memory content, and T is a short-term memory window length; specifically, the short-term emotional state memory M t is input into a multi-modal embedding model to be converted into a unified vector form z t , and the short-term memory vector is obtained by aggregating several round memory vectors in the time window.

[0044] S603, inputting the preference vector and the short-term memory vector into a multi-layer perceptron to generate a fusion query vector.

[0045] S7, retrieving in a heterogeneous task knowledge base according to the query vector to obtain a knowledge fragment set as a retrieval result; specifically, the cosine similarity is used to calculate the similarity between each knowledge vector and the query vector, and the top O most relevant knowledge fragment set D q is selected as the retrieval result.

[0046] S8, embedding the preference result, the short-term emotional state memory and the retrieval result into a prompt word template to generate a memory-enhanced template.

[0047] S9, inputting the memory-enhanced template into a fine-tuned large language model according to the task target to generate a current interaction task.

[0048] ​S10, collect behavior data of user u in the process of performing the current interaction task.

[0049] S11, evaluate and analyze the behavior data to construct a multi-dimensional reward model R; specifically including:

[0050] (1) calculate the interaction sensor force change AF;

[0051] (2) analyze the operation efficiency A by counting the total time required for operation;

[0052] (3) calculate the acceleration variance and path deviation rate of the interaction trajectory as the action stability S using a sliding window, which is expressed as: , wherein, is the variance of the interaction acceleration vector sequence collected in the ith sliding time window, is the average deviation between the actual interaction path and the expected path of user u, is the three-dimensional acceleration vector of the jth sampling point in the ith time window, is the average acceleration vector of the ith time window, and M is the number of acceleration sampling points in the window, is the Euclidean norm, and L is the number of sampling points on the interaction trajectory; is the kth sampling point in the actual operation path of user u, is the kth sampling point on the expected ideal path, and λ is the weight coefficient;

[0053] (4) construct a multi-dimensional reward according to the task completion degree, force change AF, operation efficiency A and action stability S; specifically, normalize the multi-dimensional evaluation indexes of each round of interaction task execution to a vector: , wherein, c , r t , r e , r m correspond to the task completion degree, duration A, action stability S and force change AF respectively. The task completion degree ranges from 0 to 1, and the complete set is 1, and the partial completion is calculated according to the proportion. r t is normalized according to the reference time and actual time consumption, which is expressed as: , wherein, ref is the reference time of task completion, and T obs is the actual time consumption of task completion. Combined with the long-term preference vector , the multi-dimensional reward is obtained, which is expressed as: wherein w = [w1, w2, w3, w4] is a dynamic weight vector, each element in w corresponds to a preset dynamic weight vector, w1, w2, w3 and w4 correspond to the dynamic weight vector of the task completion degree, the dynamic weight vector of the duration A, the dynamic weight vector of the action stability S and the dynamic weight vector of the force variation AF respectively; f = [f1, f2, f3, f4] T is a key indicator vector actually measured, f1, f2, f3 and f4 correspond to the key indicator vector of the task completion degree, the key indicator vector of the duration A, the key indicator vector of the action stability S and the key indicator vector of the force variation AF respectively, P is a safety or abnormal penalty term, λ P is a penalty weight dynamic weight; the dynamic weight vector is represented as: g is a meta-strategy function defined by the meta-strategy parameter θ; represents a preference vector related to the current user u; C represents the context information of the current task scene and type;

[0054] S12, using a proximal policy optimization algorithm and a REINFORCE algorithm, the long-term preference graph and the short-term emotional state memory are updated according to the multi-dimensional reward model R;

[0055] S13, judging whether the task target is completed, if yes, ending; otherwise, returning to S4.

[0056] An interactive task generation device, comprising:

[0057] a storage; the storage is used for storing a computer program;

[0058] an executor; the executor is used for executing the computer program stored in the storage, and when the computer program is executed, the interactive task generation method as described above is realized.

[0059] A computer readable storage medium, the computer readable storage medium stores a computer program, and when the computer program is executed, the interactive task generation method as described above is realized.

[0060] Embodiments, as Figure 2 shown.

[0061] The generation of the large model driven customized interactive task for cognitive rehabilitation tasks is specifically:

[0062] 1) Long-term preference extraction based on user u state modeling and spatial topology atlas enhancement: Collect multi-modal information of user u and model structured atlas, construct long-term preference atlas with semantic continuity and spatial adaptability, provide robust context basis for personalized rehabilitation task generation, generate customized main task and sub-task that meet medical standards and have higher adaptability. Through the introduction of structured spatial atlas and semantic modeling mechanism, the matching degree and stability between generated task content and user u's cognitive habits are improved, and long-term memory support is provided for subsequent semantic retrieval and prompt word construction;

[0063] Specifically includes the following steps:

[0064] Collect multi-modal input data of user u, including user u's text instructions such as "I want to learn how to take care of plants", voice input, camera collected images, and interaction behavior logs (such as clicking the "sprinkler" icon, dragging the "flowerpot" object). Combining these heterogeneous data, the system constructs the current interaction state representation vector of user u, which integrates user u's operation intention, context information, current environment perception result and historical behavior information; To ensure the consistency of the medical context, the system also guides the user u to fill in the basic information, including gender, age, occupation background, interest and hobby, cognitive impairment type, past medical history and rehabilitation stage, etc., and generates a standardized user u background description text as an auxiliary prompt input, such as: "male, 70 years old, mild Alzheimer's disease, retired teacher, hobby is gardening, currently living in urban apartment". This portrait serves as the core context of the task generation prompt, ensuring that the subsequent generated task content fully considers the user u's background and rehabilitation needs, improving the relevance and scientificity of the generated task;

[0065] A three-dimensional semantic graph modeling method with spatial structure perception ability is introduced to construct a long-term preference graph, which expresses the task preference distribution and behavior habits of user u among different spatial nodes. The long-term preference graph adopts a hierarchical topology structure of "floor-region-object": the floor node represents the physical hierarchical region of the rehabilitation center, the region node corresponds to the training scene (such as "garden cognition area" and "tool use area"), and the object node represents specific task tools and equipment (such as "flower shovel", "flower pot", "watering pot", "seed bag", etc.). The edges in the long-term preference graph represent the paths and frequencies of user u in the history task, and give weights based on the interaction intensity, such as user u has completed the "seed selection and watering task" in the "planting area" many times, forming a long-term preference triple as a preference result, such as: "<user uA, prefer to perform, planting cognition task>", "<planting cognition task, success rate higher than average, action cognition task>", "<seed identification task, belonging category, multi-modal identification task>". The preference result is used as the long-term memory part of the input prompt of the fine-tuned large language model, to guide the generation process of the large model to maintain the semantic consistency and spatial adaptability of user u's preference, and to improve the coherence and credibility of the task content.

[0066] 2) Fine-tuned large language model driven rehabilitation scene task generation and interaction: on the basis of completing the long-term preference graph modeling and the current interaction state information S t , a short-term state perception and multi-modal interaction modeling mechanism is introduced to construct a short-term emotional state memory M t , and generate a memory enhanced template, and then generate the current interaction task according to the fine-tuned large language model, realize the rehabilitation task content generation process with instant adaptability and cognitive pertinence, and provide standardized and immersive rehabilitation experience for user u;

[0067] Specifically, the steps include:

[0068] In the process of user u's gardening operation interaction, the latest round of multi-modal input information is collected in real time, including images obtained by the camera, language or text instructions issued by user u, and action operation records on the interaction interface, such as dragging, clicking, selecting, and other behaviors. The embedded visual-language model VLM is called to perform task-related semantic analysis on the current image, identify the category, position, and state information of the key objects in the scene, such as "red flowerpot has been placed" and "sprinkler is in use". On this basis, the system maintains a set of short-term memory state units in a sliding window mechanism to continuously track the behavior and intention evolution of user u in the current interaction stage. The typical information recorded in the memory cache structure includes: the current task target object (such as "seed packet" and "sprinkler"), the object state (such as "identified", "not used", and "operating"), and the spatial region or relative position associated with it (such as "garden cognitive area" and "operation table front"). The short-term emotional state memory is encoded into a short-term memory vector after semantic vector encoding, providing immediate context support for subsequent task target recognition and generation;

[0069] The short-term memory vector is encoded by a Transformer-based semantic encoder model and the preference vector are unified and mapped to a shared semantic space by an attention mechanism to generate a query vector q t , enhancing the collaborative modeling capability between different sources of information. The query vector not only carries the user u's operation target in this round (such as "I want to identify the plant name"), but also introduces high-frequency triple knowledge related to gardening tasks in the long-term preference graph (such as "user uA prefers the sowing task" and "user uA has a higher success rate in the garden cognitive area"), so that the generation process has instantaneity, stability, and personalized orientation;

[0070] The query vector q tTo retrieve the conditions, the three-stage advanced RAG fusion strategy AdvancedRAG-ToolFusion is called to perform vector matching and structured recall in a heterogeneous task knowledge base for retrieving and selecting appropriate tools or tool chains. First, the user u task goal is automatically decomposed into several sub-tasks (such as "identify plant name + prepare spray bottle + complete watering"). Second, the most suitable tool chain combination and task execution path for the current scenario are retrieved in the heterogeneous task knowledge base, and finally a set of knowledge snippets are generated as the retrieval results. The strategy is evaluated in a multi-step task and tool chain retrieval scenario, using a multi-tool collaborative operation benchmark including gardening tasks, and the Seal-Tools dataset and ToolE-Extended dataset are used for generalization testing. Experimental results show that the proposed three-stage advanced RAG fusion strategy AdvancedRAG-ToolFusion has good multi-tool task retrieval generalization ability, can effectively support multi-stage task generation in gardening interaction scenarios, and ensure the accuracy and completeness of tool chain recommendation. The experimental results are shown in Table 1.

[0071] Table 1

[0072]

[0073] According to the task goal, the memory-enhanced template input fine-tuned large language model is used to generate the current interaction task, and the 3D scene space layout, interaction mechanism binding and feedback strategy setting are completed; through the rendering engine, object layout, interaction action and dynamic feedback are realized, and trigger conditions and feedback mechanisms are defined for each interaction type, such as drag object falling into target area triggering completion judgment, single selection click showing correct or not prompt, and interactive training scene is constructed. Finally, an operable scene is formed to support the achievement of rehabilitation training goals, ensuring the continuity of interaction and the standardization of rehabilitation process.

[0074] 3) Task execution interaction and real-time feedback collection: real-time collection of user u behavior data during the user u interaction task process, feedback results are obtained through systematic evaluation, and the task generation model is optimized based on reinforcement learning mechanism to dynamically adjust task difficulty and content, to achieve personalized and adaptive generation of targeted rehabilitation training tasks;

[0075] Specifically, the following steps are included:

[0076] During the process of user u participating in task generation, key interaction data such as operation object, drag path, click selection, completion time, error times, prompt response, etc. are automatically recorded. The key interaction data is processed to extract multi-dimensional features such as operation accuracy, response speed, error mode, and prompt dependence, providing a basis for subsequent evaluation and adjustment;

[0077] Evaluation criteria are set according to different task types, such as completion time, operation accuracy, number of prompts, error rate, etc. The speech feedback of the user u is analyzed to identify the tone emotion, and information such as concentration, happiness, and frustration is inferred. After each round of interaction, the performance of the user u is automatically scored, and a structured feedback result is generated, clearly indicating the user's u ability and weaknesses in the current task. The feedback result is not only displayed on the user side, but also serves as an important basis for updating the policy of the backend reinforcement learning module;

[0078] 4) Reinforcement learning driven generation strategy optimization and personalized update: During the rehabilitation task execution process, multi-dimensional interaction behavior data and emotional state of the user u are collected as behavior data, and an interaction evaluation model is trained through supervised learning to quantify the completion quality of the user u in each task and make a comprehensive judgment. Based on the reinforcement learning mechanism, the task generation strategy is dynamically adjusted to realize personalized adaptation and progressive optimization of the task content, and to promote the continuous improvement and active participation of the user u;

[0079] During the execution of each rehabilitation task by the user u, key behavior indicators are automatically collected through sensors and interaction logs:

[0080] (1) Interaction sensor force change ΔF: assesses operation stability and precision.

[0081] (2) Duration A: statistics operation completion time, reflects operation efficiency.

[0082] (3) Action stability S: sliding window is used to calculate the acceleration variance, path deviation rate or posture change frequency of the interaction trajectory, the calculation formula is:

[0083] ;

[0084] Where, represents the variance of the interaction acceleration vector sequence collected in the i-th sliding time window, the larger the variance value, the more significant the action jitter or instability. represents the average deviation between the actual interaction path of the user u and the expected path, the larger the value, the farther the user u deviates from the operation target path, the worse the stability; represents the three-dimensional acceleration vector of the j-th sampling point in the i-th time window; represents the average acceleration vector of the i-th time window; M represents the number of acceleration sampling points in the window; represents the Euclidean norm. L represents the number of sampling points on the interaction trajectory; represents the k-th sampling point in the actual operation path of the user u, represents the k-th sampling point on the expected ideal path, and λ is the weight coefficient.

[0085] Based on the multi-round interaction data, a comprehensive reward function is constructed, a multi-dimensional reward model R is constructed by fusing task completion degree, force change Delta F, duration A and action stability S, a near-end strategy optimization algorithm and a reinforcement learning strategy gradient algorithm REINFORCE are adopted, long-term preference atlas and short-term emotional state memory are updated according to the multi-dimensional reward model R, at the same time, the short-term memory window is refreshed, the current state change of the user u is recorded, including the key objects involved in the recent interaction, the emotional fluctuation and the space position, the historical state information is replaced, the precise adaptation to the current ability level and state of the user u is realized,

[0086] Combined with the long-term preference atlas and the current task feedback, the response ability to the key demand of the user u is enhanced, the sorting priority of the task content is optimized in real time, it is ensured that the task module most suitable for the interest and ability of the user u is preferentially recommended, and the personalized recommendation effect is improved.

[0087] Based on the updated long-term preference atlas, the system adjusts the difficulty gradient and the content style of the subsequent task, realizes the progressive adjustment and personalized evolution of the task, ensures the continuity and challenge balance of the rehabilitation training process, and improves the rehabilitation motivation and overall rehabilitation effect.

[0088] The technical scheme of the present application is not limited to the above specific embodiments, and any technical modification made according to the technical scheme of the present application falls within the protection scope of the present application.

Claims

1. An interactive task generation method, characterized by, Comprise: S1, acquire the context information C of user u, task knowledge and task target, the context information C includes historical dialogue, user u preference and current knowledge demand; S2, construct long-term preference graph according to context information C, and construct heterogeneous task knowledge base according to task knowledge; the long-term preference graph is specifically constructed as follows: a user u spatial behavior graph is constructed based on a graph neural network, and is represented as: Wherein, the node set provides entity basis for the triplets; wherein, V1 is a floor node, V2 is a region node, and V3 is an object node; the edge set E includes 3D space connection edges and user u behavior edges, and is represented as: , E topo is used for connecting accessible regions, E act is used for connecting the user u and the interactive object; the user u spatial behavior graph is taken as the long-term preference graph, and the graph is encoded using a graph attention network; the embedding of the node v at the nth layer is represented as , wherein, is the attention weight between nodes, is a linear transformation matrix, and sigma is an activation function; the heterogeneous task knowledge base K is represented as: Wherein each piece d c represents a semantic unit related to the cth task, is input into a double-tower embedding model and is encoded into a knowledge vector, and is represented as: ; S3, fine-tuning is carried out to large language model; S4, obtaining interaction state information S t ; S5、 according to the current interaction state information S t constructing short-term emotional state memory, determining preference results according to task targets and long-term preference atlas; S6, generates query vector according to long-term preference graph and short-term emotional state memory; S7, according to the query vector in the heterogeneous task knowledge base retrieval, obtain the knowledge fragment set as the retrieval result; S8, the preference result, short-term emotional state memory and retrieval result are embedded into prompt word template, and memory-enhanced template is generated; S9, according to the task target, memory-enhanced template is input into fine-tuned large language model, and current interactive task is generated; S10, the behavior data in the process that user u executes current interactive task is collected; S11, the behavior data is evaluated and analyzed, and multi-dimensional reward model is constructed;Specifically, it comprises: (1) calculate the interactive sensor force change ΔF; (2) total duration required for operation is analyzed to analyze operation efficiency A; (3) The acceleration variance and path deviation rate of the interaction trajectory calculated by the sliding window are used as the action stability S, which is expressed as: wherein, is the variance of the interaction acceleration vector sequence collected in the i-th sliding time window, is the average deviation between the actual interaction path and the expected path of the user u, is the three-dimensional acceleration vector of the j-th sampling point in the i-th time window, is the average acceleration vector of the i-th time window, and M is the number of acceleration sampling points in the window, is the Euclidean norm, and L is the number of sampling points on the interaction trajectory; is the k-th sampling point in the actual operation path of the user u, is the k-th sampling point on the expected ideal path, and λ is the weight coefficient. (4) A multi-dimensional reward model R is constructed according to the task completion degree, the force variation ΔF, the operation efficiency A and the action stability S; the multi-dimensional reward model R is represented as: wherein w=[w1, w2, w3, w4] is a dynamic weight vector, each element in w corresponds to a preset dynamic weight vector, w1, w2, w3 and w4 respectively correspond to the dynamic weight vector of the task completion degree, the dynamic weight vector of the duration A, the dynamic weight vector of the action stability S and the dynamic weight vector of the force variation ΔF; f=[f1, f2, f3, f4] T is a key indicator vector actually measured, f1, f2, f3 and f4 respectively correspond to the key indicator vector of the task completion degree, the key indicator vector of the duration A, the key indicator vector of the action stability S and the key indicator vector of the force variation ΔF, P is a safety or abnormal penalty term, λ P is a penalty weight; the dynamic weight vector is represented as: g is a meta-strategy function defined by a meta-strategy parameter θ; represents a preference vector related to the current user u; C represents the context information of the current task scene and type. S12, according to the multi-dimensional reward model R, long-term preference graph and short-term emotional state memory are optimized and updated using proximal policy optimization algorithm and reinforcement learning policy gradient algorithm REINFORCE; S13, judge whether the task target is completed, if yes, end; otherwise, the behavior data is taken as the interactive state information S t and return S4.

2. The interactive task generation method of claim 1, wherein, In S3, the fine-tuning process of the large language model uses the instruction response dataset D train with the goal of maximizing the generation probability of the target output sequence y = {y1, y2, …, y Q} Optimization of the cross-entropy loss function is represented as: , wherein is a large language model containing an integrated low-rank adaptive adjustment LoRA insertion parameter; x is an input memory-enhanced template containing a task target, a preferred result, and a retrieval result; y is a current interactive task; is the probability of the qth unit token predicted by the large language model under the memory-enhanced template x.

3. The interactive task generation method of claim 1, wherein, In S6, it comprises: S601、According to the long-term preference map, a preference vector of a user u is acquired is expressed as: wherein V u is a set of historical interaction nodes of the user u, γ v is a node importance weight, B is a number of graph attention network layers, is a node vector; S602, obtaining a short-term memory vector according to the short-term emotional state memory is expressed as: wherein, z e is a multimodal memory vector of the e-th round task, is a learnable attention weight parameter, is an importance weight of the memory content, and T is a short-term memory window length; S603、obtaining the preference vector and the short-term memory vector The multi-layer perception fusion generates the query vector.

4. The interactive task generation method of claim 3, wherein, In S7, cosine similarity is used to calculate the similarity between each knowledge vector and the query vector, and the top O most relevant knowledge fragment sets D are selected q as the search results.

5. The interactive task generation method of claim 1, wherein, In S4, before the first round of task generation, an initial interactive question is initiated to the user to obtain the state of the user's reply to the interactive question as the interactive state information S t .

6. An interactive task generation apparatus characterized by comprising: Comprise: Storage; The storage is used to store computer programs; Actuator; The actuator is used to execute the computer programs stored in the storage, and when the computer programs are executed, the interactive task generation method as claimed in any one of claims 1-5 is realized.

7. A computer readable storage medium, characterized in that, The computer program is stored in the computer readable storage medium, and when the computer program is executed, the interactive task generation method as claimed in any one of claims 1-5 is realized.

Citation Information

Patent Citations

  • Reinforcement learning knowledge graph reasoning method and system guided by confrontation and attention mechanism

    CN118606485A

  • Large language model intelligent interaction method and system with memory management capability

    CN119200821A