Cross-cultural narrative adaptive generation method and device and system based on multi-modal multi-agent collaborative reasoning
The MC-MAS system, which utilizes multimodal and multi-agent collaborative reasoning, solves the problems of heterogeneous multimodal signal conflict and cultural cognition alignment in cross-cultural narrative systems, enabling adaptive generation of narrative content and improving learners' cognitive load management and interactive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EAST CHINA NORMAL UNIV
- Filing Date
- 2026-04-24
- Publication Date
- 2026-06-16
Smart Images

Figure CN122222033A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and natural language processing, specifically to a cross-cultural narrative adaptive generation method, apparatus, and system based on multimodal and multi-agent collaborative reasoning. Background Technology
[0002] Currently, the commonly used technologies in the industry are as follows:
[0003] With the rapid development of Large Language Models (LLMs) in the field of intelligent education, current education and professional training are gradually shifting from single-agent environments to multi-agent collaborative environments to simulate complex interactive dynamics. For example, SimClass constructs a comprehensive classroom environment that includes diverse peer roles and utilizes the Flanders Interaction Analysis System (FIAS) to recreate realistic teaching interactions, as discussed in the paper Zhang, Zheyuan, et al. "Simulating classroom education with LLM-empowered agents." Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. In the field of medical education, MEDCO and Agent Hospital respectively utilize the expert-patient triplet and the "Simulator-based Evolutionary Agent Learning" (SEAL) paradigm to promote the evolution of interdisciplinary reasoning and diagnostic capabilities, as discussed in the literature Wei, Hao, et al. "Medco: Medical education copilots based on a multi-agent framework." European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024 and Li, Junkai, et al. "Agent hospital: A simulacrum of hospital with evolvable medical agents." arXiv preprint arXiv:2405.02957 (2024).This collaborative model has also been applied to the cultural field. For example, TRANSAGENTS generates literary translations that emphasize cultural resonance through a multi-agent virtual company workflow that includes editors and localization experts, as discussed in the paper Wu, Minghao, et al. "(Perhaps) beyond human translation: Harnessing multi-agent collaboration for translating ultra-long literary texts." Transactions of the Association for Computational Linguistics 13 (2025): 901-922.
[0004] Meanwhile, to achieve deeper cognitive and semantic understanding, research has shifted its focus to the deep fusion of heterogeneous data and the calibration of model values, as discussed in Xie, Fang. "From words to multimodalities: Compliment perceptions across lingua cultures." Journal of Pragmatics 239(2025): 94-116. In the field of learning analytics, research by MSC-Trans and Yan et al. has integrated data streams such as facial expressions, postures, and behavioral logs to address the problem of asynchronous data and accurately detect student engagement, as discussed in Xie, Nan, et al. "Msc-trans: A multi-feature-fusion network with encoding structure for student engagement detecting." IEEE Transactions on Learning Technologies 18 (2025): 243-255 and Yan, Lijuan, Xiaotao Wu, and Yi Wang. "Student engagement assessment using multimodal deep learning." PLOS 1 20.6 (2025): e0325377. Furthermore, regarding the issue of cultural consistency in language models, CultureLLM effectively augments data using the World Values Survey, as discussed in the literature Li, Cheng, et al. "Culturellm: Incorporating cultural differences into large language models." Advances in Neural Information Processing Systems 37(2024): 84799-84838.Meanwhile, MENAValues reveals value inconsistencies in the cross-linguistic value transfer process through benchmarking, highlighting the criticality of capturing deep cultural nuances, as discussed in the literature Zahraei, Pardis Sadat, and Ehsaneddin Asgari. "I Am Aligned, But With Whom? MENAValues Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs." arXiv preprint arXiv:2510.13154 (2025).
[0005] Existing interactive narrative systems still face significant challenges in cross-cultural contexts. While existing research has focused on multimodal analysis, most systems still rely on static pre-set scripts or large monolithic models, neglecting the conflicts between cross-modal signals. Cultural misunderstandings often manifest through multimodal inconsistencies, and few studies have addressed how to simultaneously optimize narrative adaptation and cognitive regulation through real-time multi-agent reflection mechanisms. This limits the efficiency of adaptive learning methods in complex cross-cultural scenarios.
[0006] In summary, the problems with existing technologies are:
[0007] Heterogeneous multimodal signal conflicts are difficult to reconcile: Most existing interactive narrative systems rely on single text input, ignoring non-textual signals such as user behavior logs and emotional states. They cannot effectively identify and resolve cross-modal semantic conflicts between user behavior, emotions, and language, resulting in a lack of adaptability in generated feedback and the creation of a "contextual illusion."
[0008] Lack of deep cultural cognitive alignment: Monolithic models and static scripts struggle to perceive users' implicit cultural cognitive states, cannot handle deep cognitive conflicts caused by differences in cultural backgrounds, and are unable to achieve dynamic alignment in cross-cultural contexts.
[0009] Lack of feedback optimization and cognitive load adjustment: Existing technologies focus on narrative fluency rather than the user's actual level of understanding, and lack iterative reflection mechanisms to assess and optimize the user's cognitive load and the accuracy of cultural intent recognition in real time. Summary of the Invention
[0010] To address the problems of existing technologies, the present invention aims to provide a cross-cultural narrative adaptive generation method, apparatus, and system based on multimodal multi-agent collaborative reasoning. The method utilizes a large language model to construct a multimodal culture-aware multi-agent system (MC-MAS), collaboratively processing user behavior trajectories, dialogue text, and real-time facial emotion tags, and combining this with historical interaction context to perform cross-modal collaborative reasoning. In this way, the system can dynamically generate adaptive narrative content highly aligned with the user's psychological motivations and cultural background, effectively identifying and bridging the user's cultural cognitive gap, and significantly reducing the cognitive load in cross-cultural learning.
[0011] The innovation of this invention lies in the introduction of a core "Reflection-Reconstruction Loop" mechanism. Combined with the specialized analytical capabilities of behavioral, linguistic, and cultural expert agents, it enables real-time identification and iterative correction of semantic conflicts between heterogeneous modalities. This design not only effectively solves the background illusion problem that easily arises in cross-cultural narratives with monolithic large language models, but also significantly enhances the accuracy of capturing uncertain interactive intentions through the "identification-reconstruction-fusion-update" logic. Specifically, with the global control of a central coordinating agent, the system can continuously optimize narrative logic, task allocation, and AI assistant guidance strategies based on historical interaction trajectories, making each interaction step highly personalized and better suited to the deeper needs of learners from different cultural backgrounds.
[0012] Furthermore, through layered and dynamically adapted narrative text and guiding information, learners can obtain more intuitive support and cognitive completion in complex cross-cultural contexts, thereby maximizing cultural understanding and knowledge acquisition in immersive interaction. Compared to traditional teaching systems that rely solely on text responses or static scripts, this method significantly enhances the cultural sensitivity of narrative generation by deeply integrating emotional feedback and behavioral patterns, and improves learners' frustration and psychological load management. By combining multimodal collaborative reasoning and adaptive narrative generation, this invention significantly improves the interactive experience of cross-cultural learning, providing a leading technical solution for multimodal teaching with deep cultural perception capabilities in the future field of intelligent education.
[0013] The specific technical solution for achieving the objective of this invention is as follows:
[0014] A cross-cultural narrative adaptive generation method based on multimodal and multi-agent collaborative reasoning includes the following steps:
[0015] Step 1: Based on a large language model, construct a workflow for cross-cultural narrative adaptive multi-agent systems. The agents include a multimodal data acquisition module, an expert agent cluster module, a central coordinating agent module, and a personalized narrative generation module. The expert agent cluster module includes a behavior analysis agent, a language analysis agent, and a cultural context analysis agent.
[0016] Step 2: The multimodal data acquisition module acquires and aligns the user's heterogeneous interaction signals in real time. The heterogeneous interaction signals include dialogue text, operation behavior logs, and facial expression labels recorded through a discrete sequence sampling strategy. The alignment aims to provide a unified semantic and temporal benchmark for subsequent analysis, ultimately forming an aligned multimodal dataset.
[0017] Step 3: The expert agent cluster module performs parallel feature extraction on the collected data: the behavior analysis agent maps the operation logs to the Batu player type distribution; the language analysis agent quantifies cognitive load and emotional state by aligning text and emoji tags; the cultural context analysis agent combines user static profiles and dynamic behaviors to identify cultural comprehension barriers and risk levels in the narrative; finally, a preliminary analysis report is generated, which includes the Batu player type distribution output by the behavior analysis agent, the language profile including language depth and syntactic features output by the language analysis agent, and the cultural comprehension barriers and risk levels output by the literary context analysis agent;
[0018] Step 4: The central coordinating agent module summarizes the preliminary analysis report of the expert agent cluster module, calculates the consistency of cross-modal semantics and determines whether there is a conflict; if the detected semantic conflict exceeds the preset threshold, the "reflection-reconstruction" loop is triggered, generating upper-level control prompts and entering Step 3 to guide the expert agent cluster module to re-evaluate the original data until a decision consensus is reached; if there is no conflict, proceed directly to Step 5.
[0019] Step 5: Based on the decision consensus reached by the central coordinating agent module, generate personalized narrative suggestions; the suggestions include task allocation schemes customized for user player types, as well as explanatory guidance strategies and language style preference adjustment suggestions to address the cultural background gap of users;
[0020] Step 6: The personalized narrative generation module dynamically rewrites the in-game text scripts and task logic based on the generated personalized narrative suggestions. The rewriting includes adjusting the depth of metaphors, adding historical background supplements, or modifying the narrative rhythm, so as to align the narrative context with the user's real-time cultural cognition. The method ends here.
[0021] Furthermore, the behavior analysis agent described in step S1 is used to map the multimodal information contained in the discrete game behavior logs into user psychological motivation features, thereby realizing probabilistic reasoning of Batu player types.
[0022] Language analysis agent: used to identify the user's metaphorical understanding depth and real-time cognitive load status through user input text;
[0023] Cultural context analysis agent: used to assess the cognitive distance between the user's background and the target culture of the narrative, predict potential risks of cultural misunderstandings, and generate suggestions for bridging them;
[0024] Furthermore, in step S2, the formation of the multimodal dataset specifically includes:
[0025] Step 2-1: Collect real-time behavioral trajectories during user interaction, including task time, failure rate, and interaction frequency;
[0026] Step 2-2: Obtain the user's dialogue text sequence and simultaneously obtain the facial expression label sequence using a discrete sequence sampling strategy; the facial expression label sequence is obtained by sampling emotion labels every 5 seconds, and the emotion labels specifically include "confused", "happy", "nervous" and "neutral", without storing the original biometric image data;
[0027] Steps 2-3: Align user interaction data spatiotemporally using a unified time base to construct a structured input pool dataset for subsequent semantic reasoning.
[0028] Furthermore, in step S4, the central coordinating intelligent agent module executes a reflection mechanism, specifically including:
[0029] Step 4-1: The central coordinating agent module aggregates the output results of each expert agent and calculates the semantic conflict index; when the conclusions of any two expert agents are logically inconsistent and the significance exceeds the threshold, a round-robin reflection is triggered.
[0030] Step 4-2: The central coordinating agent module constructs a prompt containing conflict descriptions, original evidence, and correction instructions, and sends it to the corresponding expert agents;
[0031] Step 4-3: The expert agent rewrites its internal reasoning chain under the guidance of prompts;
[0032] Step 4-4: The central coordinating agent module performs weighted fusion on the revised expert vectors to generate the final adaptation decision.
[0033] Furthermore, in step S6, the personalized narrative generation module specifically includes:
[0034] Step 6-1: Task Allocation and Adaptation: Based on the Batu probability distribution vector, assign challenge tasks to achievement-oriented players and cultural unlocking and Easter egg guidance tasks to exploration-oriented players;
[0035] Step 6-2: Narrative style adaptation: Dynamically adjust the metaphor density and explanatory narration in the narrative text according to the risk level of cultural misunderstanding; if users have significant cultural barriers, replace abstract historical metaphors with concrete narrative descriptions;
[0036] Step 6-3: AI Assistant Tone Adaptation: Adjust the AI assistant's interactive tone and guidance depth based on the user's real-time cognitive load, and optimize the learning experience through "vertical cross-stage adaptation" and "horizontal real-time adjustment".
[0037] A cross-cultural narrative adaptive generation device based on multimodal multi-agent collaborative reasoning includes a terminal device, which is an Internet terminal device, comprising a processor and a computer-readable storage medium. The processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, which are loaded and executed by the processor as described above for the cross-cultural narrative adaptive generation method based on multimodal multi-agent collaborative reasoning.
[0038] A cross-cultural narrative adaptive generation system based on multimodal multi-agent collaborative reasoning includes a computer-readable storage medium storing multiple instructions and agent workflows. The instructions are used by the processor of a terminal device to load and execute the aforementioned cross-cultural narrative adaptive generation method based on multimodal multi-agent collaborative reasoning. The agent workflow includes a multimodal data acquisition module, an expert agent cluster module specifically including a behavior analysis agent, a language analysis agent, a cultural context analysis agent, a central coordinating agent module, and a personalized narrative generation module.
[0039] Compared with existing technologies, the cross-cultural narrative adaptive generation method, apparatus, and system based on multimodal multi-agent collaborative reasoning provided by this invention have the following beneficial effects:
[0040] This invention belongs to the field of intelligent education and cross-cultural learning, and discloses a cross-cultural narrative adaptive generation method, device, and system based on multimodal multi-agent collaborative reasoning. The method includes: constructing a multimodal culturally aware multi-agent system (MC-MAS) using a large language model to generate adaptive narrative texts and personalized learning tasks for learners; combining users' real-time operational behavior, dialogue text, and facial emotion tags, and automatically detecting and collaboratively reasoning about cross-modal semantic conflicts through a "reflection-reconstruction" loop mechanism; using a central coordinating agent to guide expert agents in iterative optimization, and applying the reached decision consensus to the dynamic reconstruction of narrative logic and guidance strategies to enhance the system's intent recognition accuracy and adaptability in complex cultural contexts. This method can generate narrative feedback highly aligned with the user's real-time cultural cognitive state, effectively reducing the learner's cognitive load through cross-modal feature fusion, thereby improving the learning effect and immersive experience in cross-cultural interactive learning processes. Attached Figure Description
[0041] Figure 1 This is a flowchart of the method of the present invention.
[0042] Figure 2 This is a structural diagram of the device of the present invention;
[0043] Figure 3 This is a system architecture diagram of the present invention. Detailed Implementation
[0044] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0045] like Figure 1 As shown, a cross-cultural narrative adaptive generation method based on multimodal and multi-agent collaborative reasoning includes the following steps:
[0046] S1: Based on a large language model, a workflow for cross-cultural narrative adaptive multi-agent is constructed by performing user profiling analysis and personalized narrative generation. The agents include a multimodal data acquisition module, an expert agent cluster module, a central coordinating agent module, and a personalized narrative generation module. The expert agent cluster module includes a behavior analysis agent, a language analysis agent, and a cultural context analysis agent.
[0047] Furthermore, in order to enhance the user profile analysis capability based on multimodal data in the intelligent agent workflow, in step S1, the intelligent agent framework for user profile analysis includes an expert intelligent agent cluster module and a central coordinating intelligent agent module.
[0048] The expert agent cluster module comprises three parts: a behavior analysis agent, a language analysis agent, and a cultural context analysis agent. The behavior analysis agent maps multimodal information from discrete game behavior logs to user psychological motivation characteristics, enabling probabilistic inference about Batu player types. The language analysis agent identifies the user's metaphorical comprehension depth and real-time cognitive load state through user input text. The cultural context analysis agent assesses the cognitive distance between the user's background and the target narrative culture, predicts potential cultural misunderstanding risks, and generates bridging suggestions.
[0049] The central coordinating agent module serves as the core of the method, responsible for calculating semantic consistency based on the heterogeneous conclusions output by each expert agent, and using the "identification-reshaping-fusion-update" loop to guide the expert agents to perform self-correction.
[0050] S2: The multimodal data acquisition module acquires and aligns the user's heterogeneous interaction signals in real time. The signals include dialogue text, operation behavior logs, and facial expression tags recorded through a discrete sequence sampling strategy. The alignment operation aims to provide a unified semantic and temporal benchmark for subsequent analysis.
[0051] Furthermore, in order to establish a unified data foundation for expert intelligent agents to perform semantic analysis and intent recognition, the generation process of the multimodal dataset in step S2 includes a real-time behavior trajectory acquisition stage, a heterogeneous interaction signal acquisition stage, and a data spatiotemporal alignment stage.
[0052] During the behavior trajectory collection phase, fine-grained operation data streams are recorded in real time during user interactions to construct user behavior trajectories. These trajectories specifically include task completion status, task duration, operation failure rate, and interaction frequency. This data provides quantitative evidence for the subsequent behavior analysis agent to identify the player's motivational spectrum and Batu game personality type.
[0053] During the heterogeneous interaction signal acquisition phase, the user's verbal expression and emotional feedback signals are acquired simultaneously. The user's dialogue text sequence is recorded, and a discrete sequence sampling strategy is used to capture the user's facial expression tag sequence. To protect user privacy while analyzing cognitive load, the expression tag sequence is generated by sampling specific emotional tags at 5-second intervals. The emotional tags include "confused," "happy," "nervous," and "neutral," and no raw biometric image or video data is stored.
[0054] In the data spatiotemporal alignment phase, the aforementioned behavioral trajectories, dialogue texts, and facial expression tags are synchronized using a unified time reference. This alignment mechanism transforms discrete, heterogeneous data streams into a structured input pool dataset with unified temporal logic, thus paving the way for subsequent expert agent cluster modules to detect cross-modal consistency conflicts, perform adaptive inference, and optimize narrative strategies.
[0055] S3: The expert agent cluster module performs parallel feature extraction on the collected data: the behavior analysis agent maps the operation logs to the Batu player type distribution; the language analysis agent quantifies cognitive load and emotional state by aligning text and emoji tags; the cultural context analysis agent combines user static profiles and dynamic behaviors to identify cultural comprehension barriers and risk levels in the narrative; finally, a preliminary analysis report is generated, which includes the Batu player type distribution output by the behavior analysis agent, the language profile with language depth and syntactic features output by the language analysis agent, and the cultural comprehension barriers and risk levels output by the literary context analysis agent.
[0056] S4: The central coordinating agent module summarizes the preliminary analysis reports of each expert agent cluster module, calculates the consistency of cross-modal semantics and determines whether there is a conflict; if the detected semantic conflict exceeds the preset threshold, the "reflection-reconstruction" loop is triggered, generating an upper-level control prompt to enter step S3 to guide the expert agents to re-evaluate the original data until a decision consensus is reached; if there is no conflict, the process proceeds directly to step S5.
[0057] Furthermore, in order to address semantic biases among expert agents in multimodal modeling and ensure the robustness of adaptive decisions, the reflection mechanism executed by the central coordinating agent module in step S4 includes a conflict assessment phase, a cue word construction phase, an inference chain reconstruction phase, and a decision fusion phase.
[0058] During the conflict assessment phase, the central coordinating agent module summarizes the analysis reports from behavioral, linguistic, and cultural expert agents and calculates the semantic conflicts between heterogeneous conclusions. When the inference results of any two expert agents show significant logical inconsistency and their semantic differences exceed a preset severity threshold, the system will automatically trigger a reflection loop to initiate cross-modal deep verification.
[0059] During the prompt word construction phase, the central coordinating agent module dynamically synthesizes a higher-level control prompt based on the detected specific conflict context. This higher-level control prompt not only includes the initial contradictory conclusions among the expert agents but also integrates relevant original interaction evidence and explicit re-evaluation instructions, thereby providing a precise environment for reflection guidance for the affected expert agents.
[0060] During the inference chain reconstruction phase, each expert agent analyzes the original multimodal input under the guidance of upper-level control prompts, thereby achieving a deeper and more accurate understanding of the user's intent.
[0061] In the decision fusion phase, the central coordinating agent module performs an aggregation operation on all iteratively revised expert output vectors, ultimately generating an adaptive narrative decision with robustness. This phase, through cross-agent deliberative negotiation, effectively eliminates inference biases that may arise from a single modality, ensuring that the generated narrative strategy maintains both plot coherence and accurate alignment with the user's cultural and cognitive state.
[0062] S5: Based on the decision-making consensus reached by the central coordinating agent module, generate personalized narrative suggestions; the suggestions include task allocation schemes customized for user player types, as well as explanatory guidance strategies and language style preference adjustment suggestions to address the cultural background gap of users;
[0063] S6: The personalized narrative generation module dynamically rewrites the in-game text scripts and task logic based on the generated personalized suggestions; the rewriting includes adjusting the depth of metaphors, adding historical background supplements, or modifying the narrative rhythm, so as to align the narrative context with the user's real-time cultural cognition.
[0064] Furthermore, in order to achieve tiered delivery of educational goals and precisely optimize learners' cross-cultural interaction experience, in step S6, the personalized narrative generation module includes a task allocation module, a narrative style mapping module, and an AI assistant adjustment module.
[0065] In the task allocation module, the system automatically allocates corresponding teaching tasks based on the probability distribution vector of Batu player types output by the behavior analysis agent.
[0066] In the narrative style mapping module, the system dynamically adjusts the presentation strategy of the narrative text based on the risk level assessed by the intelligent agent according to the cultural context analysis.
[0067] In the AI assistant adjustment module, the system dynamically fine-tunes the AI assistant's interactive tone and guidance depth based on the cognitive load and emotional state detected in real time by the language analysis agent. This adjustment process uses previous learning performance as background input through "vertical cross-stage adaptation," and coordinates with "horizontal real-time adjustment" to instantly optimize the current dialogue strategy. This aims to provide users with just the right amount of emotional support and intellectual guidance, thereby effectively reducing the frustration associated with cross-cultural learning while ensuring narrative coherence.
[0068] S7: Determine whether the current cross-cultural interaction phase has ended; if not, return to step S2 to continue the new round of data collection and dynamic adaptation; if it has ended, the cross-cultural narrative adaptation process is complete.
[0069] See Figure 2 According to another aspect of one or more embodiments of this disclosure, a cross-cultural narrative adaptive generation apparatus based on multimodal multi-agent collaborative reasoning is also provided, including a terminal device, the terminal device being an Internet terminal device, including a processor and a computer-readable storage medium, the processor being used to implement various instructions; the computer-readable storage medium being used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor for the cross-cultural narrative adaptive generation method based on multimodal multi-agent collaborative reasoning.
[0070] See Figure 3According to another aspect of one or more embodiments of this disclosure, a cross-cultural narrative adaptive generation system based on multimodal multi-agent collaborative reasoning is also provided, including a computer-readable storage medium storing a plurality of instructions and an agent workflow, the instructions being adapted to be loaded and executed by a processor of a terminal device. The agent workflow includes a multimodal data acquisition module, an expert agent cluster module comprising a behavior analysis agent, a language analysis agent, a cultural context analysis agent, a central coordinating agent module, and a personalized narrative generation module.
[0071] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code, including but not limited to disk storage, CD-ROM, optical storage, etc.
[0072] In the description of this invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0073] The above are merely embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A cross-cultural narrative adaptive generation method based on multimodal multi-agent collaborative reasoning, characterized in that, Includes the following steps: Step 1: Based on a large language model, construct a workflow for cross-cultural narrative adaptive multi-agent systems. The agents include a multimodal data acquisition module, an expert agent cluster module, a central coordinating agent module, and a personalized narrative generation module. The expert agent cluster module includes a behavior analysis agent, a language analysis agent, and a cultural context analysis agent. Step 2: The multimodal data acquisition module acquires and aligns the user's heterogeneous interaction signals in real time. The heterogeneous interaction signals include dialogue text, operation behavior logs, and facial expression labels recorded through a discrete sequence sampling strategy. The alignment aims to provide a unified semantic and temporal benchmark for subsequent analysis, ultimately forming an aligned multimodal dataset. Step 3: The expert agent cluster module performs parallel feature extraction on the collected data: the behavior analysis agent maps the operation logs to the Batu player type distribution; the language analysis agent quantifies cognitive load and emotional state by aligning text and emoji tags; the cultural context analysis agent combines user static profiles and dynamic behaviors to identify cultural comprehension barriers and risk levels in the narrative; finally, a preliminary analysis report is generated, which includes the Batu player type distribution output by the behavior analysis agent, the language profile including language depth and syntactic features output by the language analysis agent, and the cultural comprehension barriers and risk levels output by the literary context analysis agent; Step 4: The central coordinating agent module summarizes the preliminary analysis report of the expert agent cluster module, calculates the consistency of cross-modal semantics and determines whether there is a conflict; if the detected semantic conflict exceeds the preset threshold, the "reflection-reconstruction" loop is triggered, generating upper-level control prompts and entering Step 3 to guide the expert agent cluster module to re-evaluate the original data until a decision consensus is reached; if there is no conflict, proceed directly to Step 5. Step 5: Based on the decision consensus reached by the central coordinating agent module, generate personalized narrative suggestions; the suggestions include task allocation schemes customized for user player types, as well as explanatory guidance strategies and language style preference adjustment suggestions to address the cultural background gap of users; Step 6: The personalized narrative generation module dynamically rewrites the in-game text scripts and task logic based on the generated personalized narrative suggestions. The rewriting includes adjusting the depth of metaphors, adding historical background supplements, or modifying the narrative rhythm, so as to align the narrative context with the user's real-time cultural cognition. The method ends here.
2. The cross-cultural narrative adaptive generation method according to claim 1, characterized in that, The behavior analysis agent described in step S1 is used to map the multimodal information contained in discrete game behavior logs into user psychological motivation features, thereby realizing probabilistic reasoning of Batu player types. Language analysis agent: used to identify the user's metaphorical understanding depth and real-time cognitive load status through user input text; Cultural context analysis agent: used to assess the cognitive distance between the user's background and the target culture of the narrative, predict potential risks of cultural misunderstandings, and generate suggestions for bridging the gap.
3. The cross-cultural narrative adaptive generation method according to claim 1, characterized in that, In step S2, the formation of the multimodal dataset specifically includes: Step 2-1: Collect real-time behavioral trajectories during user interaction, including task time, failure rate, and interaction frequency; Step 2-2: Obtain the user's dialogue text sequence and simultaneously obtain the facial expression label sequence using a discrete sequence sampling strategy; the facial expression label sequence is obtained by sampling emotion labels every 5 seconds, and the emotion labels specifically include "confused", "happy", "nervous" and "neutral", without storing the original biometric image data; Steps 2-3: Align user interaction data spatiotemporally using a unified time base to construct a structured input pool dataset for subsequent semantic reasoning.
4. The cross-cultural narrative adaptive generation method according to claim 1, characterized in that, In step S4, the central coordinating intelligent agent module executes a reflection mechanism, specifically including: Step 4-1: The central coordinating agent module aggregates the output results of each expert agent and calculates the semantic conflict index; when the conclusions of any two expert agents are logically inconsistent and the significance exceeds the threshold, a round-robin reflection is triggered. Step 4-2: The central coordinating agent module constructs a prompt containing conflict descriptions, original evidence, and correction instructions, and sends it to the corresponding expert agents; Step 4-3: The expert agent rewrites its internal reasoning chain under the guidance of prompts; Step 4-4: The central coordinating agent module performs weighted fusion on the revised expert vectors to generate the final adaptation decision.
5. The cross-cultural narrative adaptive generation method according to claim 1, characterized in that, In step S6, the personalized narrative generation module specifically includes: Step 6-1: Task Allocation and Adaptation: Based on the Batu probability distribution vector, assign challenge tasks to achievement-oriented players and cultural unlocking and Easter egg guidance tasks to exploration-oriented players; Step 6-2: Narrative style adaptation: Dynamically adjust the metaphor density and explanatory narration in the narrative text according to the risk level of cultural misunderstanding; if users have significant cultural barriers, replace abstract historical metaphors with concrete narrative descriptions; Step 6-3: AI Assistant Tone Adaptation: Adjust the AI assistant's interactive tone and guidance depth based on the user's real-time cognitive load, and optimize the learning experience through "vertical cross-stage adaptation" and "horizontal real-time adjustment".
6. A cross-cultural narrative adaptive generation device based on multimodal multi-agent collaborative reasoning, characterized in that, The method includes a terminal device, which is an Internet terminal device, comprising a processor and a computer-readable storage medium. The processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, which are used by the processor to load and execute the cross-cultural narrative adaptive generation method based on multimodal multi-agent collaborative reasoning as described in any one of claims 1-5.
7. A cross-cultural narrative adaptive generation system based on multimodal multi-agent collaborative reasoning, characterized in that, The method includes a computer-readable storage medium storing multiple instructions and an agent workflow, the instructions being loaded and executed by a processor of a terminal device as described in any one of claims 1-5. The agent workflow includes a multimodal data acquisition module, and the expert agent cluster module specifically includes a behavior analysis agent, a language analysis agent, a cultural context analysis agent, a central coordinating agent module, and a personalized narrative generation module.