Interaction control method of foreign language culture learning system based on virtual reality

By introducing cultural interaction nodes and commitment tagging mechanisms into the virtual reality foreign language and culture learning system, the problem of unstable interactive control in existing technologies has been solved, enabling comprehensive perception and dynamic adjustment of learner behavior, thereby improving learning effectiveness and immersive experience.

CN121858003AInactive Publication Date: 2026-04-14SHIJIAZHUANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-30
Publication Date
2026-04-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing virtual reality foreign language and culture learning systems have shortcomings in terms of realistic cultural interaction modeling, learner behavior comprehension depth, and interaction result control logic. They are unable to identify the cultural pragmatic relationship between linguistic and non-linguistic behaviors, resulting in unstable interaction control and affecting learning outcomes.

Method used

Introducing cultural interaction nodes into a virtual reality environment involves collecting learners' foreign language and non-verbal interaction inputs to generate commitment actions. These commitment labels are then output through logical judgments to drive virtual character responses or scene branch changes. Dynamic adjustments are made by combining commitment label locking, window correction, debouncing merging, and adaptive cooldown periods to ensure the stability and consistency of the interaction.

Benefits of technology

It enables a comprehensive perception of learners' language, behavior, and context, identifies cultural risk situations, improves the authenticity and depth of foreign language and culture learning, enhances learning efficiency and user immersion experience, and ensures the seriousness and causal consistency of interaction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858003A_ABST
    Figure CN121858003A_ABST
Patent Text Reader

Abstract

The invention relates to an interaction control method of a foreign language culture learning system based on virtual reality, and the method is applied to a virtual reality terminal, and comprises the steps: building a foreign language culture learning scene in a virtual environment, and setting culture interaction nodes; when the terminal detects that a learner triggers a node, foreign language interaction input in a voice or text form and non-language interaction input including sight, gestures, postures, interaction distance and object delivery operation of the learner are collected; binding the two types of inputs in a preset time window to generate a commitment action; in combination with the scene attribute identifier and the social relationship attribute, carrying out logic judgment on the commitment action and outputting a commitment tag: outputting a prisoner risk tag when the directness feature reaches a first condition; outputting a dissociation risk tag when the gentality feature reaches a second condition; if the language is inconsistent with the behavior polite feature, outputting a culture inconsistent label; and finally, taking the commitment label as a subsequent interaction logic control entry, and driving virtual character response and scene branch change.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual reality and intelligent human-computer interaction technology, specifically to an interactive control method for a foreign language and culture learning system based on virtual reality. Background Technology

[0002] Current technologies, such as those published in CN108831218A and CN111028597A, show that the interactive control methods used in virtual reality-based foreign language and culture learning systems are still largely limited to scene restoration and perception enhancement. They exhibit significant shortcomings in realistic cultural interaction modeling, learner behavior comprehension depth, and interactive result control logic. The solution in CN108831218A aims to achieve spatial and perceptual synchronization between teachers and students in remote teaching through virtual classrooms, virtual characters, motion capture, and audio exchange modules. This solution focuses more on visualizing the teaching process and enhancing immersion, without distinguishing or analyzing the cultural pragmatic relationships between learners' linguistic and non-linguistic behaviors in the virtual environment. The system only treats voice, action, and location data as display or synchronization objects, lacking the ability to identify and judge the implicit cultural elements of politeness, power relations, and social distance in different cultural contexts. Consequently, the system cannot effectively respond to whether learners exhibit culturally inappropriate behavior, and its interactive control is essentially event-driven or process-driven, rather than culturally semantic-driven. Although the system event processing module in this method can map login, logout, hand-raising, and speaking operations to virtual actions, the triggering logic of these actions is discrete and rule-based, which cannot reflect the continuous behavioral changes and gradual attitude adjustments that learners often experience in real cultural communication. When learners attempt to correct their expressions multiple times within the same dialogue node, the system does not provide de-shaking merging, intent stability judgment, or interaction result locking mechanisms, which can easily cause frequent jumps in virtual character responses and affect learners' understanding of the causal relationship of cultural interaction.

[0003] The CN111028597A solution does not offer a systematic solution to the crucial issue of cultural pragmatic errors in foreign language and culture learning. The system does not perform structured analysis on whether learners' language use in AR or MR scenarios is appropriate or whether their behavior conforms to the social norms of the target culture. Instead, it relies on teachers or learners to perceive and correct these errors themselves, which is unlikely to have a substantial teaching effect in unaccompanied or self-directed learning scenarios. The interaction method in this solution mainly manifests as physical or visual interaction between learners and contextual elements. The system does not establish a mechanism for binding language input and non-verbal behavior on the same time scale, making it unable to determine whether there are cultural contradictions or conflicts between learners' language expression and body behavior. For example, situations where polite language is used but physical contact or rude actions are not recognized as cultural inconsistencies by the system, thus missing important opportunities for cultural learning feedback. The document generally lacks a system design for the stability of interaction results. Specifically, after a learner completes a key cultural interaction, the system does not lock or confirm the interaction conclusion in stages, but instead allows subsequent sporadic actions to continue affecting the system state. This design can easily lead learners to misunderstand the boundaries of cultural rules in foreign language and culture learning scenarios, making it difficult for them to distinguish which behaviors truly affect the interaction result and which are merely incidental. Furthermore, the solution generally does not consider the situation where learners attempt to correct their behavior multiple times in a short period. The system neither debouncing and merging multiple triggered behaviors nor using cooling mechanisms to suppress invalid or noisy input. This makes the interaction control logic prone to instability under complex real-world behaviors, hindering the construction of a repeatable and comprehensible cultural learning experience. Summary of the Invention

[0004] The purpose of this invention is to provide an interactive control method for a foreign language and culture learning system based on virtual reality, thereby addressing some of the drawbacks and shortcomings pointed out in the background art.

[0005] The present invention addresses the aforementioned technical problems by employing the following technical solution: an interactive control method for a virtual reality-based foreign language and culture learning system, applied to a virtual reality terminal, comprising: presenting a foreign language and culture learning scene in a virtual reality environment and setting up cultural interaction nodes; when the terminal detects that a learner has triggered the node, collecting foreign language interactive input and synchronous non-verbal interactive input, wherein the foreign language interactive input is voice or text, and the non-verbal interactive input includes gaze, gesture, posture, interaction distance, or object delivery operation;

[0006] Within a preset time window, the foreign language interactive input and non-language interactive input are bound together to generate a commitment action;

[0007] Based on scene attribute identifiers and social relationship attributes, the system makes logical judgments on commitment actions and outputs commitment labels. Specifically, when the directness feature meets the first condition, an offensive risk label is output; when the euphemism feature meets the second condition, an alienation risk label is output; and when the language politeness feature is inconsistent with the behavioral politeness feature, a cultural inconsistency label is output. The commitment label is used as the control entry point for subsequent interaction logic to drive virtual character responses or scene branch changes.

[0008] Furthermore, the conditions for satisfying the offense risk label include: the foreign language interaction input is imperative or direct rejection without buffering or euphemism; the non-verbal interaction input meets at least one of the following within a preset time window: the interaction distance is less than a first threshold, the body leans forward more than a first threshold, or the directional gesture lasts for more than a first duration; and the scene attribute is identified as formal or authoritative.

[0009] Furthermore, when outputting the cultural inconsistency label, cross-modal conflict detection is performed. When the foreign language interaction input contains at least one of the following: honorific titles, apologies, expressions of gratitude, or requests for buffering, and the non-verbal interaction input meets at least one of the following conditions within a preset time window: the interaction distance is less than a second threshold, the object delivery operation is throwing or quickly extending, the eye contact avoidance reaches a second duration, or the gesture amplitude exceeds a second threshold, the language politeness feature is determined to be inconsistent with the behavioral politeness feature, and a cultural inconsistency label is output.

[0010] Furthermore, when the commitment label is used as the entry point for subsequent interaction logic control, the commitment label is locked and a correction window is set; during the locking period, the virtual character response or scene branch is selected according to the commitment label; when the learner triggers the cultural interaction node again in the correction window and generates new foreign language interaction input and synchronous non-language interaction input, the binding and logic judgment are re-executed to update the commitment label.

[0011] Furthermore, the commitment tag is locked in two levels. When the learner triggers the cultural interaction node for the first time, it enters the first lock and opens the correction window. When the correction window detects that the delivery operation is completed or the learner leaves the node trigger area, it enters the second lock and closes the correction window. During the second lock, no further trigger input for updating the commitment tag is received.

[0012] Furthermore, during the locking period, when selecting a virtual character response or scene branch according to the commitment label, the virtual character response is added to the execution queue; when the commitment label in the correction window is updated and inconsistent with the original locking label, the unexecuted response in the queue is canceled, and the corresponding virtual character response or scene branch is reselected and replaced according to the updated commitment label.

[0013] Furthermore, when multiple triggers occur within the correction window, multiple sets of foreign language interactive inputs and synchronous non-language interactive inputs are debounced and merged. Only the input group that was last triggered and remained for a preset duration is bound and logically judged to update the commitment label. After the update, a cooldown period is set, and during the cooldown period, the update request is ignored.

[0014] Furthermore, the de-jitter merging includes validity screening. The binding and logical judgment to update the commitment label is only performed when the input group corresponding to the last trigger completes the collection of foreign language interactive input and synchronous non-verbal interactive input within a preset time window, and the non-verbal interactive input includes at least one of the following: a hand-carrying operation or an interaction distance change exceeding a threshold. Otherwise, the commitment label remains unchanged and enters a cooling-off period.

[0015] Furthermore, the cooling period is an adaptive cooling period, the duration of which is... Determine using the following formula:

[0016]

[0017] in, Indicates will Limit to range [ ]Inside; It is the natural logarithm function; This refers to the number of times the device is retried within the correction window; The maximum change in the interaction distance between the learner and the interactive object within the correction window or the input group corresponding to the last trigger; The duration of continuous stay at the cultural interaction node after the last trigger; Base cooling time; and These are the preset minimum and maximum cooling times, respectively; Preset weighting coefficients are used to make... Follow , and / or Increase as needed; during the cooling period, foreign language interactive input and synchronous non-verbal interactive input are allowed to continue to be collected, but the binding and logical judgment to update the commitment label are not performed.

[0018] Furthermore, when the validity screening is not satisfied and the commitment label remains unchanged, the terminal caches the input group corresponding to the last trigger during the cooling period; if in the first trigger after the cooling period ends, the newly collected foreign language interaction input and synchronous non-language interaction input match the cached input in terms of interaction object and trigger node, then the two are merged and the binding and logical judgment are performed to update the commitment label; otherwise, the cached input is discarded and the original commitment label is maintained.

[0019] The beneficial effects of this invention are as follows: By introducing cultural interaction nodes into a virtual reality environment and performing time-bound and joint judgment on learners' foreign language interactive input and synchronous non-verbal interactive input, a comprehensive perception of learners' language, behavior, context, and overall cultural interaction status is achieved. Compared to traditional foreign language learning methods that rely solely on linguistic correctness, this invention can identify cultural risk situations such as excessive directness, excessive euphemism, and inconsistencies between linguistic and behavioral politeness. It drives virtual character responses or scene branch changes through commitment tags, enabling learners to intuitively perceive the actual consequences of different cultural interaction methods in an immersive experience, effectively enhancing the authenticity and depth of foreign language and cultural understanding.

[0020] By employing commitment tag locking, correction windows, debouncing merging, adaptive cooldown periods, and input caching mechanisms, the system systematically controls and dynamically adjusts learners' multiple interactions, preventing interaction disorder or feedback distortion caused by frequent or invalid triggers, and ensuring the stability and continuity of the virtual reality cultural learning process. While allowing learners to make corrections and retry, the system can reasonably limit the update frequency of commitment tags, providing both error correction space and maintaining the seriousness and causal consistency of interaction results. This improves the interaction quality, learning efficiency, and user immersion experience of the foreign language and culture learning system. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating the foreign language cultural interaction logic judgment based on commitment actions in this invention.

[0022] Figure 2 This is a diagram illustrating the functional relationships of foreign language and cultural interactions based on commitment tags in this invention.

[0023] Figure 3 This is a graph showing the relationship between the multi-factor driven commitment tag adaptive cooling and cache re-evaluation of this invention.

[0024] Figure 4 This is a multi-dimensional radar chart illustrating the risk of offense in a formal meeting scenario in Embodiment 1 of the present invention.

[0025] Figure 5 This is a schematic diagram illustrating the cross-modal contradictory changes in non-verbal behavior before and after the correction in Embodiment 1 of the present invention.

[0026] Figure 6 This is a schematic diagram of the two-level locking and virtual character response queue timeline in Embodiment 1 of the present invention.

[0027] Figure 7 This is a bar chart comparing the effectiveness of the dwell time and distance changes when the corrected window is triggered multiple times in Embodiment 2 of the present invention.

[0028] Figure 8This is a visualization of the cumulative limit and formula for the adaptive cooling period Tcool in Embodiment 2 of the present invention.

[0029] Figure 9 This is a heatmap of the cached input and the first triggered input matching and feature comparison in Embodiment 2 of the present invention. Detailed Implementation

[0030] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0031] Combined with appendix Figure 1 This invention relates to an interactive control method for a virtual reality foreign language and culture learning system. When the virtual reality terminal detects that a learner has triggered a cultural interaction node, the system initiates an interactive data collection process. Triggering methods include the learner entering the virtual space area corresponding to the cultural interaction node or initiating an interactive operation with a virtual character associated with the node. After triggering, the terminal simultaneously collects the learner's foreign language interactive input and non-verbal interactive input. Foreign language interactive input includes foreign language expressions generated by the learner through speech or text. Non-verbal interactive input includes one or more of the following collected by the virtual reality terminal: gaze direction information, gesture information, body posture information, changes in the interaction distance between the learner and the virtual character, and information on the delivery of items, comprehensively recording the learner's behavioral performance during the foreign language and culture interaction process.

[0032] A preset time window is used to limit the effective association between language input and non-language behavior within the same cultural interaction event. By binding foreign language interactive input with non-language interactive input occurring within the same time window, the system generates a commitment action corresponding to that cultural interaction node. The commitment action represents the learner's comprehensive interactive stance in the current foreign language and cultural learning scenario.

[0033] Scene attribute identifiers are used to characterize the formality or authority of a cultural interaction scene. Social relationship attributes are used to characterize the type of role relationship between the learner and the virtual character. During logical judgment, the system analyzes the directness and euphemism features contained in the committed action. When the directness feature meets the preset first condition and matches the current scene attribute identifier, the system outputs an offense risk label. When the euphemism feature meets the preset second condition and matches the current social relationship attribute, the system outputs an alienation risk label. Simultaneously, the system performs consistency judgment on the linguistic politeness features reflected in foreign language interaction input and the behavioral politeness features reflected in non-linguistic interaction input. When the linguistic politeness features are inconsistent with the behavioral politeness features, the system outputs a cultural inconsistency label.

[0034] After outputting the commitment tag, the system uses it as the control entry point for subsequent interaction logic. Based on the commitment tag, the virtual reality terminal selects and controls the interaction response mode of the virtual character or the branching direction of the foreign language and culture learning scene.

[0035] Combined with appendix Figure 2 Foreign language interactive input refers to the speech or text content generated by learners at cultural interaction nodes. The system performs pragmatic feature analysis on the foreign language interactive input to identify whether it is an imperative expression or a direct refusal expression, and further determines whether the foreign language interactive input lacks buffering or euphemistic expressions to soften the tone. When the foreign language interactive input is identified as directly making a request or refusing the other party's request, and does not contain polite buffering components that conform to preset rules, the system determines that the foreign language interactive input meets the language conditions related to the risk of offense.

[0036] Nonverbal interaction input includes changes in the interaction distance between the learner and the virtual character, changes in body posture, and the duration of directional gestures. The system compares the actual interaction distance between the learner and the virtual character with a first threshold; when the interaction distance is less than the first threshold, spatial intrusion is identified. The system also calculates the learner's forward lean; when the forward lean exceeds a first threshold, an oppressive posture is identified. Simultaneously, the system monitors the duration of directional gestures; when the duration of a directional gesture exceeds a first duration, an action conveying accusation or command is identified. When at least one of the above nonverbal interaction conditions is met, the system determines that the committed action meets the offending risk-related behavioral conditions.

[0037] When both language and behavioral conditions are met, the system further considers the scenario attribute identifiers of the current foreign language and culture learning context for judgment. When the scenario attribute identifiers indicate that the current cultural interaction context is a formal or authoritative scenario, the system comprehensively determines that the learner's committed action carries an offensive risk in that context and outputs an offensive risk label accordingly.

[0038] Politeness features include the use of honorifics, apologies, expressions of gratitude, and polite requests. The system uses preset language feature rules to determine whether foreign language input contains at least one of these politeness components. When the system detects that a learner has used language forms that reflect respect or a softened tone in their foreign language expression, it classifies the input as having politeness features.

[0039] The system compares the interaction distance between the learner and the virtual character with a second threshold. When the interaction distance is less than the second threshold, it determines that there is excessive proximity. The system identifies the way objects are handed over; when the learner throws or quickly extends an object, it determines that the handing over behavior lacks cultural politeness. The system also analyzes the duration of the learner's gaze behavior; when the gaze avoidance duration reaches a second duration, it determines that there is avoidance behavior. Simultaneously, the system calculates the amplitude of hand gestures; when the gesture amplitude exceeds a second threshold, it determines that there is excessively exaggerated or intimidating behavior. When at least one of the above nonverbal interaction behaviors is satisfied, the system determines that the nonverbal interaction input lacks behavioral politeness characteristics.

[0040] When foreign language interactive input is determined to possess linguistic politeness features while non-linguistic interactive input is determined to lack behavioral politeness features, the system performs a cross-modal consistency judgment, identifying a contradiction between the linguistic and behavioral politeness features. Based on this judgment result, the system outputs a cultural inconsistency label, which is then used as the basis for subsequent interaction logic control.

[0041] When using commitment tags as the control entry point for subsequent interaction logic, a commitment tag locking and correction window mechanism is introduced to provide learners with limited correction opportunities while ensuring the stability of the interaction results. After the system completes the logical judgment based on the commitment action and outputs the commitment tag, the virtual reality terminal immediately locks the commitment tag and simultaneously opens the correction window corresponding to that cultural interaction node. The correction window limits the time range or interaction scope within which the commitment tag can be re-evaluated, preventing the commitment tag from being frequently changed in a short period and affecting the coherence of the virtual reality interaction process.

[0042] The virtual reality terminal selects the corresponding virtual character response method or branch of the foreign language and culture learning scenario based on the commitment label, and executes the relevant interactive content in a preset order. During this process, even if the learner continues to generate foreign language or non-verbal interactive input, the system maintains the current commitment label to ensure that the virtual character's behavioral feedback remains consistent with the established cultural interaction stance.

[0043] When a learner triggers a cultural interaction node again within the correction window, the system restarts the interaction acquisition and analysis process. The virtual reality terminal collects the learner's newly generated foreign language interaction input and its synchronous non-verbal interaction input, and performs binding processing on the input within a preset time window to generate a new commitment action. Subsequently, based on the scene attribute identifier and social relationship attributes of the current foreign language and culture learning scenario, the system re-executes logical judgment on the new commitment action to update the commitment label.

[0044] Once a learner triggers a cultural interaction node and completes the collection and analysis of foreign language interactive input and synchronous non-verbal interactive input, the system outputs a commitment label based on the committed action and immediately enters the first locked state, while simultaneously opening the correction window corresponding to the cultural interaction node. In the first locked state, the commitment label remains valid and is used to control subsequent interaction logic, but learners are still allowed to adjust their interactive behavior within the correction window by triggering the cultural interaction node again.

[0045] When the system detects that a learner has completed a handing-over operation or left the trigger area corresponding to a cultural interaction node, it determines that the current cultural interaction behavior is complete and switches the commitment label to the second locked state accordingly, while simultaneously closing the correction window. Once in the second locked state, the system no longer accepts further trigger inputs for updating the commitment label.

[0046] While the commitment tag is in either the first or second locked state, the system selects the virtual character's interaction response method or the branching path of the foreign language and culture learning scenario based on the current commitment tag. To avoid frequent interruptions to virtual character responses, the system adds the selected virtual character responses to the execution queue and schedules them for execution in a preset order. When the logical judgment is re-executed within the correction window and a new commitment tag is generated, and this new commitment tag is inconsistent with the original locked commitment tag, the system immediately cancels any unexecuted virtual character responses in the execution queue and reselects the corresponding virtual character response or scenario branch for replacement based on the updated commitment tag.

[0047] Combined with appendix Figure 3 When the system detects that a learner has repeatedly triggered the same cultural interaction node within the correction window, the virtual reality terminal manages the multiple sets of foreign language interaction inputs and their synchronized non-verbal interaction inputs in a unified manner, and performs de-jitter merging processing. De-jitter merging is used to filter input groups with actual cultural interaction significance from multiple triggering events, reducing redundant calculations.

[0048] During the de-jitter merging process, the system selects only the input group whose last trigger occurred and where the learner's continuous dwell time at the cultural interaction node reaches a preset duration as candidate input groups. For trigger events that do not meet the continuous dwell condition, the system does not perform commitment label update processing. For the selected candidate input groups, the system further performs validity screening to determine whether the input group meets the conditions for re-evaluating the commitment label. During validity screening, the system determines whether the candidate input group completed the collection of foreign language interactive input and synchronous non-verbal interactive input within a preset time window, and further determines whether the non-verbal interactive input includes object delivery operations or situations where the change in the interaction distance between the learner and the virtual character exceeds a preset threshold.

[0049] Only when the above validity screening conditions are met will the system re-execute the binding process for foreign language interactive input and non-verbal interactive input for the candidate input group, and perform logical judgment based on the new commitment action to update the commitment label. When the validity screening conditions are not met, the system retains the current commitment label and directly enters the cooldown period. After completing the commitment label update or keeping the original commitment label unchanged, the system sets a cooldown period. During the cooldown period, the system continues to allow the collection of learners' foreign language interactive input and synchronous non-verbal interactive input, but will no longer perform the binding and logical judgment operations for updating the commitment label.

[0050] The cooling-off period is set as an adaptive cooling-off period to dynamically adjust the update frequency of commitment labels, balancing system stability with the rationality of learner corrective behavior. The duration of the adaptive cooling-off period is denoted as: The value is calculated and determined in real time by the virtual reality terminal based on the learner's interactive behavior characteristics within the correction window, and the calculation method is as follows:

[0051]

[0052] in, This indicates that the input value will be displayed. Limited to the range inside, when Less than The time value is ,when Greater than The time value is This is to prevent the cooling period from being too short or too long, which could affect the system's user experience.

[0053] It is a natural logarithm function, used to mitigate the increase in the number of re-triggering cycles;

[0054] This indicates the number of times the cultural interaction node was triggered again within the correction window;

[0055] This represents the maximum change in the interaction distance between the learner and the virtual interactive object within the correction window or in the input group corresponding to the last trigger.

[0056] This indicates the duration a learner remains continuously at a cultural interaction node after the last triggering of that node.

[0057] This is a preset base cooling duration used to provide minimum cooling control when there are no significant disruptive behaviors.

[0058] and These are the preset minimum and maximum cooling times, respectively;

[0059] , , These are preset weighting coefficients used to unify the dimensions of various influencing factors and adjust the impact of re-triggering times, interaction distance changes, and dwell time on the cooldown period, so that the cooldown period varies with... , and / or It increases accordingly as the value increases.

[0060] When learners frequently trigger cultural interaction nodes, experience significant changes in interaction distance, or linger for extended periods within the correction window, the system extends the cooldown period to reduce the likelihood of frequent updates to commitment tags. Conversely, when learners' interaction behavior is relatively stable, the cooldown period is shortened accordingly to maintain the smoothness of the interaction process. During the cooldown period, the virtual reality terminal continues to collect learners' foreign language interaction input and synchronous non-verbal interaction input to record the learning process, but it does not perform the binding operation between foreign language interaction input and non-verbal interaction input, nor does it perform logical judgments to update commitment tags.

[0061] When the system performs validity filtering on the input group corresponding to the last trigger within the correction window and determines that the input group does not meet the conditions for updating the commitment label, the virtual reality terminal does not immediately discard the input group. Instead, it caches and saves the input group corresponding to the last trigger during the subsequent cooldown period. The cache is used to temporarily store input information that has significance for subsequent cultural interactions, so that supplementary judgments can be made in subsequent interactions.

[0062] After the cooldown period ends, when the learner triggers the cultural interaction node again for the first time, the system collects the newly generated foreign language interaction input and the synchronous non-language interaction input, and matches the collected new input with the cached input group. The matching judgment includes comparing the consistency between the new input and the cached input in terms of interaction object and trigger node to determine whether they point to the same cultural interaction context. When the system determines that the new input and the cached input are consistent in terms of interaction object and trigger node, the system merges the cached input with the newly collected input, and re-performs the binding operation on the merged foreign language interaction input and non-language interaction input within a preset time window to generate a new commitment action, and further performs logical judgment to update the commitment label.

[0063] When the system determines that the newly collected input does not match the cached input in terms of the interaction object or trigger node, the system considers that the cached input no longer has cultural interaction significance for further judgment and discards the cached input accordingly, while maintaining the current commitment label unchanged.

[0064] Example 1:

[0065] A company conference room scene was constructed in a virtual reality environment, such as... Figure 4 As shown, the meeting room is set up as a formal and authoritative cultural context, corresponding to a subordinate-to-superior reporting relationship between learners and virtual characters. The system sets up cultural interaction nodes in front of the conference table. When a learner enters this node area and faces the virtual character, the virtual reality terminal determines that the learner has triggered the cultural interaction node and begins collecting foreign language input and simultaneous non-verbal input.

[0066] Learners were asked to request changes to a project plan from a virtual character in a foreign language. The system collected the foreign language input in speech form, with the actual content being "Please change the schedule immediately." The system performed pragmatic analysis on this input, identifying it as an imperative request that lacked any softening or euphemistic expressions, such as "could you please" or "would it be possible." Therefore, the system determined that the input was highly direct and lacked softening features at the linguistic level.

[0067] Simultaneously, the system collects the learner's non-verbal interaction input within a preset time window. Using the spatial positioning module of the virtual reality terminal, the system calculates the interaction distance between the learner and the virtual character to be 0.6 meters, while the first threshold is preset to 1.0 meter, thus determining that the interaction distance is less than the safe threshold. Using the posture recognition module, the system detects that the learner's forward lean angle during speech is 18 degrees, while the corresponding threshold is 10 degrees, further determining that the forward lean significantly exceeds the safe range. The system also detects that the learner continuously points at the virtual character during speech, with the gesture lasting 2.3 seconds, while the first duration threshold is 1.5 seconds, constituting a clear directional reinforcement behavior.

[0068] exist Figure 4 In the multi-dimensional radar chart of offensive risk shown, the system maps verbal directness, spatial distance, body lean angle, and pointing gesture timing as risk dimensions, and performs normalization calculations based on a threshold safety boundary. The calculation results show that the risk coefficients for the above four items are approximately 1.5, 1.667, 1.8, and 1.533, respectively. After combining these with preset weights, the overall offensive risk score is approximately 1.62, significantly higher than the safety benchmark value of 1, indicating that the learner has significantly exceeded the safety boundaries of a formal meeting scenario in multiple dimensions of cultural interaction.

[0069] Since the current scenario is clearly identified as a formal meeting setting, and the virtual character is set as a superior with decision-making authority, the system confirms that this cultural interaction situation belongs to a formal or authoritative context. In this context, the system comprehensively determines that the learner's committed actions pose a significant risk of cultural offense in terms of linguistic directness, spatial distance, body posture, and gestures, and outputs an offense risk label. This commitment label serves as the control entry point for subsequent interaction logic. The virtual reality terminal then adjusts the virtual character's interactive response, making the virtual character's tone more indifferent, reducing explanatory information output, and guiding the dialogue in the scene branch towards a direction that requires re-clarification of the superior-subordinate relationship.

[0070] After this round of interaction, the learner triggers the same cultural interaction node again and generates new foreign language interaction input. The foreign language interaction input collected by the system is still in speech form, and the actual content is "I am sorry for my toneearlier and thank you for your time." The system performs language politeness feature recognition on this foreign language interaction input and detects that it contains both apology and thanks, thus satisfying the language politeness feature conditions.

[0071] Meanwhile, the system collects corresponding non-verbal interaction inputs within a preset time window. Using spatial positioning and motion capture modules, the system calculates the interaction distance between the learner and the virtual character to be 0.7 meters, while the second threshold is set at 1.2 meters, indicating the interaction distance is still too close. The system further detects that the learner quickly extends their arm to hand a document to the virtual character during their speech, with the action taking 0.4 seconds, below the preset stable handing time threshold of 1.0 second, thus classifying it as a rapid handing operation. Simultaneously, the system detects that the learner exhibits eye-avoidance behavior during their speech, lasting 2.1 seconds, while the second duration threshold is 1.5 seconds, meeting the eye-avoidance condition.

[0072] Based on the above data collection, the system performs cross-modal conflict detection and combines it with... Figure 5 The diagram illustrating cross-modal behavioral changes was used for quantitative analysis. The results showed that although learners expressed clear polite intentions through apologies and thanks at the linguistic level, their nonverbal behaviors, particularly in terms of interaction distance, delivery methods, and eye contact, still did not meet the requirements of formal cultural contexts. Based on this, the system identified a significant inconsistency between linguistic and behavioral politeness features and output a cultural inconsistency label.

[0073] After outputting the cultural inconsistency label, the system immediately uses the commitment label as the control entry point for subsequent interaction logic, locks the commitment label, and opens a correction window. During the lockout period, the virtual reality terminal selects the corresponding virtual character response method based on the cultural inconsistency label, causing the virtual character to exhibit an interactive state of cautiously receiving information while maintaining psychological distance.

[0074] During the correction window, the learner became aware of the problem and triggered the same cultural interaction node again. The system re-collected new foreign language interaction input and synchronous non-verbal interaction input. The new round of data collection showed that the learner's foreign language expression was "Could we continue the discussion when convenient for you," which included a request for buffering. Non-verbal interaction input showed that the learner's interaction distance with the virtual character was adjusted to 1.5 meters, which is higher than the second threshold; the object delivery operation was changed to a slow, horizontal delivery with a delivery time of 1.8 seconds; the eye contact avoidance time was reduced to 0.6 seconds; and the gesture amplitude was also lower than the second threshold. The system completed the binding process and regenerated the commitment action within the preset time window.

[0075] Combination Figure 5 Based on the calculation results of the behavioral changes before and after the correction, the system determined that all three original non-verbal risk conditions had been eliminated, the verbal politeness characteristics were consistent with the behavioral politeness characteristics, and the conditions for offense risk or alienation risk were not met. Therefore, the commitment label was updated to a risk-free commitment label. The system then canceled the original pending virtual character response and switched to a new scene branch based on the updated commitment label, allowing the virtual character to naturally receive the file and continue discussing the project plan.

[0076] In the subsequent interaction phase, the learner stands in front of the conference table again. The virtual reality system detects that the learner has triggered the cultural interaction node for the first time and immediately enters the first locked state based on the current risk-free commitment label, while simultaneously opening a correction window, such as... Figure 6 As shown. The duration of the correction window is set to 6 seconds to allow learners to make final adjustments to their interactions if necessary.

[0077] In the first locked state, the system selects the virtual character's interactive response based on the current risk-free commitment label and adds the corresponding response to the execution queue. This execution queue contains three response actions in sequence: the virtual character receiving a file, nodding in agreement, and continuing the project discussion. The total execution time is estimated to be 4.5 seconds.

[0078] During the correction window, the system continuously monitors the learner's behavior. After 3.2 seconds, the system detects that the learner has completed the object-handing operation, using both hands to hand the object horizontally, with a completion time of 1.9 seconds, which meets cultural etiquette requirements. Based on this, the system determines that the learner has completed the key interactive action and further detects that the learner subsequently moves backward 0.8 meters, leaving the trigger area of ​​the cultural interaction node. At this point, the system enters a second locking state and closes the correction window.

[0079] Upon entering the second locked state, the system explicitly commits that the tag will no longer accept any further triggering input for updates. Even if the learner continues to generate foreign language or non-language interactive input near this node, the system will not re-execute binding and logical judgment operations, ensuring the determinism and causal integrity of the interaction results. In this state, the system begins to execute the virtual character responses in the pending queue sequentially.

[0080] exist Figure 6 In the other branch scenario shown, the learner triggers the cultural interaction node again while in the first locked state and before the correction window is closed, generating new foreign language interactive input and synchronous non-verbal interactive input. After the system re-executes the binding and logical judgment, it determines that the new commitment label is an alienation risk label, and this label is inconsistent with the original locked risk-free commitment label. At this time, the system immediately cancels the virtual character responses that have not yet been executed in the pending execution queue, retaining only the initial nod that has been completed, and reselects the corresponding virtual character responses based on the updated alienation risk label, adding response actions such as delayed response, reduced interaction frequency, and early termination of dialogue to a new pending execution queue.

[0081] Example 2:

[0082] Based on Example 1, the system detected a total of 4 re-triggering actions within the 6 seconds that the correction window lasted. Figure 7 As shown, Figure 7 The bar chart displays the changes in dwell time and interaction distance for each trigger within the correction window, and uses threshold lines to indicate the dwell time threshold of 1.5 seconds and the distance change threshold of 0.5 meters, to visually reflect whether each trigger meets the validity screening criteria.

[0083] The first trigger occurred 0.8 seconds after the correction window opened. The learner simply whispered "thank you" with a brief hand gesture, pausing for 0.6 seconds. No object was handed over, and the interaction distance changed by 0.1 meters, below the preset distance change threshold of 0.5 meters. The second trigger occurred at 1.6 seconds. The learner adjusted their posture and briefly moved forward, pausing for 0.9 seconds. Although a distance change of 0.4 meters was detected, it still did not meet the threshold requirement. The third trigger occurred at 2.4 seconds. The learner again sent a brief supplementary explanation, pausing for 1.1 seconds, but without any explicit object handing over or significant spatial change. Figure 7 As can be seen, the three triggers mentioned above were all below the 1.5-second threshold in terms of dwell time, and did not meet the conditions of object delivery or distance change exceeding the threshold in terms of non-verbal behavior. Therefore, the system uniformly marked them as candidate inputs, which were only used for recording but not immediately executing the binding and logical judgment of foreign language and non-verbal inputs.

[0084] The fourth trigger occurred at 3.2 seconds into the correction window. At this point, the learner was standing stably within the cultural interaction node, remaining there continuously for 2.2 seconds, and again handed the document to the virtual character with both hands, completing the handover action in 1.7 seconds. Simultaneously, the system detected that the interaction distance between the learner and the virtual character decreased from 1.5 meters to 0.9 meters, a change of 0.6 meters, exceeding the preset threshold of 0.5 meters. Figure 7 As indicated by the red highlight, this trigger significantly exceeded the threshold requirements in both the dwell time and distance variation dimensions, and was accompanied by clear object delivery behavior. Based on this, the system determined that the input group corresponding to this trigger met the validity conditions of dwell time and non-verbal behavior.

[0085] Based on the above detection results, the system performs debouncing and merging processing on the four sets of foreign language interactive inputs and synchronous non-verbal interactive inputs collected within the correction window. Only the input group that was triggered last and whose dwell time reached the preset duration is selected as the valid input group, and the system performs a binding operation between the foreign language interactive input and the non-verbal interactive input on this input group. Subsequently, the system re-executes logical judgment based on the current scene attribute identifier and social relationship attribute to confirm that the commitment action does not introduce new risks of offense or alienation, and maintains the original risk-free commitment label unchanged.

[0086] After completing the above update or confirmation, the system immediately sets a cooldown period to suppress repeated triggering that could interfere with the interaction logic within a short period. The initial duration of the cooldown period is set to 4 seconds. During this period, the system still allows the collection of learners' foreign language interaction input and synchronous non-verbal interaction input, but explicitly prohibits binding and logical judgments to update commitment tags. The actual duration of this cooldown period will be determined by subsequent adaptive cooldown algorithms. Even if the learner whispers or makes small movements again during the cooldown period, the system will only record these as raw interaction data and will not trigger further changes to the commitment tag. Two seconds before the end of the cooldown period, if the learner attempts to provide supplementary information to the virtual character, the system will ignore this update request, maintain the current commitment tag, and continue with the established virtual character response process, as it is still within the cooldown period. Only after the cooldown period ends will the system reopen the commitment tag update entry, providing a new basis for judgment for the next round of cultural interaction.

[0087] The system employs an adaptive cooling period calculation method, dynamically determining the cooling duration based on the actual interaction intensity within the correction window. For example... Figure 8 As shown, Figure 8 The calculation process for the adaptive cooling period is presented in a visual format using itemized cumulative calculations and amplitude limiting. The duration of the cooling period is denoted as... The calculation formula is as follows:

[0088]

[0089] in, Indicates the numerical value The limit is within the range Inside; It is a natural logarithm function, used to gently modulate the increase in the number of re-triggering cycles; To correct the number of re-triggers detected within the window; To correct the maximum change in the interaction distance between the learner and the virtual character within the window or in the last triggered input group; The duration of continuous stay of the learner within the cultural interaction node after the last trigger; Base cooling time; and These are the preset minimum and maximum cooling times, respectively; , , These are preset weighting coefficients.

[0090] Base cooling time Set to 2 seconds, minimum cooldown time. Set to 3 seconds, maximum cooldown time. Set to 10 seconds. In the weighting coefficients, Set to 1.2 to characterize the instability caused by multiple triggers. Set to 2.0 to emphasize the impact of drastic changes in spatial distance on interaction stability. Set to 0.8 to reflect the intensity of the interaction intent represented by the learner's continued dwell time.

[0091] Within this correction window, the system detected a total of This is the second trigger. Calculations using the spatial positioning module show that the maximum change in interaction distance occurs during the last trigger, decreasing from 1.5 meters to 0.9 meters. Therefore:

[0092]

[0093] Meanwhile, the system detected that the learner's continuous and stable dwell time within the cultural interaction node during the last trigger was:

[0094]

[0095] Substituting the above parameters into the cooling time calculation formula, we get:

[0096]

[0097]

[0098]

[0099]

[0100] The cooling time before limiting is:

[0101]

[0102] Because it satisfies:

[0103]

[0104] Therefore, the final adaptive cooling period is determined as follows:

[0105]

[0106] After the current interaction ends, the system enters a cooldown period of approximately 6.9 seconds. During this cooldown period, the virtual reality terminal continues to collect the learner's foreign language interaction input and synchronous non-verbal interaction input for subsequent behavior analysis and learning trajectory recording. However, the system explicitly stops binding foreign language and non-verbal inputs and does not perform logical judgments to update commitment labels, ensuring the stability of the interaction results.

[0107] During the cooling period, the system also incorporates an input buffering and matching merging mechanism, the working principle of which is as follows: Figure 9 As shown. Figure 9 The normalized feature strengths of the cached input group and the first-triggered input group after the cooling period are compared and displayed in the form of heatmaps in terms of dwell time, interaction distance change, interaction node consistency and interaction object consistency.

[0108] The system detected that the foreign language input for this trigger was in speech form, containing the phrase "I just wanted to add one more point." While this statement conveyed a sense of supplementary explanation, it lacked a request for buffering or explicit politeness markers. Simultaneously detected non-verbal input showed that the learner only adjusted their position by 0.2 meters, did not perform any object delivery, and remained in the position for 0.8 seconds, less than the preset minimum effective dwell time of 1.5 seconds. Based on this, the system determined that the validity screening was insufficient, but instead of discarding the input group directly, it stored it as cached input during a cooldown period.

[0109] 0.4 seconds after the cooldown period ended, the system detected that the learner triggered the cultural interaction node again, generating new foreign language interaction input and synchronous non-verbal interaction input. The latest data collection showed that the learner's foreign language interaction input was "Could I add one more point regarding the timeline," containing a clear request. The non-verbal interaction input showed a change in interaction distance of 0.6 meters and a continuous dwell time of 1.9 seconds, meeting the validity screening criteria. The system matched the new input with the cached input, confirming that both inputs were consistent in terms of the interaction object and the triggering node, belonging to the same cultural interaction context.

[0110] like Figure 9 As shown, the normalization intensity of the first-triggered input group after cooling in the heatmap is significantly higher than that of the cached input group in terms of dwell time and distance change dimensions, and both are set to 1 in the node matching and object matching dimensions. Based on this, the system merges the two input groups, retaining the supplementary explanations and initial intent reflected in the cached input, and using the more complete language and behavioral data in the current input as the main factor to form a new merged input group.

[0111] Subsequently, the system binds the merged foreign language and non-verbal inputs and re-executes logical judgments based on the current scenario attribute identifier and social relationship attribute. The judgment results show that the merged commitment action has the characteristics of request buffering and polite expression at the linguistic level, and meets the requirements of reasonable interaction distance and stable stay at the behavioral level. It does not trigger the conditions of offensive risk or alienation risk. Therefore, the system maintains the risk-free commitment label and allows subsequent interaction logic to continue.

[0112] In another comparative scenario, if the first trigger after the cooldown ends occurs on a different interactive object or a different cultural interactive node, the system will determine that the new input does not match the cached input and immediately discard the cached input, making a judgment based only on the current input or continuing to maintain the original commitment label, thereby effectively avoiding semantic confusion across scenarios or objects.

[0113] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. An interactive control method for a foreign language and culture learning system based on virtual reality, characterized in that, Applied to virtual reality terminals, the method includes: presenting a foreign language and culture learning scenario in a virtual reality environment and setting up cultural interaction nodes; when the terminal detects that the learner triggers the node, it collects foreign language interaction input and synchronous non-verbal interaction input, wherein the foreign language interaction input is voice or text, and the non-verbal interaction input includes gaze, gesture, posture, interaction distance or object delivery operation; Within a preset time window, the foreign language interactive input and non-language interactive input are bound together to generate a commitment action; Based on scene attribute identifiers and social relationship attributes, the system makes logical judgments on commitment actions and outputs commitment labels. Specifically, when the directness feature meets the first condition, an offensive risk label is output; when the euphemism feature meets the second condition, an alienation risk label is output; and when the language politeness feature is inconsistent with the behavioral politeness feature, a cultural inconsistency label is output. The commitment label is used as the control entry point for subsequent interaction logic to drive virtual character responses or scene branch changes.

2. The interactive control method of the virtual reality-based foreign language and culture learning system according to claim 1, characterized in that... The conditions for meeting the offense risk label include: the foreign language interaction input is commanding or a direct rejection without buffering or euphemism; the non-verbal interaction input meets at least one of the following conditions within the preset time window: the interaction distance is less than a first threshold, the body leans forward more than a first threshold, or the directional gesture lasts for more than a first duration; and the scene attribute is marked as formal or authoritative.

3. The interactive control method of the virtual reality-based foreign language and culture learning system according to claim 1, characterized in that... When outputting a cultural inconsistency label, cross-modal conflict detection is performed. When the foreign language interaction input contains at least one of the following: honorific titles, apologies, expressions of gratitude, or requests for buffering, and the non-verbal interaction input meets at least one of the following conditions within a preset time window: the interaction distance is less than a second threshold, the object delivery operation is throwing or quickly extending, the eye contact avoidance reaches a second duration, or the gesture amplitude exceeds a second threshold, the language politeness feature is determined to be inconsistent with the behavioral politeness feature and a cultural inconsistency label is output.

4. The interactive control method of the virtual reality-based foreign language and culture learning system according to claim 1, characterized in that... When the commitment label is used as the entry point for subsequent interaction logic control, the commitment label is locked and a correction window is set; during the locking period, the virtual character response or scene branch is selected by the commitment label; When the learner in the correction window triggers the cultural interaction node again and generates new foreign language interaction input and synchronous non-language interaction input, the binding and logical judgment are re-executed to update the commitment label.

5. The interactive control method of the virtual reality-based foreign language and culture learning system according to claim 4, characterized in that... The commitment tag is locked in two levels. When the learner triggers the cultural interaction node for the first time, it enters the first lock and opens the correction window. When the correction window detects that the delivery operation is completed or the learner leaves the node trigger area, it enters the second lock and closes the correction window. During the second lock, no further trigger input for updating the commitment tag will be received.

6. The interactive control method of the virtual reality-based foreign language and culture learning system according to claim 4, characterized in that... When selecting a virtual character response or scene branch according to the commitment label during the locking period, the virtual character response is added to the execution queue; When the commitment label in the correction window is updated and inconsistent with the original locked label, cancel the unexecuted responses in the queue, and reselect and replace the corresponding virtual character response or scene branch according to the updated commitment label.

7. The interactive control method for a virtual reality-based foreign language and culture learning system according to claim 4, characterized in that... When multiple triggers occur within the correction window, debouncing and merging are performed on multiple groups of foreign language interactive inputs and synchronous non-language interactive inputs. Only the input group that was last triggered and remained for a preset duration is bound and logically judged to update the commitment label. After the update, a cooldown period is set, and subsequent update requests are ignored during the cooldown period.

8. The interactive control method for a virtual reality-based foreign language and culture learning system according to claim 7, characterized in that... The de-jitter merging includes validity screening. The binding and logical judgment to update the commitment label is only performed when the input group corresponding to the last trigger completes the collection of foreign language interactive input and synchronous non-verbal interactive input within a preset time window, and the non-verbal interactive input includes at least one of the following: handing over an object or the change in interaction distance exceeds a threshold. Otherwise, the commitment label remains unchanged and enters the cooling-off period.

9. The interactive control method of the virtual reality-based foreign language and culture learning system according to claim 7, characterized in that... The cooling period is an adaptive cooling period, the duration of which increases with the number of times it is triggered again within the correction window; during the cooling period, foreign language interactive input and synchronous non-verbal interactive input are allowed to continue to be collected, but the binding and logical judgment to update the commitment label are not performed.

10. The interactive control method of the virtual reality-based foreign language and culture learning system according to claim 8, characterized in that... When the validity screening is not satisfied and the commitment label remains unchanged, the terminal caches the input group corresponding to the last trigger during the cooling period; if in the first trigger after the cooling period ends, the newly collected foreign language interaction input and synchronous non-language interaction input match the cached input in terms of interaction object and trigger node, then the two are merged and the binding and logic judgment are performed to update the commitment label; otherwise, the cached input is discarded and the original commitment label is maintained.

Citation Information

Patent Citations

  • Remote teaching system based on virtual reality

    CN108831218A

  • Mixed reality foreign language scene, environment and teaching aid teaching system and method

    CN111028597A