Emotion parameter-driven autonomous approach and hug interaction control system for virtual characters in virtual space

JP7900630B1Active Publication Date: 2026-08-04佐藤 景虎
View PDF 13 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
佐藤 景虎
Filing Date
2026-04-09
Publication Date
2026-08-04

AI Technical Summary

Benefits of technology

【0023】 本発明によれば、体調が優れないキャラクターが慰めを求めて近寄ってくるといった、現実の人間関係に近い自発的·感情的インタラクションが実現される。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007900630000001_ABST
    Figure 0007900630000001_ABST
Patent Text Reader

Abstract

This VR experience integrates character-driven, emotionally autonomous approach and embrace experiences with natural, non-forceful re-guidance when the story deviates from the narrative. [Solution] An information processing method and system that integrates multidimensional internal state parameter composite evaluation value-driven autonomous approach control, four-zone distance-linked AI voice dialogue, hug sequence, multimodal rejection detection, as well as three types of deviation detection (gaze deviation, speech deviation, and no response) and a three-stage escalation type natural return guidance (gentle verbal prompting → curiosity / empathy induction → approach-type re-guidance) according to the duration of the deviation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of interactive experience technology in a virtual reality (VR) space. In particular, triggered by changes in multi-dimensional internal emotional parameters possessed by a virtual character, the virtual character autonomously approaches the user independently of the user's operation input, performs a hugging interaction while conducting real-time AI voice dialogue using a large language model (LLM), and when the user deviates from the expected progression of the story, the virtual character re-induces the user into the story in a natural rather than forced manner. The present invention relates to an integrated control system and method therefor.

[0002] The story natural return induction mechanism of the present invention is technically distinct from a rejection detection mechanism (responding to physical and vocal rejections). Rejection is a response to the state where "the user intentionally rejects contact with the character", whereas deviation return is a gentle re-induction by the character when "the user's attention has wandered away from the story", and its technical feature is a natural pull-back that utilizes curiosity, empathy, and intimacy rather than forced control.

Background Art

[0003] In recent years, with the rapid development and popularization of virtual reality (VR) technology using a head-mounted display (HMD), an interactive experience with virtual characters in a virtual space is being commercially developed. In particular, VR content themed on characters from animation video works and games has gained a certain demand in the market, and improving its immersion and bidirectionality has become an important industrial issue.

[0004] One type of prior art is a system in which a character reacts and activates voice input when the direction of the user's gaze matches the direction of the character (see Patent Document 1). This system uses an eye-tracking sensor to acquire the user's gaze vector, and the character notices and reacts when the angle difference with the virtual character's position vector falls below a predetermined value. However, this technology is merely a passive reaction of the character to the user's gaze, and does not realize the character actively approaching the user based on changes in its own internal state, nor does it provide natural guidance to return the user to the story if they deviate from it.

[0005] A second type of conventional technology is known in which a character expresses emotion when the distance between the character and the user in a virtual space falls below a certain level (see Patent Document 2). This technology only allows the character to react when the user approaches the virtual character on their own, and does not support actions such as the character walking towards the user on their own, or natural re-guiding actions by the character when the story deviates.

[0006] A third type of prior art is known to be technology related to the movement control of non-player characters (NPCs) in games (see Patent Document 3). This technology discloses a method for NPCs to move on a map based on a pathfinding algorithm in the context of a game, but it does not disclose the activation of approach behavior based on a composite evaluation value of the character's internal emotional state, integration with distance-linked AI voice, or story deviation detection and gradual natural return guidance.

[0007] A fourth type of prior art has been proposed: an AI character control system using a large language model (see Patent Document 4). This technology dynamically generates the character's speech content using an LLM, but it does not disclose how to control the character's physical movement, how to change voice parameters in conjunction with distance changes, how to integrate with hug animation sequences, or how to guide a gradual return to a natural story when the story deviates.

[0008] A fifth type of prior art is known to be a technology for reproducing the feeling of being hugged using a haptic feedback device for VR (see Patent Document 5). This technology uses a dedicated device to provide the user with a sense of physical pressure, but it has high dissemination costs because it requires additional hardware, and it does not address spontaneous approaching actions from the character or re-guidance control when the story deviates.

[0009] None of the conventional technologies mentioned above have a proactive deviance recovery mechanism that "naturally brings back users who have strayed from the story through character-driven means." As a result, even when users become distracted, the content continues to progress unilaterally, creating a discontinuity in the user experience. To solve this problem, the present invention provides a three-stage escalation type natural recovery guidance (minor verbal prompting → curiosity stimulation → approach-type re-induction) that corresponds to the duration of the deviance. [Prior art documents] [Patent Documents]

[0010] [Patent Document 1] U.S. Patent No. 9250703

[0011] [Patent Document 2] U.S. Patent No. 10668382

[0012] [Patent Document 3] Japanese Patent Publication No. 2015-024159

[0013] [Patent Document 4] U.S. Patent Application Publication No. 2023 / 0351118

[0014] [Patent Document 5] U.S. Patent No. 11231781 [Overview of the project] [Problems that the invention aims to solve]

[0015] The primary problem that this invention aims to solve is the realization of a proactive system in which a virtual character proactively and autonomously initiates actions to approach the user based on its own internal state parameters. In order to realize emotional interactions that are close to real human relationships, such as a character who is feeling unwell approaching for comfort or a tired character asking for support, a mechanism is needed to constantly manage the character's internal state and for changes in that state to trigger spontaneous actions.

[0016] The second problem that this invention aims to solve is the calculation of a composite evaluation value corresponding to multidimensional internal state parameters and ensuring the reliability of approach activation by determining its threshold.

[0017] The third problem that this invention aims to solve is to realize integrated control that simultaneously and seamlessly performs autonomous movement of a virtual character and AI voice dialogue, and changes the content, volume, and language style of the voice dialogue in real time in a multidimensional manner according to changes in distance during movement.

[0018] The fourth problem that this invention aims to solve is to realize a mechanism in which, when a user deviates from the expected progression of the story, the virtual character uses curiosity, empathy, and intimacy to naturally guide the user back to the story, rather than using coercive control. The core of this problem is to ensure an experience quality in which the user feels "the character cared about them" rather than "they were pulled back to the story." Through a three-stage escalation (minor prompting → eliciting curiosity and empathy → approach-based re-guidance) depending on the duration of the deviation, appropriate intensity of re-guidance according to the severity of the deviation is achieved. [Means for solving the problem]

[0019] The first aspect of the present invention solves the first and second problems by managing a plurality of internal state parameters, including a group of physical condition parameters (health level H, fatigue level F) and a group of environmental parameters (external stress index S, external event flag E), at a 500-millisecond cycle, calculating a weighted sum V = wH × (1-H) + wF × F + wS × S + wE × E, and initiating autonomous movement when a first threshold θ1 (default 0.65) is exceeded.

[0020] A second aspect of the present invention solves the third problem by constructing multi-layered contextual prompts to the LLM simultaneously with the start of autonomous movement and realizing AI voice dialogue control that automatically changes volume, speech frequency, and language style for each distance zone (D>3.0m: Zone 1, 1.5~3.0m: Zone 2, 0.5~1.5m: Zone 3, D≦0.5m: Zone 4).

[0021] A third aspect of the present invention solves the fourth problem by having the story deviation detection unit 95 monitor three types of deviation states (gaze deviation, speech deviation, and no response) in parallel, and having the natural return guidance unit 96 execute a three-stage escalation-type guidance according to the duration of the deviation. In the first stage (deviation 10-30 seconds), the LLM generates a light verbal utterance that is in line with the current story context. In the second stage (deviation 30-60 seconds), the LLM generates an utterance that promotes empathy for unresolved mysteries in the story, anticipation for future developments, and the emotional state of the character. In the third stage (deviation exceeding 60 seconds), an autonomous movement step is activated to physically bring the character closer to the user while generating a gentle re-guidance utterance. Each stage is implemented not as a forced restart of story progression but as a natural expression of the character's emotions, and is designed so that the user does not feel that control has been taken away.

[0022] The fourth aspect of the present invention is to independently provide a physical rejection detection unit 81 (IMU 90Hz, backward speed exceeding 15 cm / second or head rotation angular velocity exceeding 50° / second) and an audio rejection detection unit 82 (keyword list matching and LLM intention analysis parallel processing), and immediately stop upon detection by either, increase the emotion parameter Sd by +0.2, and output the corresponding animation and speech, thereby ensuring a safe and comfortable experience. Note that rejection detection and deviation detection are independent processing systems, and when rejection is detected, rejection handling takes precedence over deviation recovery handling.

Effects of the Invention

[0023] According to the present invention, spontaneous and emotional interactions similar to real human relationships are realized, such as a character with poor physical condition approaching seeking comfort.

[0024] According to the present invention, by integrating seamless AI voice dialogue from autonomous movement to hugging, the experience is given emotional depth and narrativity.

[0025] According to the present invention, the story natural return induction mechanism provides the experience of "the character cares about me" even when the user's attention is distracted from the story, and promotes immersion in the story without a sense of coercion. With a three-stage escalation design according to the deviation duration, it can handle from minor inattention to long-term detachment in a step-by-step and appropriate intensity.

[0026] According to the present invention, the multi-modal rejection control mechanism immediately stops the approach of the character when the user does not desire it, thus ensuring a safe and comfortable experience. By handling different nature states of rejection and deviation with independent processing systems, optimal responses for each are realized.

Brief Description of the Drawings

[0027] [Figure 1] System overall configuration block diagram in an embodiment of the present invention. It shows the data flow between all processing units including the deviation detection unit (95) and the natural return induction unit (96). [Figure 2] Detailed flowchart of internal state parameter management and composite evaluation value calculation process. [Figure 3] Overall flowchart of the autonomous movement control unit and the embrace activation process. Shows coordination with rejection detection and deviation detection. [Figure 4] Processing flowchart for distance-linked AI voice interaction control. Includes details of four-zone control. [Figure 5] A flowchart for multimodal rejection detection, emotion reflection, story deviation detection, and three-stage natural return induction processing. [Figure 6] A flowchart illustrating the processing steps for the voluntary approach behavior detection algorithm. It shows the integrated decision-making process using IMU data and eye-tracking data. [Figure 7] Detailed configuration diagram of distance and velocity sensing method. Shows integrated processing of inside-out tracking and IMU. [Figure 8] Detailed diagram of the specific processing flow in the scenario of deteriorating health. It shows the series of processes from the calculation of the composite evaluation value V to the activation of autonomous movement. [Figure 9] A diagram illustrating the criteria for judgment and escalation processing in multi-stage control of deviation recovery. [Figure 10] Detailed configuration diagram of the haptic feedback implementation. It shows the control parameters for the LRA vibration motor and the multi-point force feedback device. [Modes for carrying out the invention]

[0028] Embodiments of the present invention will be described in detail below with reference to the drawings. In this embodiment, a VR interactive experience application using animated characters will be described as a specific example, but the scope of application of the present invention is not limited thereto.

[0029] <Overall System Configuration> As shown in Figure 1, the VR interaction system 10 of this embodiment comprises an HMD 20, an internal state management unit 30, a composite evaluation value calculation unit 40, an autonomous movement control unit 50, an AI dialogue control unit 60, a hug sequence control unit 70, a rejection detection unit 80, an emotion reflection unit 90, a story deviation detection unit 95, a natural return guidance unit 96, and a cloud server 100 as its main components. These achieve low-latency communication by combining WebSocket (wss: / / ) and Cloudflare Workers edge processing.

[0030] <Configuration of Internal State Management Unit 30> The Internal State Management Unit 30 manages a group of physical condition parameters 31 (health level H, fatigue level F) and a group of environmental parameters 32 (external stress index S, external event flag E). Each parameter receives update information from the story engine 101 at 500-millisecond intervals and writes it to Redis.

[0031] <Operation of the Composite Evaluation Value Calculation Unit 40> As shown in Figure 2, the Composite Evaluation Value Calculation Unit 40 starts up every 500 milliseconds and calculates V = wH × (1-H) + wF × F + wS × S + wE × E (default: wH=0.3, wF=0.3, wS=0.3, wE=0.1). If V > θ1 (0.65), it sends an activation signal to the autonomous movement control unit 50. Even if V ≤ θ1, it also sends an activation signal if any of the composite trigger conditions (V ≥ 0.455 and TS > 120 seconds, TG > 10 seconds, or speech detection) are met.

[0032] <Operation of Autonomous Movement Control Unit 50> As shown in Figure 3, after receiving the activation signal, the user position Pu is acquired, the distance D = |Pc - Pu| is calculated, and a path calculation using the A* algorithm is performed. The character's movement speed is 1.2 m / sec by default. The path is updated with feedback every 33 milliseconds. When D ≤ θ2 (0.5 m) is reached, an activation signal is sent to the embrace sequence control unit 70. During movement, the rejection detection unit 80 and the story deviation detection unit 95 operate in parallel.

[0033] <Operation of AI Dialogue Control Unit 60> As shown in Fig. 4, the AI dialogue control unit 60 starts up simultaneously with the start of autonomous movement, acquires the distance D at a cycle of 33 milliseconds, and executes zone determination. Zone 1 (D > 3.0 m): Volume 100% · Polite form · Interval of 30 seconds or more. Zone 2 (1.5 - 3.0 m): Volume 80% · Plain form · Interval of 15 seconds · Conversation history 8T added. Zone 3 (0.5 - 1.5 m): Volume 60% · Intimate form · Interval of 8 seconds · Emotional state details added. Zone 4 (D ≤ 0.5 m): Volume 40% (whisper) · Highest intimacy style · Interval of 3 seconds · Hugging phase information added.

[0034] <Operation of Hugging Sequence Control Unit 70> It is composed of four phases. Phase 1 (Approaching phase, 0 - 0.5 seconds): Walk the final 0.5 m while spreading the arms. Phase 2 (Hugging start phase, 0.5 - 0.8 seconds): The arms wrap around the user. Phase 3 (Hugging maintenance phase, 0.8 - 5.8 seconds): Continue the AI voice dialogue while maintaining the hugging posture. Phase 4 (Detachment phase, 5.8 - 6.3 seconds): A natural way to separate. The phase information field of the LLM prompt is updated at each phase transition.

[0035] <Operation of Story Deviation Detection Unit 95> The story deviation detection unit 95 operates constantly throughout the VR experience. It executes the following three types of deviation detections in parallel. (1) Gaze deviation detection: Acquire the user's gaze vector from the gaze tracking sensor (90 Hz), and set the gaze deviation flag when the state where the angle difference from any of the "expected attention areas" (a list of areas in the VR space that the user should pay attention to based on the current story progress) acquired from the story engine 101 exceeds 15° and continues for 10 seconds or more. (2) Utterance deviation detection: Calculate the cosine similarity between the vector representation of the recognized text received from the ASR engine and the vector representation of the current story context, and set the utterance deviation flag when an utterance with a similarity less than 0.30 (statistically deviated from the story context) continues. (3) Non - response detection: Detect a state where the amount of position change from the IMU data is less than 3 cm per unit time and the ASR does not detect an utterance for 30 seconds or more as non - response.

[0036] If one or more of the three systems set a flag, a deviation detection flag is set and a notification is sent to the natural recovery guidance unit 96. If multiple systems set flags simultaneously, the deviation score (weighted sum of the number of flags × the duration coefficient of each system) increases, and the escalation rate increases.

[0037] <Details of the operation of the natural recovery guidance unit 96> As shown in Figure 5, the natural recovery guidance unit 96 measures the deviation duration TD and performs a three-stage escalation type guidance. The reference time for stage determination can be changed in the setting file of the story engine 101 (default: first reference time T1 = 30 seconds, second reference time T2 = 60 seconds).

[0038] First stage (TD: 10~T1 seconds): The system instruction is added to the prompt to the LLM: "The user is slightly distracted. Gently attract their attention without being forceful, using natural language that fits the context of the current story scene (<scene description>). Generate a short, character-like utterance." Then, the utterance is generated. Example: "Um... are you listening?" "I'd like you to look over here" (example). At this stage, the character's position does not change, and only the utterance is performed.

[0039] Second stage (TD: T1-T2 seconds): Add the following instruction to the prompt for LLM: "Generate utterances that rekindle the user's interest by creating unresolved mysteries in the story, anticipation for future developments, and empathy for the characters' current emotional states. Leverage curiosity and empathy to guide the user to spontaneously become interested." Examples: "Actually, I've been wondering about that from earlier... Shall I tell you?" "I'm feeling really anxious right now. There's something I want to talk to you about." (Examples). The LLM automatically extracts and utilizes story-specific mysteries and emotional hooks from the context.

[0040] Third stage (TD: T2 seconds or more): An approach activation signal is sent to the autonomous movement control unit 50 (a trigger path independent of the threshold determination of the composite evaluation value V), and the character generates a gentle re-induction utterance while physically approaching the user. The prompt to the LLM includes the instruction, "The user has been away from the story for a long time. The character should approach and speak to the user gently and emotionally sincerely to re-engage them. This should be expressed as the character's earnest desire, not as coercion." Example: "Hey, I want you to stay by my side. Please." (Example).

[0041] <Determination of Deviance Recovery End> The natural recovery guidance unit 96 monitors the user's behavior after each stage of guidance is performed. If any of the following occurs within 30 seconds, such as the user's gaze returning to the expected attention area, the speech context similarity exceeding 0.50, or the user's action being detected, the deviation detection flag is reset and the guidance ends. After completion, the deviation duration and guidance effect data are recorded in the story engine 101 and used to optimize the timing of future guidance.

[0042] <Priority control with rejection detection> The rejection detection unit 80 operates continuously even during the execution of the third stage of story deviation (approach-type re-guidance). If physical rejection (retreat speed exceeding 15 cm / sec, head rotation exceeding 50° / sec) or auditory rejection is detected, approach is immediately stopped, and the emotion reflection unit 90 executes the corresponding emotion change and animation. If rejection occurs during deviation return guidance, guidance is completely interrupted, a cooldown period (default 180 seconds) is established, and then re-evaluation is performed.

[0043] <Changing the settings of each parameter> The first threshold θ1, the second threshold θ2, each weight coefficient, and each threshold for deviation detection (reference time T1·T2, gaze angle threshold, utterance OOD threshold, no-response judgment time) can be changed on the configuration screen of Story Engine 101. Customization is possible according to the character's personality (set T1 shorter for an assertive personality, set T1 longer for a cautious personality) and the nature of the content (set deviation detection loosely for free exploration type content, and strictly for story type content that seeks strong immersion). [Examples]

[0044] Examples of the present invention are described below. [Examples]

[0045] A completely spontaneous experience of approaching and embracing due to a deterioration in physical condition.

[0046] With the character "Aoi" (fictional name) at H=0.4, F=0.92, S=0.3, and E=1, V=0.18+0.276+0.09+0.10=0.646≈0.65, reaching θ1 and activating autonomous movement. Aoi approaches at a speed of 1.2m / s while AI voice dialogue is activated in parallel. Upon reaching D=0.5m, the embrace sequence is activated, and during phase 3 (maintenance), whisper-volume speech is generated under zone 4 control. [Examples]

[0047] First stage: deviance recovery experience (gentle verbal encouragement)

[0048] This scenario describes a situation where, during a normal scene, the user continues to look near the ceiling in the VR space for 15 seconds (gaze deviation detection). The first stage is activated at TD=15 seconds. The first stage instruction is added to the prompt to the LLM, generating a natural-sounding phrase such as "Um...is there anything that's bothering you?" (example). The character does not move and only delivers the line. If the user responds to this line and looks at the character (gaze within 15° of the character's direction for 5 seconds), the deviation flag is reset and the guidance ends. [Examples]

[0049] Second stage: Deviance recovery experience (stimulation of curiosity and empathy)

[0050] If there is no response to the first stage prompt and the deviation duration TD exceeds 30 seconds (T1), the second stage is activated. The LLM extracts the "unsolved mystery" element of the current scene from the story context and generates an utterance such as, "Actually... I've been wondering about that door from earlier. Shall we check it out together?" (example). By automatically utilizing story-specific hooks with the LLM, a highly versatile guidance system is achieved that works without content creators having to design deviation scenarios in advance. [Examples]

[0051] Third stage deviation recovery experience (approach-type re-induction)

[0052] If TD exceeds 60 seconds (T2), the third stage is activated. An activation signal is sent to the autonomous movement control unit 50, and the character begins to approach the user at a gentle speed of 1.0 m / s. The LLM generates an emotionally charged utterance such as, "Hey... I want you to stay here. I'll be so lonely if you leave" (example). If rejection detection is performed during the approach and the user backs away, the process immediately stops and transitions to emotion reflection processing. If there is no rejection and D=1.5m (zone 2) is reached, the LLM switches to an intimate style of utterance, creating an experience where the user regains interest in the story.

[0053] In the first embodiment described above, the state acquisition step (step (a)) acquires a set of physical condition parameters, including at least health level H and fatigue level F, as the emotional state of the autonomous agent, and a set of environmental parameters, including an external stress index S and an external event flag E. Health level H is a continuous value from 0 (extremely unhealthy) to 1 (perfectly healthy), and fatigue level F is a continuous value from 0 (no fatigue) to 1 (extreme fatigue). The external stress index S represents the mental burden caused by events in the story as a scalar value from 0 to 1, and the external event flag E is a binary flag that becomes 1 when a specific event occurs in the progression of the story. These parameters are acquired periodically based on update information from a scenario database managed by the story engine.

[0054] The weighted combination operation in the control value calculation step (step (b)) in the first embodiment described above is configured to allow individual adjustment of the contribution of each state parameter, thereby realizing diverse behavioral patterns according to the personality setting of the autonomous agent. For example, in an agent with a highly empathetic personality, setting a large amount of wH (weight for health) can impart behavioral characteristics that react sensitively to changes in the user's physical condition. Furthermore, in the separation-inducing intervention embodiment among the intervention embodiments, the autonomous agent moves to a position a certain distance (e.g., 2.0m or more) away from the user, and then changes its gaze direction or speaks to itself to arouse the user's curiosity and induce spontaneous approach behavior.

[0055] The relationship adjustment determination step (step (c)) in the first embodiment described above includes, as adjustment of spatial relationships, a change in physical distance due to a change in the virtual space coordinates of the autonomous agent; as adjustment of perceptual relationships, a change in the focal length of the virtual camera, a change in depth of field, or the application of a background blur effect; and as adjustment of the theatrical relationships, a change in the color temperature of ambient lighting, a change in the tempo of ambient music, or the generation of particle effects. These relationship adjustments are performed individually or in combination, and the intensity of the adjustments changes in steps according to the magnitude of the intervention control value.

[0056] In the output expression execution step (step (d)) of the first embodiment described above, the dynamic changes based on distance information include the following attributes: For voice attributes, volume level (decreases inversely proportional to distance), speech rate (changes more gradually at closer distances), voice quality parameters (fine-tuning of pitch and formant frequency), and speech interval (shortened at closer distances) change in conjunction with distance. For motion attributes, walking speed (decelerates when transitioning to a close distance), trunk tilt angle (increases forward tilt at close distances), and hand gesture amplitude (decreases at close distances) change. For environmental effect attributes, background music volume (attenuates at close distances), ambient lighting warming (increases at close distances), and particle density (increases at close distances) change.

[0057] In the first embodiment described above, the response adjustment step (step (e)) classifies the user's response into three categories: positive responses (approaching behavior, fixed gaze, positive utterances), neutral responses (no response, minor movements), and negative responses (retreating behavior, avoiding gaze, rejecting utterances), and applies adjustment rules corresponding to each category. If a positive response is detected, relation adjustment control is maintained or strengthened; if a neutral response continues for a predetermined time (e.g., 15 seconds) or longer, the output expression is changed to draw attention; and if a negative response is detected, relation adjustment control is immediately suppressed or stopped.

[0058] <Details of the detection algorithm for voluntary approach behavior> In this embodiment, the algorithm for detecting a user's voluntary approach behavior is as follows: The user's head position coordinates are obtained from the IMU (Inertial Measurement Unit, sampling rate 90Hz) built into the HMD20, and the movement direction vector and movement speed are estimated by applying linear regression to the position data of the past 1 second (90 samples). If the angle difference between the estimated movement direction vector and the direction vector from the user to the character is within 30°, and the movement speed is 5cm / second or more, this condition is detected as a voluntary approach behavior by the user for 2 seconds or more. Furthermore, if the controller stick input direction matches the character direction within 45°, it is also detected as a voluntary approach. These two detection systems are integrated by an OR condition, and if either one is true, the approach behavior flag is set.

[0059] <Specific Examples of Distance and Velocity Sensing Methods> Real-time distance information between the autonomous agent and the user is acquired using the following methods. As the first method, the Euclidean distance between the user's head position coordinates, acquired from the HMD20's inside-out tracking camera (stereo camera method, resolution 640 x 480, frame rate 30 fps or 60 fps), and the autonomous agent's coordinates in virtual space is calculated at a 33-millisecond interval. As the second method, when using an external base station tracking system (laser method or infrared method, tracking accuracy of 1 mm or less), the distance is calculated from the 6 degrees of freedom (6DoF) position and orientation information. In both methods, noise reduction processing using a Kalman filter is applied to exclude rapid fluctuations in distance values ​​(inter-frame changes of 0.5 m or more) as outliers. Velocity information is calculated by the time derivative of the distance and smoothed by applying a moving average filter (window width 0.5 seconds).

[0060] <Specific Processing Flow in the Health Deterioration Scenario> The processing flow for the health deterioration scenario in this embodiment is described in detail below. An example is given where the story engine 101 gradually decreases the character's health level H from 1.0 based on the scenario script. First stage (H=0.8~0.6): The internal state management unit 30 detects the decrease in H and reflects a mild fatigue expression (reduced eye opening to 95%, slight drooping of the corners of the mouth) in the character's facial expression parameters. At this stage, the composite evaluation value V does not reach the threshold θ1, so autonomous movement is not activated, but the context of "not feeling well" is added to the prompt of the AI ​​dialogue control unit 60, and signs of mild poor health are reflected in the utterances. Second stage (H=0.6~0.4): The facial expression parameters reflect a moderate fatigue expression (reduced eye opening to 85%, slight wrinkles between the eyebrows, decrease in cheek redness). The composite evaluation value V approaches the lower limit of the composite trigger condition, and autonomous movement can be activated depending on the composite conditions with other parameters. Third stage (H≦0.4): The composite evaluation value V reaches the threshold θ1, and an activation signal is sent to the autonomous movement control unit 50. The character begins approaching at a speed slightly slower than the normal movement speed (1.2 m / sec) (0.8 m / sec), and the prompt to the LLM is updated with the context that "the character is feeling unwell and is seeking comfort." The animation during movement also reflects the character's illness, and the amplitude of the gait fluctuations is set to 1.5 times the normal value.

[0061] <Detailed Criteria for Multi-Stage Control of Deviation Recovery> The details of the criteria for each stage of the three-stage escalation type guidance performed by the natural recovery guidance unit 96 are shown below. The activation conditions for the first stage (gentle verbal encouragement) are that the deviation duration TD is 10 seconds or more and less than T1 (default 30 seconds), and the deviation score DS (calculated as a weighted sum of the presence or absence of flags and duration for each detection system, DS = α1 × f1 × t1 + α2 × f2 × t2 + α3 × f3 × t3, where α1 = 0.4, α2 = 0.35, α3 = 0.25, fi is the flag value, and ti is the normalized duration) is 0.2 or more. In the speech generation of the first stage, a simple verbal encouragement is generated by giving an instruction to the LLM to limit the upper limit of the number of tokens to 50 tokens. The activation conditions for the second stage (curiosity stimulation) are that TD is T1 or more and less than T2 (default 60 seconds), and no user recovery behavior is detected after the guidance of the first stage. In the second stage, the "list of unresolved elements in the current scene" and the "description of the character's current emotional state," obtained from story engine 101, are injected into the prompts for the LLM, and the token limit is expanded to 150 tokens to generate detailed guiding utterances. The activation conditions for the third stage (approach-type re-guidance) are that TD is T2 or higher and no rejection detection has occurred within the last 180 seconds. In the third stage, the autonomous movement speed is set to a slower 1.0 m / s than the normal speed (1.2 m / s) to ensure that the approach action itself does not intimidate the user.

[0062] <Specific Implementation of Haptic Feedback> In this embodiment, the hug sequence control unit 70 generates additional haptic output when a haptic feedback compatible device is connected. As a first implementation example, vibration pattern control is performed using a vibration motor (LRA method: linear resonant actuator method, frequency band 150~250Hz) built into a standard VR controller. In hug phase 2 (hugging start), three short pulse vibrations (pulse width 100 milliseconds, pulse interval 200 milliseconds) with a frequency of 200Hz and an amplitude of 0.3G are generated to create the sensation of arms touching. In hug phase 3 (hugging maintenance), continuous low vibrations with a frequency of 150Hz and an amplitude of 0.1G are output in sync with the character's breathing cycle (sine wave amplitude modulation with inhalation 2.0 seconds, exhalation 3.0 seconds) to create the effect of feeling breathing. In hug phase 4 (releasing), the amplitude is faded out from 0.1G to zero over 2.0 seconds. As a second implementation example, if a multi-point force feedback device such as a tactile glove or tactile vest is connected, the output value of each actuator is mapped to the bone coordinates of the character's arm, and the contact area is controlled to increase in stages as the embrace phase progresses (Phase 2: 2 points on the front of the chest, Phase 3: 2 points on the front of the chest + 2 points on the back + 2 points on the upper arm, for a total of 6 points). The pressure value at each point is controlled within the range of 0.5N to 2.0N.

[0063] <Specific Method for User State Estimation> User state estimation in the response adjustment step is performed by integrating the following multiple sensor inputs. Firstly, based on gaze data acquired from the eye-tracking sensor (infrared camera type, sampling rate 90Hz, accuracy within 1°) built into the HMD20, changes in pupil diameter (dilation is used as an indicator of interest / surprise, constriction as an indicator of discomfort), fixation stability (if the amplitude of fixation micro-movement is within 0.5°, it is determined to be a state of concentration), and blink frequency (more than 25 times / minute is detected as a sign of fatigue / boredom, compared to the normal value of 15-20 times / minute). Secondly, the trunk posture tilt angle is estimated from the IMU data, and a forward-leaning posture (forward tilt of 5° or more) is used as an indicator of active engagement, and a backward-leaning posture (backward tilt of 10° or more) is used as an indicator of a passive or rejecting attitude. Thirdly, the system obtains utterance text and prosodic features (average pitch, pitch variation, speech rate, and pause length) from the ASR (Automatic Speech Recognition) engine, and performs sentiment classification using LLM (three-level classification of positive, neutral, and negative, and confidence scores for each class). These estimation results are integrated at 0.5-second intervals and output as a user state vector (three-dimensional representation of engagement, comfort, and focus). This user state vector is used as an input parameter for relational adjustment control or dynamic adjustment of output representation in the response adjustment step.

[0064] <List of Numerical Examples> The following is a list of the main numerical parameters in this embodiment. Note that the following values ​​are all preferred examples, and the present invention is not limited to these values. Distance threshold: θ2 = 0.5m (hugging activation distance), zone boundary = 3.0m, 1.5m, 0.5m. Speed ​​threshold: Normal movement speed = 1.2m / sec, movement speed when physical condition deteriorates = 0.8m / sec, approach-type re-guidance speed = 1.0m / sec, user spontaneous approach detection speed threshold = 5cm / sec, user reverse rejection detection speed threshold = 15cm / sec. Reaction time: State parameter update cycle = 500 milliseconds, distance update cycle = 33 milliseconds (30fps), rejection detection response time = within 100 milliseconds (immediate stop). Deviation detection time: Gaze deviation flag setting time = 10 seconds, no response detection time = 30 seconds, first stage activation time = 10 seconds, second stage activation time = 30 seconds (T1), third stage activation time = 60 seconds (T2), recovery confirmation waiting time = 30 seconds. Angle threshold: Gaze deviation angle threshold = 15°, spontaneous approach direction angle threshold = 30°, head rotation rejection threshold = 50° / second.

[0065] <Modification 1: Alternative Sensor Type> In the above embodiment, the HMD's built-in inside-out tracking and IMU were described as the main sensing methods, but the present invention is not limited thereto. As a first modification, an externally installed depth camera (ToF sensor method or structured light method, distance measurement range 0.5~10m, distance measurement accuracy ±10mm) may be used to acquire the user's whole-body skeletal information in real time, and the distance and posture may be estimated from the three-dimensional coordinates of the head, chest, and hands. As a second modification, in a configuration using AR glasses (transmissive head-mounted display), the distance may be calculated from the user's position on a spatial map acquired by the SLAM (Simultaneous Localization and Mapping) function mounted on the AR glasses and the autonomous agent's display position. As a third modification, a biosignal sensor (photoplethysmography sensor for acquiring heart rate variability HRV or skin electrical response GSR sensor) may be added, and the user's autonomic nervous system activity index may be used as an auxiliary input for state parameters. In this case, if the ratio of low-frequency components to high-frequency components (LF / HF ratio) of heart rate variability exceeds a predetermined value (e.g., 3.0), it may be determined that the user is in a stressed state, and the sensitivity of deviation detection may be temporarily reduced (the frequency of prompts is suppressed for users in a stressed state).

[0066] <Modification 2: Alternatives to Haptic Devices> In the above embodiment, an LRA vibration motor and a multi-point force feedback device built into the VR controller were exemplified, but the method for generating haptic feedback is not limited to these. As a first alternative, an ultrasonic haptic feedback device (airborne ultrasonic phased array method, output frequency 40kHz, modulation frequency 200Hz, focal diameter approximately 10mm) may be used to present haptic stimuli to the user's hand or forearm without contact. This method has the advantage of providing haptic feedback without compromising immersion, as it does not require the attachment of a device. As a second alternative, an electrotactile stimulator (transcutaneous electrical stimulation method, stimulation current 0.5~5mA, pulse width 0.1~1millisecond) may be attached to the user's arm as a thin patch electrode to electrically reproduce pressure and temperature sensations. As a third alternative, haptic feedback may be replaced by visual effects (superimposing a ripple effect that suggests touch on the surface of the user's avatar) and auditory effects (spatial acoustic rendering of environmental sounds such as rustling clothes and breathing sounds) without relying on physical devices. In this case, a simulated haptic impression can be provided without requiring any additional hardware. [Explanation of Symbols]

[0067] 10 VR Interaction Systems 20 HMD (Head-Mounted Display Device) 30 Internal Condition Management Department 40. Composite Evaluation Value Calculation Unit 50 Autonomous Mobile Control Unit 60 AI Dialogue Control Unit 70 Embrace Sequence Control Unit 80 Rejection detection unit 81. Physical Rejection Detection Module 82. Voice rejection detection module 90 Emotion reflection part 95 Story Deviation Detection Unit (New) 96 Natural Recovery Induction Section (New) 100 Cloud Servers 101 Story Engine V Composite evaluation value θ1 First threshold (0.65) θ2 Second threshold (0.5m) Sd (Sadness Parameter) TD deviation duration T1 First reference time (30 seconds) T2 Second reference time (60 seconds)

Claims

1. A method for controlling the interaction of an autonomous agent in a computer-generated virtual environment, wherein the virtual environment is configured as a space where a user's avatar exists via a head-mounted display, and the computer performs the following steps: (a) a state acquisition step in which the computer acquires a group of physical condition parameters of the autonomous agent itself, which includes at least health level and fatigue level, and a group of environmental parameters including an external stress index and an external event flag, at predetermined intervals; (b) a control value calculation step in which the computer calculates a composite evaluation value by weighted sum by applying individual weight coefficients to each parameter of the physical condition parameter group and the environmental parameter group, and when the composite evaluation value exceeds a first threshold, it determines to initiate autonomous movement of the autonomous agent to the user's avatar; and (c) based on the decision to initiate autonomous movement, the autonomous agent An interaction control method comprising: (d) a relationship adjustment determination step that transitions the spatial relationship between the agent and the user's avatar, based on coordinates in virtual space, to a target state by relationship adjustment control that includes a change in physical distance due to a change in the coordinates of the autonomous agent; (d) an output expression execution step that, simultaneously with the start of the autonomous movement, acquires real-time distance information between the autonomous agent and the user's avatar, and generates an output expression including the motion and voice of the autonomous agent while performing AI voice dialogue control that automatically changes the volume, speech frequency and language style for each of the multiple distance zones, and dynamically modifies it in accordance with the progress of the relationship adjustment control; and (e) a response adjustment step that dynamically adjusts the manner of the relationship adjustment control or the output expression based on data regarding the user's response during the execution of the relationship adjustment control or the control of the output expression.

2. An interaction control method according to claim 1, wherein the relationship adjustment control further includes adjusting a perceptual relationship, including changing the focal length of a virtual camera, changing the depth of field, or applying a background blur effect.

3. An interaction control method according to claim 1, wherein the relationship adjustment control further includes adjusting the theatrical relationship, including changing the color temperature of ambient lighting in the virtual environment, changing the tempo of ambient music, or generating particle effects.

4. An interaction control method according to claim 1, wherein the relationship adjustment control includes at least one of an affiliative behavior, an empathetic behavior, or a relationship-reinforcing behavior by the autonomous agent toward the user.

5. An interaction control method according to claim 1, wherein the change in the relationship in the relationship adjustment control is realized by at least one of the following: changing the position of the autonomous agent, moving the viewpoint of the user's avatar, changing the spatial scale of the virtual environment, changing the display size of the autonomous agent, or changing the camera work.

6. An interaction control method according to claim 1, wherein the relationship adjustment control includes changing a perceptual relationship or a theatrical relationship by a method that does not involve changing the position on the coordinate system in the virtual environment.

7. An interaction control method according to claim 1, wherein the control value calculation step selects at least one of the following based on the composite evaluation value: an intervention that strengthens the relationship between the autonomous agent and the user's avatar; an intervention that arouses the user's attention or interest by separating the autonomous agent, changing the display mode or changing environmental elements; and an intervention that suppresses or stops the current action of the autonomous agent.

8. An interaction control method according to claim 1, wherein the control value calculation step dynamically selects or combines, based on the composite evaluation value, which of a plurality of intervention modes having different directions or properties from each other, the plurality of intervention modes includes at least an intervention mode that brings the autonomous agent closer to the user's avatar and an intervention mode that changes the behavior pattern of the autonomous agent in order to draw the user's attention to the autonomous agent.

9. An interaction control method according to claim 1, wherein the group of physical condition parameters and the group of environmental parameters in the state acquisition step are periodically acquired based on update information from a scenario database managed by the story engine.

10. An interaction control method according to claim 1, wherein the response adjustment step processes at least one of the user's physical movements and the voice.

11. An interaction control method according to claim 1, wherein the data relating to the response includes at least one of the following: changes in body posture, head movement, gaze direction, hand movements, prosodic features of speech, speech content, biosignals, and controller operation patterns.

12. An interaction control method according to claim 1, wherein the dynamic adjustment in the response adjustment step includes at least one of the relation adjustment control or the output expression suppression, stop, mitigate, change, strengthen, or step transition.

13. An interaction control method according to claim 1, wherein the response adjustment step includes stopping, delaying, or canceling the current action of the autonomous agent when it is determined that the user's response is rejecting.

14. An interaction control method according to claim 1, further comprising a re-intervention decision step of recalculating the composite evaluation value and determining whether or not to perform the relationship adjustment control again when the effect of the relationship adjustment control has diminished over time.

15. An interaction control method according to claim 1, wherein, when a plurality of autonomous agents exist in the virtual environment, at least two of the plurality of autonomous agents cooperate to perform the relationship adjustment control or the output expression control.

16. An interaction control method according to claim 1, wherein the plurality of distance zones in the output expression execution step include four distance zones: a far distance zone, a medium distance zone, a near distance zone, and a very near distance zone, the volume in the very near distance zone is set to be lower than the volume in the far distance zone, and the language style in the very near distance zone is set to the highest intimacy style.

17. An interaction control method according to claim 1, further comprising a hug sequence control step of initiating a hug sequence by the autonomous agent when the distance between the autonomous agent and the user's avatar reaches a predetermined second threshold or less, wherein the hug sequence consists of a plurality of phases including an approach phase, a hug start phase, a hug maintenance phase and a release phase, and the context information of the AI ​​voice dialogue control is updated according to the transition of each phase.

18. In the interaction control method according to claim 1, when interactive content is operating in the virtual environment, the computer further performs a state deviation detection step in which (a) parallel execution of three systems of deviation detection: gaze deviation detection, speech deviation detection, and no response detection, and sets a deviation detection flag when one or more of the three systems set a flag, and calculates a deviation score as a weighted sum of the number of flags in the multiple systems and the duration coefficient of each system, and (b) a first step in which, if the deviation duration is less than the first reference time, a voice utterance is generated without changing the position of the autonomous agent, and the deviation duration is less than the first reference time An interaction control method comprising: (c) an adaptive induction step of three stages, including a second stage of generating a curiosity-stimulating utterance using unresolved elements in the story and the emotional state of the character when the deviation duration is greater than or equal to the second reference time and less than the second reference time; and a third stage of generating a re-induction utterance involving the autonomous agent's physical approach to the user's avatar when the deviation duration is greater than or equal to the second reference time; and an adaptive control step of maintaining, stopping, changing, adding, switching or adjusting the intensity of the return induction means in operation, based on the user's response during or after the execution of the adaptive induction step.

19. An interaction control method according to claim 18, wherein if a physical or audible rejection by the user is detected during the execution of the third stage of the adaptive guidance step, the approach is immediately stopped, and a re-evaluation is performed after a predetermined cool-down period has elapsed.

20. An interaction control method according to claim 1, wherein the response adjustment step classifies the user's response into three categories: positive response, neutral response, and negative response; maintains or strengthens the relationship adjustment control when a positive response is detected; changes the form of the output expression to draw attention when a neutral response continues for a predetermined time or longer; and suppresses or stops the relationship adjustment control when a negative response is detected.

21. An information processing apparatus comprising a processor and a memory, wherein the processor executes a program stored in the memory to perform each step of the interaction control method described in claim 1.

22. An information processing apparatus comprising a processor and a memory, wherein the processor executes a program stored in the memory to perform each step of the interaction control method described in claim 18.

23. A program for causing a computer to function as a means for performing each step of the method according to claim 1.

24. A program for causing a computer to function as a means for performing each step of the method according to claim 18.