Immersive holographic theater-based virtual-real character dynamic interaction control system and method

By working together with spatial positioning, motion capture, emotion capture, and dynamic rendering modules, the problem of unnatural interaction between the audience and virtual characters in immersive holographic theaters has been solved, achieving high-precision, real-time interaction between virtual and real characters and dynamic rendering, thus enhancing the immersive experience.

CN121937682BActive Publication Date: 2026-06-12HHTC (XIAMEN) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HHTC (XIAMEN) TECH CO LTD
Filing Date
2026-03-30
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

In traditional immersive holographic theaters, the interaction between the audience and virtual characters lacks natural, real-time, and precise interaction. The spatial positioning and emotion capture technologies are not precise enough, resulting in poor virtual-real fusion. Dynamic rendering cannot be adjusted according to the audience's state, and there is a lack of immersive experience.

Method used

The system employs a spatial positioning module that works in collaboration with UWB beacons and tags to accurately acquire the audience's location; a motion capture module and an emotion capture module that collect the audience's body movements and emotional characteristics in real time; a virtual-real interaction module that controls the virtual character's feedback based on this data; and a dynamic rendering module that adjusts rendering parameters according to the audience's location and emotions, enabling multiple modules to work together.

Benefits of technology

It achieves a highly realistic and interactive immersive environment for the audience and virtual characters. The virtual characters can react naturally according to the audience's real-time location and emotions, enhancing the sense of immersion and interactivity, and improving the quality of visual presentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937682B_ABST
    Figure CN121937682B_ABST
Patent Text Reader

Abstract

This invention relates to a dynamic interactive control system and method for virtual and real characters based on immersive holographic theaters, and belongs to the field of electronic digital data processing technology. The system includes: a spatial positioning module for collecting the position information of the audience within a virtual-real fusion space constructed using light field reconstruction technology; a virtual-real interaction module for adjusting the state of virtual characters within the virtual-real fusion space based on the audience's position information; a motion capture module for collecting the audience's body movement data; an emotion capture module for extracting the audience's emotional features from facial images and voice data; the virtual-real interaction module is also used to control the virtual characters to provide interactive feedback based on the audience's body movement data and emotional features; and a dynamic rendering module for adjusting rendering parameters of the virtual-real fusion space in real time based on the audience's position information and emotional features, which has the advantages of realizing natural, real-time, and accurate interaction between the audience and the virtual characters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, and in particular to a dynamic interactive control system and method for virtual and real characters based on immersive holographic theater. Background Technology

[0002] In the current booming development of the cultural and entertainment industry, immersive experiences have gradually become a key element in attracting audiences and enhancing viewing value. As an emerging product under this trend, immersive holographic theaters integrate advanced technologies such as light field reconstruction, virtual reality, and motion capture, bringing audiences an unprecedented audio-visual feast.

[0003] In traditional theatrical performances, audience interaction with actors is often limited to simple responses or limited on-site participation. The interaction between characters is relatively simple and fails to meet the audience's growing demand for personalized and in-depth experiences. With the continuous advancement of technology, although some theaters have begun to introduce virtual characters to enrich the performance content, most of these virtual characters are controlled by preset programs and lack the ability to interact with the audience in real time and naturally. They cannot respond flexibly according to the audience's position, movements, and emotional state, which greatly diminishes the effect of blending the virtual and real in the performance, making it difficult for the audience to truly immerse themselves in the experience.

[0004] In terms of spatial positioning, previous technologies struggled to accurately capture audience location information in complex virtual-real hybrid spaces, preventing virtual characters from accurately sensing the audience's location and thus affecting the timeliness and accuracy of interaction. While motion capture technology can capture audience body movements to some extent, it suffers from insufficient precision and high latency, failing to capture the details of audience movements in real time and with fine detail, thus limiting the accurate imitation and feedback of audience movements by virtual characters. In the field of emotion capture, existing technologies mostly focus on the analysis of single-modal data, such as inferring audience emotions solely from facial images or voice data. This makes it difficult to comprehensively and accurately extract the complex emotional characteristics of the audience, resulting in less realistic and vivid emotional responses from virtual characters.

[0005] Furthermore, as a crucial element in creating an immersive atmosphere, traditional methods of dynamic rendering fail to adequately consider the audience's position and emotional factors. With fixed rendering parameters, the visual effects of the virtual-real fusion space cannot be adjusted according to the audience's real-time state, resulting in a monotonous and layered scene presentation that cannot create an immersive environment that closely matches the performance content and makes the audience feel as if they are there.

[0006] Therefore, there is a need to provide a dynamic interactive control system and method for virtual and real characters based on immersive holographic theaters, so as to realize natural, real-time and accurate interaction between the audience and virtual characters. Summary of the Invention

[0007] This invention provides a dynamic interactive control system for virtual and real characters based on an immersive holographic theater, comprising: a spatial positioning module for collecting audience position information within a virtual-real fusion space constructed using light field reconstruction technology; a virtual-real interaction module for adjusting the state of virtual characters within the virtual-real fusion space based on the audience's position information; a motion capture module for collecting audience body movement data with user authorization; an emotion capture module for collecting audience facial images and voice data with user authorization, and extracting audience emotional features from the facial images and voice data; the virtual-real interaction module is also used to control virtual characters to provide interactive feedback based on the audience's body movement data and emotional features; and a dynamic rendering module for adjusting rendering parameters of the virtual-real fusion space in real time based on the audience's position information and emotional features.

[0008] Furthermore, the spatial positioning module is further configured to: divide the virtual-real fusion space into multiple regions; for each region, calculate the comprehensive weight of the region based on the historical movement trajectories of multiple historical viewers in the virtual-real fusion space obtained with user authorization, determine the density of UWB beacons set in the region based on the comprehensive weight of the region, and set multiple UWB beacons in the virtual-real fusion space based on the density of UWB beacons set in each region.

[0009] Furthermore, the spatial positioning module is further configured to: for each region, determine the percentage of time spent in the region based on the historical movement trajectories of multiple historical viewers in the virtual-real fusion space, and calculate the first weight of the region based on the percentage of time spent in the region; for any two regions, determine the synchronization ratio of time spent in the two regions based on the historical movement trajectories of multiple historical viewers in the virtual-real fusion space; for each region, calculate the second weight of the region based on the synchronization ratio of time spent in the region with any other region; and calculate the comprehensive weight of the region based on the first weight and the second weight of the region.

[0010] Furthermore, the spatial positioning module is further configured to: determine multiple regional positioning UWB beacons from multiple UWB beacons; determine the area where the audience is located using UWB tags and multiple regional positioning UWB beacons; filter multiple target UWB beacons from multiple UWB beacons based on the comprehensive weight of the audience's area, the performance of multiple UWB beacons, and the signal transmission data of multiple UWB beacons and UWB tags; and generate the audience's location information based on the signal transmission data of multiple target UWB beacons and UWB tags.

[0011] Furthermore, the virtual-real interaction module is further used to: determine the adjustment frequency based on the audience's trajectory information; calculate the distance between the audience and the virtual character based on the audience's position information and the current position of the virtual character according to the adjustment frequency; calculate the angular deviation between the audience and the virtual character based on the audience's position information, the current position and orientation of the virtual character; and adjust the state of the virtual character in the virtual-real fusion space based on the distance and angular deviation between the audience and the virtual character.

[0012] Furthermore, the virtual-real interaction module is further used to: determine similar historical movement trajectories based on the audience's trajectory information; and determine the adjustment frequency based on the audience's trajectory information and similar historical movement trajectories.

[0013] Furthermore, the virtual-real interaction module is further used to: match response actions based on the audience's body movement data and emotional characteristics; match response emotions based on the audience's body movement data and emotional characteristics; and control the virtual character to provide interactive feedback based on the response actions and response emotions.

[0014] Furthermore, the dynamic rendering module is further used to adjust the texture resolution of the virtual character based on the distance between the viewer and the virtual character.

[0015] Furthermore, the dynamic rendering module is further used to adjust lighting rendering parameters based on the audience's emotional characteristics.

[0016] This invention provides a method for dynamic interactive control of virtual and real characters based on an immersive holographic theater, comprising: collecting the position information of the audience within a virtual-real fusion space constructed using light field reconstruction technology; adjusting the state of the virtual characters within the virtual-real fusion space based on the audience's position information; collecting the audience's body movement data; collecting the audience's facial images and voice data, and extracting the audience's emotional features from the audience's facial images and voice data; controlling the virtual characters to provide interactive feedback based on the audience's body movement data and emotional features; and adjusting the rendering parameters of the virtual-real fusion space in real time based on the audience's position information and emotional features.

[0017] Compared to existing technologies, the virtual-real character dynamic interaction control system and method based on immersive holographic theater provided in this specification has at least the following beneficial effects:

[0018] 1. By analyzing historical audience movement trajectories to divide areas and determine UWB beacon density, beacons can be rationally deployed according to the characteristics of different areas, achieving high-precision acquisition of audience positions. The virtual-real interaction module, based on audience position information and trajectory, determines the adjustment frequency and calculates the distance and angle deviation with the virtual character, thereby precisely adjusting the virtual character's state so that the virtual character can make natural and reasonable movements and reactions based on the audience's real-time position. Motion capture and emotion capture modules respectively collect audience body movements and emotional characteristics, enabling the virtual-real interaction module to control the virtual character to provide targeted interactive feedback. This multi-module collaborative work creates a highly realistic and interactive immersive environment for the audience, making them feel as if they are in a real storyline, fully engaged in the experience.

[0019] 2. The texture resolution of the virtual characters is adjusted based on the distance between the audience and the virtual characters. When the audience is close, the resolution is increased to make character details clearer; when the audience is far away, the resolution is decreased to save rendering resources and ensure smooth visuals. Simultaneously, lighting rendering parameters are adjusted based on the audience's emotional characteristics. When the audience expresses positive emotions, bright and warm lighting effects are rendered; when the audience expresses negative emotions, a dark and cold atmosphere is created. This dynamic rendering method can change the scene's visual effects in real time according to the audience's state, enhancing the audience's emotional resonance with the storyline and further improving the visual presentation quality of the immersive holographic theater.

[0020] 3. Based on the audience's body language data and emotional characteristics, the system intelligently matches responses to actions and emotions, and controls the virtual characters to provide interactive feedback. This transforms the virtual characters from simple programmed settings into diverse and personalized responses tailored to different audience reactions, enriching the interactive layers of the storyline. Audiences can engage deeply with the virtual characters through their own actions and expressions, experiencing their emotional changes and responses as if communicating with real characters. This intelligent matching and feedback mechanism increases audience participation and interactivity, making the immersive holographic theater's storyline more vivid and engaging, and providing viewers with a completely new entertainment experience. Attached Figure Description

[0021] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0022] Figure 1 This is a block diagram of a virtual and real character dynamic interaction control system based on an immersive holographic theater, as shown in one embodiment of this application;

[0023] Figure 2 This is a flowchart illustrating the generation of audience location information in one embodiment of this application;

[0024] Figure 3 This is a flowchart illustrating a method for dynamic interaction control of virtual and real characters based on an immersive holographic theater, as shown in one embodiment of this application. Detailed Implementation

[0025] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below.

[0026] Figure 1 This is a block diagram of a dynamic interactive control system for virtual and real characters based on an immersive holographic theater, as shown in one embodiment of this application. Figure 1 As shown, the virtual-real character dynamic interaction control system based on immersive holographic theater can include a spatial positioning module, a virtual-real interaction module, a motion capture module, an emotion capture module, and a dynamic rendering module.

[0027] The spatial positioning module is used to collect the location information of the audience within a virtual-real fusion space constructed using light field reconstruction technology.

[0028] Specifically, the spatial positioning module includes multiple UWB beacons positioned within the virtual-real fusion space and UWB tags worn by audience members. The UWB beacons, serving as fixed reference points, are strategically placed at key locations throughout the theater, such as the surrounding walls and stage perimeter. They continuously and stably transmit ultra-wideband signals at specific frequencies, constructing a precise positioning network for the entire space. The positions of these beacons are precisely determined and recorded in the system beforehand, serving as crucial benchmarks for subsequent positioning calculations. The UWB tags worn by audience members are compact devices with signal transceiver capabilities. They receive signals from multiple surrounding UWB beacons and, using a built-in high-precision algorithm, accurately calculate their own position coordinates relative to each beacon based on parameters such as the time difference of arrival (TDOA). Because multiple beacons are deployed within the theater, the tags acquire multiple sets of positioning data. By fusing and optimizing this data, the system effectively eliminates errors, further improving the accuracy and stability of the positioning.

[0029] Through this collaborative working method of UWB beacons and tags, the spatial positioning module can obtain the audience's three-dimensional position information in the virtual-real fusion space in real time and accurately, providing a reliable data foundation for the virtual-real interaction module. This enables virtual characters to make corresponding interactive actions according to the audience's actual position, thereby achieving a more natural and smooth dynamic interaction effect between virtual and real characters and bringing the audience an immersive experience.

[0030] In some embodiments, the spatial positioning module is further used for:

[0031] The virtual-real integrated space is divided into multiple areas;

[0032] Acquire the historical movement trajectories of multiple historical viewers within a virtual-real integrated space;

[0033] For each region, the comprehensive weight of the region is calculated based on the historical movement trajectories of multiple historical viewers in the virtual-real fusion space. The density of UWB beacons set up in the region is determined according to the comprehensive weight of the region.

[0034] Based on the density of UWB beacons set in each region, multiple UWB beacons are set up in the virtual-real fusion space.

[0035] Specifically, the virtual-real fusion space can be divided into multiple regions in various ways, such as dividing the virtual-real fusion space into multiple regions.

[0036] All user information (such as the historical movement trajectories of multiple historical viewers in the virtual-real fusion space, viewers' body movement data, viewers' facial images and voice data, etc.) is obtained with the user's authorization.

[0037] In some embodiments, the spatial positioning module is further configured to:

[0038] For each region, the percentage of time spent in the region is determined based on the historical movement trajectories of multiple historical viewers within the virtual-real fusion space, and the first weight of the region is calculated based on the percentage of time spent in the region.

[0039] For any two areas, determine the dwell synchronization ratio between the two areas based on the historical movement trajectories of multiple historical viewers in the virtual-real fusion space;

[0040] For each region, calculate the region's second weight based on the ratio of the region's dwell time to that of any other region.

[0041] Calculate the overall weight of a region based on its first and second weights.

[0042] Specifically, the first weight of the region can be calculated according to the following process:

[0043] For each historical visitor's historical movement trajectory within the virtual-real fusion space, calculate the time the historical visitor spends in each area within the historical movement trajectory, and calculate the ratio between the historical visitor's time spent in a single area and the total time corresponding to the historical movement trajectory, as the percentage of time spent in that area corresponding to the historical movement trajectory;

[0044] For each region, the average percentage of stays for each historical movement trajectory in that region is calculated as the region's percentage of stays. Based on this percentage of stays, the first weight of the region is calculated using the following formula:

[0045]

[0046] in, The first weight of the i-th region, Let i be the percentage of stay in the i-th region. The percentage of stay in the nth region. This represents the total number of regions.

[0047] The second weight of the region can be calculated according to the following process:

[0048] For each historical viewer's historical movement trajectory in the virtual-real fusion space and any two regions, calculate the ratio of the historical viewer's dwell time in the two regions, which is used as the dwell synchronization ratio of the corresponding historical movement trajectory in the two regions;

[0049] For any two regions, the average of the dwell synchronization ratios for each historical movement trajectory corresponding to the two regions is used as the dwell synchronization ratio between the two regions, and the second weight of the region is calculated based on the following formula;

[0050]

[0051] in, The second weight for the i-th region. Let be the synchronization ratio of the stay in region i to that in region m. This represents the synchronization ratio of stay in the nth region to that in the jth region.

[0052] The above formula aims to calculate the second weight of each region by comprehensively considering the synchronization of dwell time between regions. The numerator is the sum of the dwell time synchronization ratios of a specific region with all other regions. This reflects the overall situation of the region's correlation with other regions in terms of audience dwell time; the larger the value, the closer the correlation between the region and other regions in terms of audience dwell time. The denominator is the double sum of the dwell time synchronization ratios of all pairwise combinations of regions. It represents the sum of the dwell time synchronization ratios of all regions within the entire virtual-real integrated space, serving as a global reference. By using the ratio of the numerator to the denominator, the influence of differences in the number of different region combinations can be eliminated, allowing the degree of correlation between a specific region and other regions to be measured within the broader context of the entire space, thus obtaining a relatively reasonable second weight for that region.

[0053] The average or weighted sum of the first and second weights of a region can be used to obtain the overall weight of the region.

[0054] The density of UWB beacons set up within a region can be determined using the following formula:

[0055]

[0056] in, The density of UWB beacons set in the i-th region. Let i be the comprehensive weight of the i-th region. The initial density of UWB beacons. To Round down.

[0057] From the perspective of positioning accuracy, the comprehensive weighting system fully considers multiple dimensions of factors, including the percentage of time viewers spend within a given area and the synchronization of time spent between areas. A high percentage of time spent indicates frequent viewer activity in that area, requiring higher positioning accuracy; the synchronization ratio between areas reflects the correlation between areas and the patterns of viewer movement. Determining beacon density based on the comprehensive weighting system allows for the deployment of more beacons in areas with high viewer activity and interaction needs, reducing blind spots and achieving more accurate positioning. This enables virtual characters to make more fitting interactive actions based on the viewer's actual location, enhancing immersion.

[0058] In terms of resource utilization efficiency, this approach avoids the blind deployment of beacons. Without considering comprehensive weighting, some areas may have too many beacons, leading to resource waste, while some key areas may have insufficient beacons, affecting positioning performance. Determining beacon density through comprehensive weighting allows for the rational allocation of limited beacon resources, reducing costs and improving the economic efficiency of resource utilization while meeting positioning needs.

[0059] From an adaptability and flexibility perspective, as the content of the theater performance and audience behavior patterns change, historical movement trajectory data is dynamically updated, and the overall weighting changes accordingly, thereby adjusting the beacon density. This allows the spatial positioning system to quickly adapt to changes in different scenarios and audience needs, maintaining excellent positioning performance and providing audiences with a stable and high-quality immersive experience.

[0060] Figure 2 This is a flowchart illustrating the generation of audience location information in one embodiment of this application, as shown below. Figure 2 As shown, in some embodiments, the spatial positioning module is further used for:

[0061] Identify multiple regional positioning UWB beacons from multiple UWB beacons;

[0062] The location of the audience can be determined by using UWB tags and multiple regional UWB beacons;

[0063] Based on the comprehensive weight of the audience's location, the performance of multiple UWB beacons, and the signal transmission data of multiple UWB beacons and UWB tags, select UWB beacons for multiple targets from multiple UWB beacons;

[0064] Based on the signal transmission data of UWB beacons and UWB tags from multiple targets, the location information of the audience is generated.

[0065] Specifically, multiple UWB beacons can be selected from a pool of UWB beacons using any method. For example, through testing, the human body is positioned in different areas each time. Signal transmission data from multiple UWB beacons and UWB tags are analyzed to filter areas and determine the most effective UWB beacon combinations, thus identifying multiple UWB beacons for area positioning. For instance, during testing, a human body (representing the audience) is positioned in different areas. The human body wears a UWB tag that continuously emits specific signals, which are received by multiple surrounding UWB beacons. The different locations of the human body in each test allow for independent evaluation of signal transmission in each area. Using specialized signal acquisition equipment, detailed signal transmission data between multiple UWB beacons and UWB tags is recorded, including key information such as signal strength, time of arrival, and signal-to-noise ratio. This massive amount of signal transmission data is then analyzed in depth. Signal strength directly reflects the communication quality between beacons and tags; higher strength indicates more stable and reliable signal transmission. The accuracy of signal arrival time is crucial for positioning calculations; smaller errors result in higher positioning accuracy. The signal-to-noise ratio (SNR) reflects the degree of interference during signal transmission; a high SNR means less interference and more accurate data. Based on a comprehensive analysis of this data, the UWB beacon combinations that provide the most stable and accurate signal transmission in each area are selected. For example, in an area with frequent audience activity and numerous obstacles, only a few specific UWB beacon combinations may be found to overcome signal attenuation and interference problems, achieving accurate positioning. These UWB beacon combinations that perform best in different areas are identified as the area-specific positioning UWB beacons.

[0066] The UWB beacons located in this way take into full account the actual environment and signal transmission characteristics of each area, providing a solid foundation for accurately determining the location of the audience. This ensures that, in the immersive holographic theater, the audience can be quickly and accurately located regardless of their position, thereby enhancing the audience's immersive experience and the overall operational effectiveness of the theater.

[0067] As audience members move within the theater, their UWB tags continuously emit signals, which are simultaneously received by surrounding UWB beacons. Due to the time delay in signal propagation and the varying distances between different beacons and tags, the arrival times of the signals at each beacon also differ.

[0068] Each UWB beacon, upon receiving a signal, records its arrival time and transmits this data to the central processing system. Using any feasible algorithm, such as the Time Difference of Arrival (TDOA) algorithm, the arrival times recorded by multiple beacons are analyzed and calculated. By measuring the time difference of signal propagation between different beacons and combining this with known beacon location information, the distance between the UWB tag (i.e., the viewer) and each beacon can be accurately calculated, thus determining the viewer's location.

[0069] The performance of multiple UWB beacons encompasses indicators such as transmit power, receive sensitivity, and signal stability. Higher transmit power enhances signal coverage, higher receive sensitivity allows for precise capture of weak signals, and better signal stability reduces positioning errors. A higher-performing beacon ensures better positioning quality. Furthermore, the signal transmission data between multiple UWB beacons and UWB tags, such as signal strength, signal-to-noise ratio (SNR), and time of arrival (TOA), are crucial for evaluating the communication quality between beacons and tags. High signal strength, a high SNR, and accurate TOA indicate good communication between the beacons and tags, and reliable positioning data.

[0070] The following process can be used to filter UWB beacons from multiple UWB beacons for multiple targets:

[0071] The number of UWB beacons for the target is determined based on the overall weight of the area where the audience is located. The higher the overall weight, the more UWB beacons the target will have.

[0072] The distance between multiple UWB beacons and UWB tags is determined based on the signal transmission data of multiple UWB beacons and UWB tags;

[0073] Based on the performance of multiple UWB beacons and the distance between multiple UWB beacons and UWB tags, the multiple UWB beacons are sorted to generate a sorting result. The higher the performance and the shorter the distance, the higher the sorting result.

[0074] Based on the number of UWB beacons for the target, the top-ranked UWB beacons are selected from the sorting results. Any feasible algorithm, such as the Time Difference of Arrival (TDOA) algorithm, is used to analyze and calculate the signal arrival times of the UWB beacons for the multiple targets. By measuring the time difference of signal propagation between different beacons and combining this with the location information of the target's UWB beacons, the audience's location information is generated.

[0075] From the perspective of selecting target UWB beacons, the number of beacons is determined based on the comprehensive weight of the audience's location. This allows for the allocation of more target beacons in key areas or areas with high population density and high positioning accuracy requirements, ensuring the accuracy and reliability of positioning in critical areas. Conversely, the number of beacons is reasonably reduced in relatively less important areas to avoid resource waste. The distance between the beacon and the tag is determined based on signal transmission data, and beacon performance is prioritized. This allows for the selection of high-performance beacons that are close to the target, effectively reducing interference and errors during signal transmission. Closer beacons have less signal attenuation, and high-performance beacons have stronger data processing capabilities, thus improving the quality of positioning data.

[0076] The virtual-real interaction module is used to adjust the state of virtual characters within a virtual-real integrated space based on the audience's location information.

[0077] Specifically, it includes:

[0078] The adjustment frequency is determined based on the audience's trajectory information;

[0079] Based on the adjustment frequency, the distance between the audience and the virtual character is calculated using the audience's location information and the virtual character's current location.

[0080] Based on the audience's location information, the virtual character's current position and orientation, the angular deviation between the audience and the virtual character is calculated;

[0081] Based on the distance and angle deviation between the audience and the virtual character, the state of the virtual character is adjusted in the virtual-real integrated space.

[0082] In some embodiments, the virtual-real interaction module is further used for:

[0083] Based on the audience's trajectory information, similar historical movement trajectories are identified;

[0084] The adjustment frequency is determined based on the audience's trajectory information and similar historical movement trajectories.

[0085] Specifically, a trajectory similarity matching algorithm based on distance metrics is used to determine similar historical movement trajectories from a historical trajectory database. Specifically, the viewer's current trajectory information is discretized into a series of timestamped location point sequences, and the Euclidean distance or Dynamic Time Warping (DTW) distance between it and each sequence in the historical trajectories is calculated. By setting a distance threshold or using the K-Nearest Neighbors (KNN) algorithm, several historical trajectories most similar to the current trajectory are selected.

[0086] To quantify the fluctuation of the current trajectory and similar historical trajectories, their positional variances are calculated separately. For the current trajectory, the variance is the average of the squared Euclidean distances from each location point to its mean location; the same calculation method is used for similar historical trajectories. Positional variance reflects the dispersion of trajectory data; the larger the variance, the more drastic the trajectory fluctuations and the higher the uncertainty. The adjustment frequency is determined based on the two variances, using a linear weighting or nonlinear mapping algorithm. A simple and effective approach is to construct a linear relationship between the adjustment frequency and the variance, i.e., Adjustment Frequency = Base Frequency + k × (Positional variance of the viewer's trajectory information + Positional variance of similar historical movement trajectories), where the base frequency is the system's preset minimum adjustment frequency, and k is a proportionality coefficient that can be adjusted according to the actual application scenario and interaction requirements. The larger the two variances, the higher the calculated adjustment frequency.

[0087] Based on the set adjustment frequency, the system periodically acquires the audience's position information in the virtual-real fusion space. Taking the audience's position as the starting point and the virtual character's position as the ending point, the straight-line distance between the two is calculated using the distance calculation method between two points in spatial geometry.

[0088] When calculating the angular deviation, first determine the virtual character's current orientation, usually with the character's front as the reference direction. Decompose the orientation of the viewer's position relative to the virtual character's position into horizontal and vertical components. The angular deviation is obtained by comparing the angle between the viewer's direction and the virtual character's orientation. In practice, vector cross product and dot product operations can be used to first calculate the angle between the two direction vectors, and then the sign of the angular deviation is determined based on the virtual character's coordinate system to distinguish whether the viewer is to the left, right, above, or below the virtual character.

[0089] In a space that blends the virtual and the real, adjusting the state of the virtual character based on the distance and angle deviation between the audience and the virtual character can greatly enhance the realism and immersion of the interaction.

[0090] When the viewer is close to the virtual character, a small angle deviation means the viewer is facing the virtual character directly and at a close distance. In this case, the virtual character can display a friendly and approachable demeanor, such as slightly bowing its head, showing a gentle smile, and slowing down its movement, or even remaining still, to create a natural and harmonious interactive atmosphere. If the angle deviation is large, it means the viewer is close but to the side or behind the virtual character. The virtual character will quickly adjust its orientation, turning its body to face the viewer directly. During the turning process, the amplitude and speed of the movement will vary appropriately according to the distance; when the distance is close, the turning will be relatively smooth to avoid the movement being too abrupt.

[0091] When the audience is far from the virtual character, if the angular deviation is small, the virtual character will maintain a certain speed and move towards the audience, adjusting the size of its steps according to the distance. Larger steps indicate a faster approach, while smaller steps indicate a more relaxed and natural movement. If the angular deviation is large, the virtual character will first quickly adjust its orientation, taking the shortest path to the audience's location. After turning, it will then choose an appropriate movement strategy based on the distance, gradually approaching the audience.

[0092] Throughout the adjustment process, changes in distance and angle deviations are continuously monitored in real time, and the virtual character's state is dynamically optimized based on the latest data. For example, as the distance and angle deviations change during the viewer's movement, the virtual character will flexibly adjust its movement speed, direction, and posture to maintain an appropriate interactive relationship with the viewer, making the viewer feel as if they are in a real and vibrant virtual-real fusion world, gaining a richer and more vivid interactive experience.

[0093] The motion capture module is used to collect data on the audience's body movements.

[0094] Specifically, the motion capture module uses an array of high-definition cameras to capture the audience from multiple angles, ensuring that subtle movements of all parts of the audience's body are fully captured and avoiding data loss due to blind spots. The captured images are transmitted to the processing unit at high speed to ensure the real-time nature and continuity of the movements. In the image processing stage, object detection algorithms from the field of computer vision are used to quickly locate the audience's body regions in the image and accurately separate them from the background. This effectively eliminates background interference, allowing subsequent processing to focus on the audience's limbs themselves. Next, pose estimation technology is used to identify and locate key points of the audience's body through a deep learning model. These key points cover important joints such as the head, shoulders, elbows, wrists, hips, knees, and ankles. The model has been trained on a large number of samples and has high-precision recognition capabilities, accurately identifying key points even when the audience's movements are large and the posture is complex. After determining the key points, 3D reconstruction is performed based on the coordinate information of these points in the image, combined with camera parameters and spatial geometric relationships. Through complex mathematical calculations, the coordinates of key points in a two-dimensional image are converted into actual position coordinates in three-dimensional space, thereby obtaining the motion trajectory and posture data of the audience's limbs in three-dimensional space.

[0095] The emotion capture module is used to collect facial images and voice data of the audience, and extract the audience's emotional features from the facial images and voice data.

[0096] Specifically, for facial images, the emotion capture module uses a high-resolution camera to quickly capture images of the viewer's face. Then, using advanced computer vision algorithms, the image is preprocessed, such as adjusting brightness and contrast, and removing noise to improve image quality. Next, facial feature point localization technology is used to accurately identify key facial features, such as the contours and positions of eyebrows, eyes, and mouth. Based on these feature points, a deep learning model analyzes changes in facial muscle movements, such as the degree of upward movement of the corners of the mouth, and the extent to which eyebrows are raised or furrowed, to determine whether the viewer is in an emotional state such as happy, surprised, angry, or sad. Because different emotions trigger specific facial muscle movement patterns, the model, trained on a large number of facial images labeled with emotions, can accurately identify these patterns and extract corresponding emotional features.

[0097] Regarding voice data, the emotion capture module collects audience voice data in real time via microphone. The voice signal is first processed through pre-emphasis, framing, and windowing to enhance high-frequency components and stabilize the signal. Then, speech recognition technology is used to convert the speech into text, simultaneously extracting acoustic features such as pitch, speech rate, volume, and sound quality. Pitch variations reflect the level of emotional excitement, speech rate suggests tension or relaxation, and volume reflects emotional intensity. By analyzing these acoustic features and combining them with text content, machine learning algorithms are used to determine the audience's emotional state and extract emotion-related features.

[0098] The virtual-real interaction module is also used to control virtual characters to provide interactive feedback based on the audience's body movement data and emotional characteristics.

[0099] Specifically, it includes:

[0100] Match response actions based on audience body movement data and emotional characteristics;

[0101] Based on the audience's body movement data and emotional characteristics, match the response emotions;

[0102] Based on the responding actions and emotions, control the virtual character to provide interactive feedback.

[0103] Specifically, in matching response actions, it selects the most suitable response action from a vast pre-set action library based on the viewer's intention in their body language and the needs of the scene. For example, when a viewer makes an inviting gesture, the virtual character might respond with an elegant bow; if the viewer waves their arm quickly, the virtual character might jump excitedly. This matching is not a simple mechanical correspondence, but fully considers the fluidity, naturalness, and coordination with the overall scene of the action to ensure that the response action realistically and vividly reflects the virtual character's understanding and feedback to the viewer's actions. As an example, by analyzing and learning from a large number of sample actions using deep learning algorithms, an action intent recognition model is built, which can accurately determine the viewer's action intent based on the captured data, such as distinguishing the meaning behind different actions like inviting, waving, and pointing. At the same time, combined with scene perception technology, using environmental sensors or pre-set scene information, it clarifies the current scene requirements, such as whether it is a warm social scene or an intense competitive scene. Then, it selects response actions from a vast pre-set action library, which is created using computer graphics technology and contains various styles and types of actions, each with detailed parameter descriptions. The selection process employs an intelligent matching algorithm that comprehensively considers the action's intent, scene requirements, and the smoothness, naturalness, and harmony of the action with the overall scene. The algorithm scores and ranks the actions in the action library, selecting the highest-scoring action as the response action to ensure the virtual character's response is realistic and vivid.

[0104] The virtual-real interaction module also matches responses to the viewer's emotional characteristics. It identifies whether the viewer is currently happy, surprised, angry, or sad, and then makes the virtual character display the corresponding emotion. Emotional mapping technology is used to convert these emotions into expressible emotional parameters for the virtual character. For example, happiness is mapped to specific contractions of the virtual character's facial muscles and brightness of its eyes, resulting in a bright smile and joyful eyes. Sadness is mapped to changes in parameters such as a slight bow of the head and dimmed eyes, allowing the virtual character to express concern and sympathy.

[0105] The dynamic rendering module is used to adjust rendering parameters in real time based on the audience's location information and emotional characteristics to create a virtual-real fusion space.

[0106] Specifically, it includes:

[0107] Adjust the texture resolution of the virtual character based on the distance between the viewer and the virtual character;

[0108] Adjust lighting and rendering parameters based on the audience's emotional characteristics.

[0109] Specifically, when the viewer is far from the virtual character, the dynamic rendering module automatically reduces the texture resolution of the virtual character because the human eye has limited ability to distinguish details at long distances. This not only reduces the system's consumption of graphics processing resources and improves rendering efficiency, avoiding stuttering caused by over-rendering distant details, but also ensures the smoothness of the overall image. As the viewer gradually approaches the virtual character, the module rapidly increases the texture resolution, making the surface details of the virtual character, such as skin texture and clothing wrinkles, more clearly visible, allowing the viewer to truly feel the realistic texture of the virtual character and enhancing the immersive experience during close-range interaction.

[0110] When the module detects that the viewer is in a happy mood, it increases the brightness and saturation of the lighting to create a bright, warm, and cheerful atmosphere, making the colors of the virtual characters more vibrant and eye-catching, matching the viewer's joyful mood. If the viewer expresses sadness, the module reduces the lighting intensity, using softer, cooler-toned light to create a slightly depressing and quiet environment, making the virtual characters seem to be immersed in sadness as well, evoking emotional resonance from the viewer. Through this lighting adjustment based on emotional characteristics, the virtual-real integrated space can better match the viewer's emotional changes, further enhancing the realism and appeal of the interaction, and bringing the viewer a comprehensive and immersive virtual-real interactive experience.

[0111] Figure 3 This is a flowchart illustrating a method for dynamic interaction control of virtual and real characters based on an immersive holographic theater, as shown in one embodiment of this application. Figure 3 As shown, the method for dynamic interaction control of virtual and real characters based on immersive holographic theater may include the following steps:

[0112] Within a virtual-real fusion space constructed using light field reconstruction technology, the location information of the audience is collected;

[0113] Based on the audience's location information, the state of the virtual character is adjusted within the virtual-real integrated space;

[0114] Collect audience body movement data;

[0115] Collect facial images and voice data of the audience, and extract the audience's emotional features from the facial images and voice data;

[0116] Based on the audience's body movement data and emotional characteristics, control the virtual character to provide interactive feedback;

[0117] Based on the audience's location information and emotional characteristics, the rendering parameters of the virtual-real fusion space are adjusted in real time.

[0118] The method for dynamic interaction control of virtual and real characters based on immersive holographic theater can be applied to the dynamic interaction control system of virtual and real characters based on immersive holographic theater, which will not be elaborated here.

[0119] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.

Claims

1. A dynamic interactive control system for virtual and real characters based on an immersive holographic theater, characterized in that, include: The spatial positioning module is used to collect audience location information within a virtual-real fusion space constructed using light field reconstruction technology. This includes: dividing the virtual-real fusion space into multiple regions; for each region, calculating a comprehensive weight based on the historical movement trajectories of multiple historical audience members within the virtual-real fusion space obtained with user authorization; determining the density of UWB beacons set within the region based on the comprehensive weight; setting multiple UWB beacons within the virtual-real fusion space based on the density of UWB beacons in each region; filtering multiple target UWB beacons from the multiple UWB beacons based on the signal transmission data between the multiple UWB beacons and UWB tags placed on the audience members; and generating audience location information based on the signal transmission data between the multiple target UWB beacons and UWB tags. The virtual-real interaction module is used to adjust the state of virtual characters within the virtual-real fusion space based on the audience's location information; The motion capture module is used to collect the audience's body movement data with the user's authorization; The emotion capture module is used to collect facial images and voice data of the audience with the user's authorization, and extract the audience's emotional features from the facial images and voice data; The virtual-real interaction module is also used to control the virtual character to provide interactive feedback based on the audience's body movement data and emotional characteristics. The dynamic rendering module is used to adjust rendering parameters in real time based on the audience's location information and emotional characteristics to blend the virtual and real spaces. The calculation of the comprehensive weight of a region based on the historical movement trajectories of multiple historical viewers within the virtual-real fusion space includes: For each historical visitor's historical movement trajectory within the virtual-real fusion space, the dwell time of the historical visitor in each area within the historical movement trajectory is calculated. The ratio between the dwell time of the historical visitor in a single area and the total time corresponding to the historical movement trajectory is calculated as the dwell percentage of that area corresponding to that historical movement trajectory. For each area, the average dwell percentage of that area corresponding to each historical movement trajectory is calculated as the dwell percentage of that area. Based on the dwell percentage of the area, the first weight of the area is calculated. For any two areas, determine the dwell synchronization ratio between the two areas based on the historical movement trajectories of multiple historical viewers in the virtual-real fusion space; For each historical viewer's historical movement trajectory in the virtual-real fusion space and any two regions, calculate the ratio of the historical viewer's dwell time in the two regions, which is used as the dwell synchronization ratio of the corresponding historical movement trajectory in the two regions; for any two regions, calculate the average of the dwell synchronization ratios of each historical movement trajectory in the two regions, which is used as the dwell synchronization ratio of the two regions, and calculate the second weight of the region. Calculate the overall weight of a region based on its first and second weights.

2. The virtual and real character dynamic interaction control system based on immersive holographic theater according to claim 1, characterized in that, The spatial positioning module is further used for: Identify multiple regional positioning UWB beacons from multiple UWB beacons; The location of the audience can be determined by using UWB tags and multiple regional UWB beacons; Based on the comprehensive weight of the audience's location, the performance of multiple UWB beacons, and the signal transmission data of multiple UWB beacons and UWB tags, select UWB beacons for multiple targets from multiple UWB beacons; Based on the signal transmission data of UWB beacons and UWB tags from multiple targets, the location information of the audience is generated.

3. The virtual-real character dynamic interaction control system based on an immersive holographic theater according to claim 1 or 2, characterized in that, The virtual-real interaction module is further used for: The adjustment frequency is determined based on the audience's trajectory information; Based on the adjustment frequency, the distance between the audience and the virtual character is calculated using the audience's location information and the virtual character's current location. Based on the audience's location information, the virtual character's current position and orientation, the angular deviation between the audience and the virtual character is calculated; Based on the distance and angle deviation between the audience and the virtual character, the state of the virtual character is adjusted in the virtual-real integrated space.

4. The virtual and real character dynamic interaction control system based on immersive holographic theater according to claim 3, characterized in that, The virtual-real interaction module is further used for: Based on the audience's trajectory information, similar historical movement trajectories are identified; The adjustment frequency is determined based on the audience's trajectory information and similar historical movement trajectories.

5. The virtual and real character dynamic interaction control system based on immersive holographic theater according to claim 3, characterized in that, The virtual-real interaction module is further used for: Match response actions based on audience body movement data and emotional characteristics; Based on the audience's body movement data and emotional characteristics, match the response emotions; Based on the responding actions and emotions, control the virtual character to provide interactive feedback.

6. The virtual and real character dynamic interaction control system based on immersive holographic theater according to claim 3, characterized in that, The dynamic rendering module is further used for: The texture resolution of the virtual character is adjusted based on the distance between the viewer and the virtual character.

7. The virtual and real character dynamic interaction control system based on immersive holographic theater according to claim 6, characterized in that, The dynamic rendering module is further used for: Adjust lighting and rendering parameters based on the audience's emotional characteristics.

8. A method for dynamic interaction control of virtual and real characters based on immersive holographic theater, characterized in that, The virtual and real character dynamic interaction control system based on immersive holographic theater as described in claim 1 includes: Within a virtual-real fusion space constructed using light field reconstruction technology, the location information of the audience is collected; Based on the audience's location information, the state of the virtual character is adjusted within the virtual-real integrated space; With user authorization, collect audience body movement data; With user authorization, facial images and voice data of the audience are collected, and emotional features of the audience are extracted from the facial images and voice data. Based on the audience's body movement data and emotional characteristics, control the virtual character to provide interactive feedback; Based on the audience's location information and emotional characteristics, the rendering parameters of the virtual-real fusion space are adjusted in real time.

Citation Information

Patent Citations

  • VR large-space positioning interaction system based on multi-modal perception

    CN120469587A

  • Immersive concert experience system based on virtual reality

    CN120765890A