Immersive spatial three-dimensional interaction method and system based on artificial intelligence

By using multimodal signal fusion processing and AI decision-making strategies, the coupling interference problem between visual data and infrared touch signals was solved, improving the accuracy of interactive response and user experience in immersive spaces.

CN121505151APending Publication Date: 2026-02-10GUANGZHOU WUKONG CREATIVE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511518086.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-02-10

Smart Images

  • Figure CN121505151A_ABST
    Figure CN121505151A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of three-dimensional interaction, in particular to an immersive space three-dimensional interaction method and system based on artificial intelligence. A three-dimensional model of a four-fold-screen or five-fold-screen immersive space is constructed through a UE5 ground editing technology, three-dimensional visualization is achieved, at least eight 4K industrial cameras are adopted to collect user visual data, and at least 32 infrared touch sensors are adopted to collect user touch signals; performing fusion weight calculation on the visual data and the touch signal to realize multi-mode signal fusion processing, and eliminating coupling interference of visual data missing and infrared touch false triggering caused by user shielding; generating an AI interaction decision strategy by combining the user position, the user behavior intention and the space environment parameters based on the fused signal; according to the invention, the coupling interference of visual data and infrared touch signals can be effectively solved, and the interactive response reliability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of three-dimensional interactive technology, and in particular to an immersive spatial three-dimensional interactive method and system based on artificial intelligence. Background Technology

[0002] Immersive spaces, with their panoramic visual presentation and interactive experience, have been widely used in cultural tourism exhibitions, science education, virtual simulation, and other fields. Currently, such spaces typically use UE5 ground-editing technology to construct three-dimensional visual scenes, and use multiple cameras to collect user visual data and infrared touch sensors to collect user touch signals, in order to enable interaction between users and the scene and digital humans.

[0003] However, in existing technologies, visual data and infrared touch signals are processed independently, without considering the coupling interference between the two in practical applications. When there are many users in the space or when users obstruct the view, the camera is prone to visual data loss or distortion due to obstructed vision. At the same time, infrared touch signals are easily affected by factors such as changes in ambient light intensity and differences in the distance between the user and the screen, resulting in false triggers or weak signals. Because the two signals are processed independently, the system cannot dynamically complement each other based on signal quality, leading to a significant decrease in the accuracy of interactive response. This can even result in problems such as no response when the user initiates a touch or the system misinterpreting non-interactive behavior, severely disrupting the interaction between the user and the immersive scene and digital human, and weakening the continuity and intelligence of the immersive experience.

[0004] Based on the above problems, there is an urgent need for a technical solution that can effectively solve the coupling interference between visual data and infrared touch signals and improve the reliability of interactive response, so as to improve the interactive experience of immersive spaces. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and propose an immersive spatial three-dimensional interactive method based on artificial intelligence, comprising: The 3D model of a four-fold or five-fold immersive space is constructed and visualized using UE5 ground editing technology. At least 8 4K industrial cameras are used to collect user visual data and at least 32 infrared touch sensors are used to collect user touch signals. The visual data and the touch signal are fused together using a fusion weight calculation to achieve multimodal signal fusion processing, thereby eliminating the coupling interference between visual data loss caused by user occlusion and false triggering of infrared touch. Based on the fused signal, AI interactive decision-making strategies are generated by combining user location, user behavior intent and spatial environment parameters. The AI ​​interactive decision-making strategy dynamically adjusts the lighting parameters, plot progress, and digital human interaction behavior of the 3D scene, while simultaneously adjusting the vertex count and rendering frame rate of the 3D model to solve the problem of disconnect between users and scenes and digital human interaction in immersive spaces, thereby enhancing the intelligence of the immersive experience.

[0006] Preferably, the multimodal signal fusion processing further includes a preprocessing step, in which the visual data is processed by a modal decomposition algorithm in the field of mechanical vibration monitoring to extract effective visual features of user contour and gaze direction to achieve noise reduction, and the touch signal is processed by an interference feature library matching algorithm to filter effective touch signals of user touch pressure and position to achieve filtering. The interference feature library is constructed by pre-collecting waveform features of common interference signals in the immersive space. The modal decomposition algorithm and the interference feature library matching algorithm are stored in the storage unit of the FPGA processor and are called and executed by the FPGA processor.

[0007] In a further preferred embodiment, in the step of generating an AI interactive decision-making strategy by combining user location, user behavioral intent, and spatial environment parameters, the user behavioral intent is obtained by weighting the user's gesture dwell time and gaze focus time, with a weighting coefficient of 0.6 for gesture dwell time and 0.4 for gaze focus time; the user group interaction needs are obtained by weighting the number of users and the consistency of group actions, with a weighting coefficient of 0.3 for the number of users and 0.7 for the consistency of group actions; the spatial environment parameters include ambient light intensity and color deviation at the screen splicing point, wherein the ambient light intensity is used to calculate the difference between the ambient light and the screen display brightness, and the color deviation is used to calculate the color uniformity at the screen splicing point.

[0008] Further preferably, in the step of collaboratively adjusting the number of vertices and the rendering frame rate of the 3D model, the adjusted number of vertices and the rendering frame rate data of the 3D model are collected in real time and fed back to the multimodal signal fusion processing step to form a closed-loop processing flow, with a closed-loop processing cycle of no more than 50ms; the rendering frame rate adjustment is based on no less than 30fps, combined with dynamic adaptation of the number of 3D model vertices, and the adjustment of the number of 3D model vertices does not exceed 2×10 6 The upper limit is set to ensure a coordinated balance between the accuracy of the 3D model and the smoothness of rendering.

[0009] More preferably, the formula for calculating the fusion weight in the multimodal signal fusion processing is: ; in: The multimodal signal fusion weights at time t are... This is the basic weighting coefficient for the signal, and its value ranges from 0.4 to 0.6. The effective visual data intensity at time t is expressed in cd / m². Let be the ambient light interference coefficient at time t, with a value between 0.1 and 1.0, and n be the total number of cameras, where n ≥ 8. Let be the raw visual data intensity of the i-th camera at time t, in cd / m². The effective infrared touch signal strength at time t is expressed in V. Let t be the user touch distance coefficient, with a value between 0.2 and 1.0, and m be the total number of infrared touch sensors, where m ≥ 32. Let be the original touch signal strength of the j-th infrared sensor at time t, expressed in V. To avoid the minimum value of denominator 0 and take the value 10 -6 .

[0010] More preferably, the decision priority calculation formula for the AI ​​interactive decision-making strategy is as follows: ; in: Let t be the priority of the interactive decision and take a value between 0 and 100. For multimodal signal fusion weights, This is the behavioral intent weighting coefficient, with a value ranging from 0.5 to 0.8. Let t be the intensity of the user's behavioral intent, with a value between 0 and 50. Let t be the intensity of user group interaction demand, with a value ranging from 0 to 50. The ambient light influence coefficient ranges from 0.6 to 0.9. Let be the ambient light intensity required for adaptation at time t, and let its value be 0-50. The required intensity of screen display deviation correction at time t is 0-50.

[0011] More preferably, the formula for calculating the vertex simplification rate of the collaboratively adjusted 3D model's vertex count and rendering frame rate is: ; In the formula Let be the vertex simplification factor of the 3D model at time t, and take a value between 0 and 1. is the maximum vertex simplification rate, with a value of 0.8, and k is the collaborative control coefficient, with a value between 0.02 and 0.05. Prioritizing interactive decision-making Render the load coefficient for the model at time t, with a value between 0 and 1. Let t be the number of vertices in the current model, expressed in units of vertices. The maximum number of vertices in the model, with a value of 2 × 10. 6 indivual, Let t be the current rendering frame rate in fps. The minimum guaranteed frame rate is set to 30fps.

[0012] An immersive spatial 3D interactive system based on artificial intelligence, applied to any of the above-described immersive spatial 3D interactive methods based on artificial intelligence, includes: a 3D model construction module, a multimodal acquisition module, a signal fusion processing module, an AI decision-making module, and a scene execution module; The 3D model building module uses UE5 ground editing technology to build a 3D model of a four-fold or five-fold immersive space and realize 3D visualization; The multimodal acquisition module includes at least 8 4K industrial cameras and at least 32 infrared touch sensors. The 4K industrial cameras are used to acquire visual data of the user's contour and gaze direction, and the infrared touch sensors are used to acquire touch signals of the user's touch pressure and position. The signal fusion processing module includes an FPGA processor, which is used to perform multimodal signal fusion processing on the visual data and the touch signal; The AI ​​decision-making module includes an edge computing server, which is used to generate AI interactive decision-making strategies based on the fused signals. The scene execution module includes a UE5 rendering engine, a screen display controller, and a digital human rendering unit. The scene execution module is used to adjust the interaction behavior between the 3D scene and the digital human according to the AI ​​interaction decision strategy. Each module realizes signal interaction through the PCIe 4.0 bus.

[0013] More preferably, the FPGA processor of the signal fusion processing module is a Xilinx Kintex UltraScale, and the FPGA processor is also connected to a storage unit, which is used to store the mode decomposition algorithm program, interference feature library data and multi-mode signal fusion weight calculation program; The FPGA processor is connected to the multimodal acquisition module via the SPI bus, receives visual data and touch signals output by the multimodal acquisition module, and calls the program in the storage unit to preprocess and fuse the visual data and touch signals.

[0014] More preferably, the screen display controller of the scene execution module adopts a DMX512 protocol controller, which is used to adjust the screen brightness, color temperature and other light and shadow parameters. The digital human rendering unit includes a GPU rendering card, which is an NVIDIA RTX A6000. The GPU rendering card is used to drive the digital human to initiate interactive behaviors such as gestures and voice. The GPU rendering card is connected to the edge computing server of the AI ​​decision module via a PCIe 4.0 bus, and receives interactive decision commands output by the edge computing server.

[0015] Technical effects: This invention can eliminate the coupling interference between visual data and infrared touch signals, solving the problem of low interactive response accuracy caused by the independent processing of the two in the background technology; at the same time, it combines user position, behavioral intention and spatial environment parameters to generate AI interactive decision-making strategies, and coordinates the adjustment of the number of vertices of the 3D model and the rendering frame rate, so that the interactive response fits the user's needs, avoids the fragmentation of the immersive experience, significantly enhances the reliability and intelligence of immersive spatial interaction, and ensures the continuity of the immersive experience. Attached Figure Description

[0016] Figure 1 This is a flowchart of an immersive spatial 3D interactive method based on artificial intelligence, as described in this application. Figure 2 This is a connection diagram of an immersive spatial three-dimensional interactive system based on artificial intelligence, as described in this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0018] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0019] Traditional technical solutions suffer from the following problems: fragmented multimodal signal processing, with visual data and infrared touch signals processed independently, failing to eliminate the coupling interference between missing visual data caused by user occlusion and false infrared triggering; AI decision-making lacks multi-dimensional parameter support, relying solely on single user location data without combining behavioral intent and spatial environment parameters, resulting in a disconnect between scene adjustments and user needs; and difficulty in coordinating 3D model accuracy and rendering frame rate, either resulting in insufficient frame rate due to high accuracy, or sacrificing accuracy for high frame rate, thus disrupting the immersive experience.

[0020] Based on this, please refer to Figure 1 This embodiment provides an immersive spatial 3D interaction method based on artificial intelligence, including: S1: Construct a 3D model of a four-fold or five-fold immersive space using UE5 ground-editing technology and achieve 3D visualization. Use at least 8 4K industrial cameras to collect user visual data and at least 32 infrared touch sensors to collect user touch signals. S2: Multimodal signal fusion processing is achieved by calculating fusion weights for the visual data and the touch signal to eliminate the coupling interference between visual data loss caused by user occlusion and false triggering of infrared touch. S3: Based on the fused signal, combined with user location, user behavior intent and spatial environment parameters, an AI interactive decision-making strategy is generated; S4: Dynamically adjust the lighting parameters, plot progress, and digital human interaction behavior of the 3D scene according to the AI ​​interactive decision-making strategy, and simultaneously adjust the number of vertices and rendering frame rate of the 3D model to solve the problem of disconnect between users and scenes and digital human interaction in immersive spaces, and enhance the intelligence of the immersive experience.

[0021] This solution breaks down data fragmentation through multimodal signal fusion, supports AI decision-making with multi-dimensional parameters, and forms a closed loop of signal processing, decision generation, and final scene execution through coordinated adjustment of accuracy and frame rate. This effectively solves the problems of interaction fragmentation, decision lag, and accuracy-frame rate contradiction in existing solutions, making the interactive response of immersive space more in line with user behavior and the scene presentation smoother and more accurate.

[0022] Traditional technical solutions suffer from the following problems: visual data denoising does not incorporate cross-domain algorithms and only uses conventional filtering, which fails to accurately extract key features such as user contours and gaze direction; infrared touch signal filtering lacks targeted interference matching and has poor suppression of common interference signals in space; and the storage and execution carriers of preprocessing algorithms are not clearly defined, making it difficult to implement the solutions.

[0023] Based on this, the multimodal signal fusion processing also includes a preprocessing step. The visual data is processed using a modal decomposition algorithm from the field of mechanical vibration monitoring to extract effective visual features of the user's contour and gaze direction to achieve noise reduction. The touch signal is processed using an interference feature library matching algorithm to filter effective touch signals of the user's touch pressure and position to achieve filtering. The interference feature library is constructed by pre-collecting waveform features of common interference signals in the immersive space. The modal decomposition algorithm and the interference feature library matching algorithm are stored in the storage unit of the FPGA processor and are called and executed by the FPGA processor.

[0024] This solution introduces a modal decomposition algorithm from the field of mechanical vibration to process visual data, which can accurately separate effective features from noise; it improves the purity of touch signals by selectively filtering infrared signals through a pre-built interference feature library; and it specifies FPGA as the storage and execution carrier of the algorithm to ensure that the preprocessing process is feasible.

[0025] Ultimately, this improves the accuracy of the preprocessed data, providing high-quality input for subsequent multimodal fusion and avoiding interactive misjudgments caused by noise in the original data.

[0026] Traditional technical solutions have the following technical problems: the calculation of user behavior intent lacks clear quantitative logic and relies solely on subjective judgment, which cannot accurately reflect the user's proactive needs; the needs of user group interaction are not included in the decision-making process, ignoring the characteristics of group behavior in multi-person scenarios; the application of spatial environmental parameters is too general, and the correlation between ambient light intensity, color deviation and decision-making is not clearly defined, resulting in a disconnect between decision-making and the actual state of the space.

[0027] Based on this, in the step of generating AI interactive decision-making strategies by combining user location, user behavioral intent, and spatial environment parameters, user behavioral intent is obtained by weighting the user's gesture dwell time and gaze focus time, with a weighting coefficient of 0.6 for gesture dwell time and 0.4 for gaze focus time; user group interaction needs are obtained by weighting the number of users and the consistency of group actions, with a weighting coefficient of 0.3 for the number of users and 0.7 for the consistency of group actions; spatial environment parameters include ambient light intensity and color deviation at the screen splicing point, where ambient light intensity is used to calculate the difference between ambient light and screen display brightness, and color deviation is used to calculate the color uniformity at the screen splicing point.

[0028] This solution quantifies individual and group user needs through weighted calculations, making decisions more aligned with users' actual intentions. It transforms environmental parameters into specific computational indicators, such as brightness difference and color uniformity, enabling decisions to adapt to changes in the spatial environment. The resulting AI decision-making strategy is more accurate, avoiding interaction gaps caused by misjudgments of needs or poor environmental adaptation, and improving the fit between users and the scenario.

[0029] Traditional technical solutions have the following technical problems: lack of closed-loop feedback mechanism, the adjusted vertex count and frame rate data cannot be fed back to the front-end processing in real time, resulting in adjustment lag; there is no clear benchmark for adjusting the rendering frame rate and model vertex count, relying only on empirical values, making it difficult to balance accuracy and smoothness; the adjustment logic is not related to interaction requirements, and cannot be dynamically adapted according to interaction priority, resulting in insufficient frame rate or insufficient accuracy during high-priority interactions.

[0030] Based on this, in the step of collaboratively adjusting the number of vertices and rendering frame rate of the 3D model, the adjusted number of vertices and rendering frame rate data of the 3D model are collected in real time and fed back to the multimodal signal fusion processing step to form a closed-loop processing flow with a closed-loop processing cycle of no more than 50ms; the rendering frame rate adjustment is based on no less than 30fps, combined with dynamic adaptation of the number of vertices of the 3D model, and the number of vertices of the 3D model is adjusted to no more than 2×106, so as to ensure the coordinated balance between the accuracy of the 3D model and the smoothness of rendering.

[0031] This solution achieves real-time adjustments through closed-loop feedback within 50ms to avoid lag; it clearly defines the adjustment boundaries using 30fps as the frame rate benchmark and 2×106 vertices as the upper limit; and it links accuracy and smoothness through dynamic adaptation logic. Ultimately, this ensures that the 3D model maintains appropriate accuracy and frame rate in different interactive scenarios, avoiding a decline in immersive experience due to improper adjustments and ensuring stable and smooth scene presentation during high-priority interactions.

[0032] Traditional technical solutions have the following technical problems: they do not dynamically combine ambient light and touch distance factors, the weight allocation is fixed and cannot adapt to different spatial states and user positions; the parameter definitions are vague and lack clear dimensions and value ranges, resulting in unreliable calculation results; they do not consider the aggregation and processing of data from multiple devices, and only rely on data from a single device, resulting in one-sided fusion results.

[0033] Based on this, the formula for calculating the fusion weight in the multimodal signal fusion processing is as follows: ; In the formula The multimodal signal fusion weight at time t is dimensionless and ranges from 0 to 1. It is used to allocate the fusion ratio of visual and touch signals. The fundamental weighting coefficient for the signal is dimensionless and ranges from 0.4 to 0.6, balancing the fundamental contributions of visual and touch signals. The effective visual data intensity at time t is expressed in cd / m², reflecting the clarity of the user's visual features. The ambient light interference coefficient at time t is dimensionless and ranges from 0.1 to 1.0. The stronger the ambient light, the smaller the coefficient, which weakens the weight of visual data. n is the total number of cameras, dimensionless, and n≥8 to ensure spatial visual coverage. Let be the raw visual data intensity of the i-th camera at time t, with the dimension cd / m², representing the basic value of the visual data collected by a single camera. Let t be the effective infrared touch signal strength at time t, with the dimension V, which reflects the effective signal amplitude of the user's touch behavior; The user touch distance coefficient at time t is dimensionless and ranges from 0.2 to 1.0. The closer the user is to the screen, the larger the coefficient becomes, thus enhancing the weight of the infrared signal. m is the total number of infrared touch sensors, dimensionless, and m ≥ 32 to ensure touch detection accuracy. Let be the original touch signal intensity of the j-th infrared sensor at time t, with the dimension V, representing the basic value of the touch data collected by a single infrared sensor; To avoid the minimum value where the denominator is 0, it is dimensionless and takes the value 10. -6 This ensures the validity of the calculations.

[0034] In this embodiment, The basic weighting coefficient for the signal is set at 0.4-0.6, which is not set arbitrarily, but is based on the interactive characteristics of immersive space: visual data such as user outline and gaze are the core of judging user intent, while infrared touch signals such as user touch actions are direct interactive commands. The two are similar in basic importance but need to be slightly emphasized to avoid deviation caused by a single signal completely dominating.

[0035] This combination of numerator and denominator is a key design element for achieving dynamic modulation of visual signals: [The molecule contains...] Effective visual data intensity and Multiplying the ambient light interference coefficient essentially corrects the effectiveness of visual data by using the ambient light factor.

[0036] When the ambient light is too strong, such as when the glass curtain wall of the exhibition hall reflects light, It will drop to the 0.1-0.3 range, even if Because the camera hardware performance remains high, the product of the two will also be significantly reduced, thereby weakening the weight of the visual signal in the fusion and avoiding misjudgment caused by blurred visual data under strong light; the denominator is the sum of the original visual data from all cameras. This is to avoid the problem of partial data loss caused by user obstruction, such as when multiple people are interacting simultaneously, and to improve the overall reliability of visual data by aggregating data from multiple devices. Take 10 -6 This is a necessary design in engineering implementation to prevent extreme situations, such as when all cameras are temporarily blocked. A denominator of 0 causes a calculation crash, ensuring the formula's executability across all scenarios.

[0037] The second half of the formula A logic that forms a symmetry and complement to the visual signal: within the molecule The user touch distance coefficient is designed to take advantage of the characteristic that infrared touch signals are more accurate the closer they are. When the user approaches the screen, such as reaching out to touch a four-fold screen, Increased to 0.8-1.0, enhancing effective touch signals. The weighting of the data prevents the infrared signal from being easily interfered with when the user is away from the screen, such as accidental triggering caused by other users touching the screen. The denominator summarizes all infrared sensor data, which is also to improve the reliability of the touch signal through multi-device redundancy. This works in conjunction with the denominator logic of the visual signal part to jointly solve the technical problem of unreliable data from a single device.

[0038] This formula incorporates key factors affecting signal quality, such as ambient light interference, user location, and multi-device redundancy, into the weight calculation through the coordinated design of various parameters. This ensures that the fused data can truly reflect the user's interaction intent and fully supports the technical feature of eliminating the coupling interference between visual data loss caused by user occlusion and false triggering of infrared touch.

[0039] The formula is passed and Dynamically adjust weights to adapt to changes in environment and user location; clearly define the dimensions and value ranges of each parameter to ensure reliable calculations; aggregate data from multiple devices... This ensures a comprehensive fusion result. The final calculated fusion weights can accurately allocate the proportion of visual and touch signals, eliminate coupling interference, improve the fusion accuracy of multimodal data, and provide reliable input for subsequent decision-making.

[0040] There are technical problems with the current AI interactive decision priority calculation: First, the multimodal signal fusion weights are not associated, and the decision priority is decoupled from the signal quality, resulting in decisions based on low-quality signals still being executed first; Second, the weight allocation of behavioral intentions, group needs, and environmental parameters has no clear basis, the parameter definitions are vague, and it is impossible to quantify and calculate; Third, the priority value range is not clear, and it is impossible to judge the urgency of the decision, resulting in a chaotic execution order.

[0041] Based on this, the formula for calculating the decision priority of the AI ​​interactive decision-making strategy is as follows: ; in, The priority of interactive decision at time t is dimensionless and ranges from 0 to 100. The higher the value, the higher the priority of the decision. The multimodal signal fusion weights are dimensionless and correlated with signal quality, ensuring that decisions supported by high-quality signals are more reliable. This is a weighting coefficient for behavioral intent, dimensionless, ranging from 0.5 to 0.8, prioritizing user-initiated behavior; Let t be the intensity of user behavior intent at time t, which is dimensionless and ranges from 0 to 50, quantifying the user's need for proactive interaction. Let t represent the intensity of user group interaction demand, which is dimensionless and ranges from 0 to 50, quantifying the group demand in a multi-person scenario. The ambient light influence coefficient is dimensionless and ranges from 0.6 to 0.9, taking into account the impact of ambient light on the scene. The ambient light adaptation requirement intensity at time t is dimensionless, ranging from 0 to 50, quantifying the adaptation requirements between ambient light and screen brightness. The intensity of the screen display deviation correction requirement at time t is dimensionless, ranging from 0 to 50, and quantifies the correction requirement for screen color deviation.

[0042] in the formula and The weight allocation is the core of achieving signal quality-oriented decision-making: when When the visual signal quality is high, such as 0.7-0.9, and there is no obstruction, the decision-making process will prioritize user-related needs. In some cases, the user's proactive behaviors, such as gestures and eye focus, are the core of the decision-making process; when At lower values, such as 0.1-0.3, when visual signals are interfered with and infrared signals dominate, decision-making will shift towards space environment-related factors. In some areas, priority should be given to addressing environmental compatibility issues, such as screen brightness and color deviation, to avoid making incorrect decisions based on low-quality signals.

[0043] in, Take 0.5-0.8 and emphasize The intensity of user behavioral intent is based on the characteristics of immersive spaces, which are centered on individual user-initiated interaction. The duration of user gestures and the duration of eye focus are direct indicators of whether a user has a need for interaction. For example, if a user stares at a digital avatar for more than 3 seconds, It will rise to the 40-50 range, at which point... A high value ensures that the request is responded to first; while The introduction of the intensity of user group interaction needs solves the problem of ignoring group needs in multi-person scenarios, such as when multiple people wave at the same time, ensuring the consistency of group actions. Approaching 1, If the focus is too high, the decision-making process will adjust the storyline or lighting to be more suitable for group interaction, avoiding a decline in the group experience caused by focusing only on individuals.

[0044] Take 0.6-0.9 and emphasize Ambient light adaptation requirements address the technical issue that has the greatest impact on screen display performance; when the difference between ambient light and screen display brightness... When it exceeds 500 cd / m², It will rise to the 40-50 range, at which point... A high value ensures that screen brightness is adjusted first, preventing a break in the immersive experience caused by users not being able to see the content on the screen; while The design, which addresses the need for screen display deviation correction, solves the common problem of uneven color when splicing multi-fold screens, such as color deviation at the splicing point of a four-fold screen. When it exceeds 30ΔE, If the setting is raised, the decision-making process will initiate color calibration to ensure consistency in scene display. The quantitative design, with values ​​ranging from 0 to 100, provides clear priority standards for decision-making execution, such as... The system automatically implements digital human-based proactive interaction and emergency adjustments to lighting and shadows. The plot progresses gradually while regular lighting and shadow adjustments are made. The model accuracy is optimized in real time, and those skilled in the art can directly design execution logic based on this value range to ensure that the decision is implemented.

[0045] This formula, through the linkage of various parameters, quantifies and integrates factors from three dimensions: signal quality, user needs, and environmental conditions. It fully supports the technical characteristics of generating AI interactive decision-making strategies based on the fused signal, combined with user location, user behavioral intent, and spatial environmental parameters. The formula is based on By linking signal quality with decision priority, the system avoids prioritizing decisions based on low-quality signals. It clarifies the weighting coefficients and value ranges of each parameter to achieve quantitative priority calculation. A value range of 0-100 clearly defines the urgency of decisions and standardizes the execution order. The resulting decision priorities are more accurate, dynamically sorting based on signal quality, user needs, and environmental conditions. This ensures that high-demand, highly adaptable decisions are executed first, improving the relevance of interactions and the adaptability to different scenarios.

[0046] Traditional technical solutions have the following technical problems: First, they do not take into account the priority of AI interaction decisions, and cannot adjust the model accuracy according to the urgency of the interaction, resulting in the use of low-precision models even when there is a high priority interaction; second, the calculation of the rendering load coefficient has no clear logic, and does not relate it to the number of model vertices and the frame rate, so it cannot accurately reflect the load status; third, the maximum limit of the vertex simplification rate is not clear, which may lead to excessively low accuracy and damage to scene details.

[0047] Based on this, the formula for calculating the vertex simplification rate of the collaboratively adjusted 3D model's vertex count and rendering frame rate is as follows: ; In the formula The simplification factor of the 3D model at time t is dimensionless, ranging from 0 to 1. A lower value indicates higher model accuracy. The number of vertices = ); The maximum vertex simplification rate is dimensionless and takes a value of 0.8 to avoid excessively low precision; k is the cooperative control coefficient, dimensionless, and takes a value of 0.02-0.05 for adjustment. and The degree of influence on the simplification rate; The interactive decision priority in claim 6 is dimensionless, related to interactive needs, and reduces the simplification rate to improve accuracy when the priority is high. Render the load coefficient for the model at time t. It is dimensionless and takes a value of 0-1. The higher the load, the closer the coefficient is to 1. Let t be the number of vertices in the current model at time t, expressed in units of "vertices", reflecting the current accuracy status of the model; The maximum number of vertices in the model is expressed in units of , with a value of 2 × 10^6, and the upper limit of precision is defined. Let t be the current rendering frame rate, measured in fps, reflecting the smoothness of the current scene; To ensure the minimum guaranteed frame rate, the unit is fps, with a value of 30fps, clearly defining the lower limit of smoothness.

[0048] in the formula The innovation of this load factor calculation formula lies in integrating the two core indicators of model accuracy and rendering frame rate into a single load factor: Reflects the impact of model accuracy on load; when the current number of model vertices Approaching the maximum number of vertices When the ratio is close to 1, it indicates that the excessively high precision leads to an increase in load. Reflects the frame rate's feedback to the load; when the current frame rate... Below the minimum guaranteed frame rate When the frame rate is less than 1, it indicates that the load is too high, causing a drop in frame rate; the product of the two is... A value between 0 and 1 can comprehensively reflect the current load status of the rendering system, such as... This indicates that the load is too high. This indicates that the load is too low. This quantitative design solves the technical problem of ambiguity in load judgment and provides a clear basis for subsequent vertex simplification rate adjustment.

[0049] formula This is the key to achieving accurate coordination between load and decision-making: The setting is based on a balance between UE5 rendering performance and scene accuracy; when At that time, the number of vertices in the model is This number of vertices ensures that the scene details displayed on the four-fold / five-fold screen, such as the digital human facial texture and the outline of scene props, are preserved, while avoiding the problem of blurry images due to too few vertices and avoiding oversimplification. A coefficient of 0.02-0.05 is used as the collaborative control coefficient, which is used to adjust... and right The degree of influence. When High, such as 90-100, high-priority interaction and Low, such as 0.2-0.3, when the load is low. The value is approximately 0.18-0.45. After substituting it into the formula... The approximate value is 0.8×(1-0.18 / 1.18)≈0.68-0.8×(1-0.45 / 1.45)≈0.52. At this point, the number of model vertices is approximately 6.4×10⁵-9.6×10⁵. This improves accuracy while maintaining a frame rate ≥45fps, meeting the demands of high-priority interactions for scene detail. Low, such as 30-40, low-priority interactions and High, such as 0.7-0.8, under high load, The value is approximately 0.042-0.16. After substituting it into the formula... The approximate values ​​are 0.8 × (1 - 0.042 / 1.042) ≈ 0.77 - 0.8 × (1 - 0.16 / 1.16) ≈ 0.71. At this point, the number of vertices in the model is approximately 4.8 × 10⁵ - 5.8 × 10⁵. By appropriately simplifying the vertices to reduce the load, a frame rate of ≥ 60fps is ensured, avoiding stuttering caused by excessive load during low-priority interactions. The linkage logic of this formula completely solves the technical problem of the contradiction between accuracy and frame rate, and all parameters have clear technical basis and value ranges, which can be directly applied to actual system development by those skilled in the art.

[0050] At the same time, it passed By linking AI decision priorities, the model adjustment is synchronized with user interaction needs, fully supporting the technical feature of collaboratively adjusting the number of vertices and rendering frame rate of the 3D model in claim 1.

[0051] The formula is based on Associating interaction priority with vertex simplification rate, higher priority interactions improve model accuracy; through Combining vertex count and frame rate to quantize load ensures accurate load calculation; Limiting the simplification rate avoids excessively low accuracy. The final calculated vertex simplification rate dynamically adapts to interaction requirements and load conditions, achieving a coordinated balance between 3D model accuracy and rendering frame rate. This ensures sufficient scene accuracy and stable frame rate during high-priority interactions, while maintaining high smoothness during low-priority interactions, thus enhancing the overall immersive experience.

[0052] Traditional technical solutions suffer from the following problems: the core module functions are vaguely defined, and the specific role and processing objects of each module are not clearly defined, resulting in the modules being unable to work together; the equipment specifications and acquisition content of the multimodal acquisition module are not clearly defined, making it impossible to ensure the accuracy and coverage of data acquisition; the signal interaction methods between modules are not clearly defined, and only vague descriptions such as connection are used, making it impossible to implement the system architecture.

[0053] Based on this, please refer to Figure 2This embodiment provides an immersive 3D interactive spatial system based on artificial intelligence, including: a 3D model construction module, a multimodal acquisition module, a signal fusion processing module, an AI decision-making module, and a scene execution module. The 3D model construction module uses UE5 ground editing technology to construct a 3D model of a four-fold or five-fold immersive space and achieve 3D visualization. The multimodal acquisition module includes at least eight 4K industrial cameras and at least 32 infrared touch sensors. The 4K industrial cameras are used to acquire visual data of user contours and gaze direction, and the infrared touch sensors are used to acquire touch signals of user touch pressure and position. The signal fusion processing module includes an FPGA processor, which is used to perform multimodal signal fusion processing on the visual data and the touch signals. The AI ​​decision-making module includes an edge computing server, which is used to generate AI interactive decision-making strategies based on the fused signals. The scene execution module includes a UE5 rendering engine, a screen display controller, and a digital human rendering unit, which is used to adjust the interaction behavior between the 3D scene and the digital human according to the AI ​​interactive decision-making strategy. Each module achieves signal interaction through a PCIe 4.0 bus.

[0054] This solution clearly defines the specific functions and processing objects of each module. For example, the 3D model building module is based on UE5 ground programming, and the multimodal acquisition module specifies the equipment specifications and acquisition content. The PCIe 4.0 bus defines the signal interaction method between modules to ensure efficient and stable data transmission. The final system architecture is clear and implementable, with modules working collaboratively to achieve a complete process from data acquisition, fusion, decision-making to execution. This solves the problem of disorganized and unimplementable modules in existing systems, providing stable hardware support for immersive interaction.

[0055] Traditional technical solutions have the following technical problems: the FPGA processor has no specific model, making it impossible to determine whether its processing performance meets the real-time processing requirements of multimodal signals; the storage content of the storage unit is not clear, resulting in the inability to store algorithm programs and data reasonably, and low retrieval efficiency; the connection method between the FPGA and the multimodal acquisition module is not clear, and the stability and real-time performance of signal transmission cannot be guaranteed.

[0056] Based on this, the FPGA processor of the signal fusion processing module is a Xilinx Kintex UltraScale. The FPGA processor is also connected to a storage unit, which is used to store the modal decomposition algorithm program, interference feature library data, and multimodal signal fusion weight calculation program. The FPGA processor is connected to the multimodal acquisition module via the SPI bus, receives the visual data and touch signals output by the multimodal acquisition module, and calls the program in the storage unit to preprocess and fuse the visual data and the touch signals.

[0057] The solution specifies the FPGA model as Xilinx Kintex UltraScale to ensure it has sufficient parallel processing capabilities to meet the real-time processing requirements of multimodal signals; it also specifies the specific storage content of the storage unit to achieve orderly storage of algorithms and data and improve retrieval efficiency; and it defines the connection method between the FPGA and the acquisition module using the SPI bus to ensure stable and real-time signal transmission.

[0058] The final signal fusion processing module can efficiently and stably complete data preprocessing and fusion, avoiding processing delays or data loss caused by insufficient hardware performance, storage chaos, or unstable connection, and providing high-quality fused data for subsequent AI decision-making.

[0059] Traditional technical solutions have the following technical problems: the screen display controller has no specific protocol, which makes it impossible to accurately control the light and shadow parameters, resulting in poor ambient light adaptation; the hardware model of the digital human rendering unit is not clear, the rendering performance is insufficient, and it cannot support complex digital human interactive behaviors; the definition of digital human interactive behaviors is vague and the specific interaction forms are not clear, resulting in simple interaction and lack of immersion; and the connection method between the rendering unit and the AI ​​decision module is not clear, resulting in high latency in the transmission of decision commands.

[0060] Based on this, the screen display controller of the scene execution module adopts a DMX512 protocol controller, which is used to adjust the brightness and color temperature of the screen. The digital human rendering unit includes a GPU rendering card, specifically an NVIDIA RTX A6000, which drives the digital human to initiate interactive behaviors such as gestures and voice. The GPU rendering card is connected to the edge computing server of the AI ​​decision module via a PCIe 4.0 bus, receiving interactive decision commands output by the edge computing server. This solution uses a DMX512 protocol controller to ensure accurate adjustment of lighting parameters, such as brightness and color temperature, to adapt to changes in ambient light; it specifies the GPU model as NVIDIA RTX A6000 to ensure powerful graphics rendering capabilities to support complex digital human interactions; it defines specific interactive behaviors of the digital human, such as gestures and voice, to enrich the forms of interaction; and it connects the GPU and the edge computing server via a PCIe 4.0 bus to reduce the latency of decision command transmission.

[0061] The final scene execution module can accurately adjust the lighting and shadows and smoothly drive digital human interaction, avoiding problems such as poor scene adaptation and insufficient immersion caused by unclear control protocols, insufficient hardware performance, simple interaction, and instruction delay, thus improving the user's interactive experience with the scene and digital human.

[0062] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. An immersive spatial three-dimensional interactive method based on artificial intelligence, characterized in that, include: The 3D model of a four-fold or five-fold immersive space is constructed and visualized using UE5 ground editing technology. At least 8 4K industrial cameras are used to collect user visual data and at least 32 infrared touch sensors are used to collect user touch signals. The visual data and the touch signal are fused together using a fusion weight calculation to achieve multimodal signal fusion processing, thereby eliminating the coupling interference between visual data loss caused by user occlusion and false triggering of infrared touch. Based on the fused signal, AI interactive decision-making strategies are generated by combining user location, user behavior intent and spatial environment parameters. The AI ​​interactive decision-making strategy dynamically adjusts the lighting parameters, plot progress, and digital human interaction behavior of the 3D scene, while simultaneously adjusting the vertex count and rendering frame rate of the 3D model to solve the problem of disconnect between users and scenes and digital human interaction in immersive spaces, thereby enhancing the intelligence of the immersive experience.

2. The immersive spatial three-dimensional interactive method based on artificial intelligence according to claim 1, characterized in that, The multimodal signal fusion processing also includes a preprocessing step. The visual data is processed using a modal decomposition algorithm from the field of mechanical vibration monitoring to extract effective visual features of the user's contour and gaze direction to achieve noise reduction. The touch signal is processed using an interference feature library matching algorithm to filter effective touch signals of the user's touch pressure and position to achieve filtering. The interference feature library is constructed by pre-collecting waveform features of common interference signals in the immersive space. The modal decomposition algorithm and the interference feature library matching algorithm are stored in the storage unit of the FPGA processor and are called and executed by the FPGA processor.

3. The immersive spatial three-dimensional interactive method based on artificial intelligence according to claim 1, characterized in that, In the step of generating an AI interactive decision-making strategy by combining user location, user behavioral intent, and spatial environment parameters, user behavioral intent is calculated by weighting the duration of user gesture dwell time and the duration of gaze focus, with weighting coefficients of 0.6 for gesture dwell time and 0.4 for gaze focus time; user group interaction needs are calculated by weighting the number of users and the consistency of group actions, with weighting coefficients of 0.3 for the number of users and 0.7 for the consistency of group actions; spatial environment parameters include ambient light intensity and color deviation at the screen splicing point, where ambient light intensity is used to calculate the difference between ambient light and screen display brightness, and color deviation is used to calculate the color uniformity at the screen splicing point.

4. The immersive spatial three-dimensional interactive method based on artificial intelligence according to claim 1, characterized in that, In the step of collaboratively adjusting the vertex count and rendering frame rate of the 3D model, the adjusted vertex count and rendering frame rate data of the 3D model are collected in real time and fed back to the multimodal signal fusion processing step to form a closed-loop processing flow with a closed-loop processing cycle of no more than 50ms; the rendering frame rate adjustment is based on a minimum of 30fps, combined with dynamic adaptation of the 3D model vertex count, and the 3D model vertex count adjustment is no more than 2×10 6 The upper limit is set to ensure a coordinated balance between the accuracy of the 3D model and the smoothness of rendering.

5. The immersive spatial three-dimensional interactive method based on artificial intelligence according to claim 1, characterized in that, The formula for calculating the fusion weight in the multimodal signal fusion processing is as follows: ; in: The multimodal signal fusion weights at time t are... This is the basic weighting coefficient for the signal, and its value ranges from 0.4 to 0.

6. The effective visual data intensity at time t is expressed in cd / m². Let be the ambient light interference coefficient at time t, with a value between 0.1 and 1.0, and n be the total number of cameras, where n ≥ 8. Let be the raw visual data intensity of the i-th camera at time t, in cd / m². The effective infrared touch signal strength at time t is expressed in V. Let t be the user touch distance coefficient, with a value between 0.2 and 1.0, and m be the total number of infrared touch sensors, where m ≥ 32. Let be the original touch signal strength of the j-th infrared sensor at time t, expressed in V. To avoid the minimum value of denominator 0 and take the value 10 -6 .

6. The immersive spatial three-dimensional interactive method based on artificial intelligence according to claim 5, characterized in that, The formula for calculating the decision priority of the AI ​​interactive decision-making strategy is as follows: ; in: Let t be the priority of the interactive decision and take a value between 0 and 100. For multimodal signal fusion weights, This is the behavioral intent weighting coefficient, with a value ranging from 0.5 to 0.

8. Let t be the intensity of the user's behavioral intent, with a value between 0 and 50. Let t be the intensity of user group interaction demand, with a value ranging from 0 to 50. The ambient light influence coefficient ranges from 0.6 to 0.

9. Let be the ambient light intensity required for adaptation at time t, and let its value be 0-50. The required intensity of screen display deviation correction at time t is 0-50.

7. The immersive spatial three-dimensional interactive method based on artificial intelligence according to claim 6, characterized in that, The formula for calculating the vertex simplification rate of the collaboratively adjusted 3D model in terms of vertex count and rendering frame rate is as follows: ; In the formula Let be the vertex simplification factor of the 3D model at time t, and take a value between 0 and 1. is the maximum vertex simplification rate, with a value of 0.8, and k is the collaborative control coefficient, with a value between 0.02 and 0.

05. Prioritizing interactive decision-making Render the load coefficient for the model at time t, with a value between 0 and 1. Let t be the number of vertices in the current model, expressed in units of vertices. The maximum number of vertices in the model, with a value of 2 × 10. 6 indivual, Let t be the current rendering frame rate in fps. The minimum guaranteed frame rate is set to 30fps.

8. An immersive spatial three-dimensional interactive system based on artificial intelligence, applied to the immersive spatial three-dimensional interactive method based on artificial intelligence as described in any one of claims 1-7, characterized in that, It includes: a 3D model building module, a multimodal acquisition module, a signal fusion and processing module, an AI decision-making module, and a scene execution module; The 3D model building module uses UE5 ground editing technology to build a 3D model of a four-fold or five-fold immersive space and realize 3D visualization; The multimodal acquisition module includes at least 8 4K industrial cameras and at least 32 infrared touch sensors. The 4K industrial cameras are used to acquire visual data of the user's contour and gaze direction, and the infrared touch sensors are used to acquire touch signals of the user's touch pressure and position. The signal fusion processing module includes an FPGA processor, which is used to perform multimodal signal fusion processing on the visual data and the touch signal; The AI ​​decision-making module includes an edge computing server, which is used to generate AI interactive decision-making strategies based on the fused signals. The scene execution module includes a UE5 rendering engine, a screen display controller, and a digital human rendering unit. The scene execution module is used to adjust the interaction behavior between the 3D scene and the digital human according to the AI ​​interaction decision strategy. Each module realizes signal interaction through the PCIe 4.0 bus.

9. The immersive spatial three-dimensional interactive system based on artificial intelligence according to claim 8, characterized in that, The FPGA processor of the signal fusion processing module is a Xilinx Kintex UltraScale. The FPGA processor is also connected to a storage unit, which is used to store the mode decomposition algorithm program, interference feature library data and multi-mode signal fusion weight calculation program. The FPGA processor is connected to the multimodal acquisition module via the SPI bus, receives visual data and touch signals output by the multimodal acquisition module, and calls the program in the storage unit to preprocess and fuse the visual data and touch signals.

10. The immersive spatial three-dimensional interactive system based on artificial intelligence according to claim 8, characterized in that, The screen display controller of the scene execution module adopts a DMX512 protocol controller, which is used to adjust the screen brightness, color temperature and other light and shadow parameters. The digital human rendering unit includes a GPU rendering card, which is an NVIDIA RTX A6000. The GPU rendering card is used to drive the digital human to initiate interactive behaviors such as gestures and voice. The GPU rendering card is connected to the edge computing server of the AI ​​decision-making module via a PCIe 4.0 bus, and receives interactive decision-making instructions output by the edge computing server.