Naked eye 3D visual display method and system based on AI intelligent image adaptation
By using AI-powered intelligent image adaptation technology, the problems of distortion and disconnect between multiple user perspectives in naked-eye 3D displays have been solved, enabling distortion-free naked-eye 3D viewing and a naturally integrated immersive experience for multiple users, thus enhancing the interactive effect in multi-user scenarios.
Patent Information
- Application Number
- CN202511720162.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-02-13
AI Technical Summary
Existing naked-eye 3D display technologies suffer from problems such as perspective distortion, interaction-visual disconnect, and low integration of digital humans with the scene in multi-user scenarios, failing to meet the needs of multiple people watching and interacting simultaneously.
By adopting an AI-powered intelligent image adaptation method, the system collects naked-eye 3D screen parameters and scene data, uses an AI multi-view image adaptation engine to decompose and dynamically reconstruct the viewpoint, and combines camera tracking of the audience's position and interaction commands to achieve real-time viewpoint adaptation and interactive-visual linkage for multiple users. Furthermore, the naked-eye 3D digital human collaboration module enables the sharing of spatial coordinates and visual parameters between the digital human and the scene.
It achieves a distortion-free naked-eye 3D viewing experience for multiple users, enhances the sense of interactive immersion and the natural integration of digital humans with the scene, and improves the immersion and interactive effect in multi-user scenarios.
Smart Images

Figure CN121531111A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of naked eye 3D visual display, and particularly relates to a naked eye 3D visual display method and system based on AI intelligent image adaptation. BACKGROUND
[0002] The existing naked eye 3D display technology generally has a multi-user experience bottleneck. The traditional scheme adopts a single-view optimization method, and the generated view image can only adapt to the fixed position of the audience. When multiple audiences are at different positions, the audiences far from the fixed position will see distorted and blurred 3D images, and cannot obtain clear viewing effect. At the same time, the interactivity and visual display are insufficiently linked, and the interactive instructions such as touch and gesture triggered by the audience cannot be adjusted in real time, which leads to disconnection between interactive operation and visual effect, and poor immersion. In addition, the fusion degree of naked eye 3D digital person and scene is low, and the spatial coordinates and visual parameters of the digital person model and the scene are independent of each other. When the digital person is displayed, it is easy to appear suspended, and lacks real spatial correlation and occlusion relationship with the 3D model in the scene, which destroys the overall visual coherence. These problems lead to poor applicability of the existing technology in a multi-user scene, and cannot meet the actual needs of outdoor large-screen multi-person viewing, exhibition hall multi-person interaction, educational scene multi-student operation, etc.
[0003] Based on the above problems, there is an urgent need for a naked eye 3D visual display scheme that can realize multi-user real-time view adaptation, interactive-visual deep linkage, and natural fusion of digital person and scene. SUMMARY
[0004] The purpose of the present application is to solve the shortcomings in the prior art, and to propose a naked eye 3D visual display method based on AI intelligent image adaptation, comprising: S1: collecting naked eye 3D screen parameters and scene data, wherein the naked eye 3D screen parameters include screen resolution, display area size, and pixel arrangement mode, and the scene data includes viewing area range, initial audience position, environmental lighting information, and interactive device deployment position; S2: inputting the naked eye 3D screen parameters and scene data into an AI multi-view image adaptation engine, wherein the AI multi-view image adaptation engine performs view decomposition and dynamic reconstruction on 3D models or videos based on deep learning, and automatically generates a multi-view image sequence covering the viewing area; S3: dynamically tracking the audience position by a camera, and transmitting the audience position coordinates to the AI multi-view image adaptation engine in real time, wherein the AI multi-view image adaptation engine updates the view image sequence in real time according to the audience position coordinates; S4: synchronously triggering 3D visual parameter adjustment by an interactive instruction, wherein the AI multi-view image adaptation engine reversely optimizes the interactive area rendering parameters; S5: The naked eye 3D digital human cooperation module makes the digital human model share the space coordinates and visual parameters with the scene, and the digital human model enters the interaction area with a naked eye 3D posture during interaction, and the AI multi-view image adaptation engine adjusts the parallax and occlusion relationship of the digital human model in real time; S6: Form a closed-loop process of scene data, AI adaptation, naked eye 3D display, and signal feedback to AI dynamic adjustment, and the signal feedback includes audience position feedback signal, interaction instruction feedback signal, and display effect feedback signal.
[0005] Preferably, the environmental light information in S1 includes light intensity, light direction, and light spectrum distribution, the initial audience position is initially collected by a camera array, and the interaction device deployment position includes a touch device coordinate, a gesture recognition device detection range, and a line-of-sight tracking device effective monitoring angle.
[0006] Further preferably, the deep learning model of the AI multi-view image adaptation engine in S2 adopts an encoder-decoder architecture, the encoder extracts depth features of a 3D model or a video, the decoder dynamically reconstructs the extracted depth features according to the naked eye 3D screen parameters and the scene data, and the generated multi-view image sequence covers all possible audience observation angles in the viewing area, and the resolution of each view image matches the naked eye 3D screen parameters.
[0007] Further preferably, the camera array in S3 adopts a distributed deployment manner, there is an overlapping area between the detection ranges of adjacent cameras, the camera array collects audience position coordinates in real time at a frequency consistent with the frequency at which the AI multi-view image adaptation engine updates the view image sequence, and the audience position coordinates include audience eye three-dimensional coordinates and head rotation angle.
[0008] Further preferably, when the AI multi-view image adaptation engine updates the view image sequence according to the audience position coordinates in S3, a multi-dimensional view optimization weight calculation formula is adopted, and the multi-dimensional view optimization weight calculation formula is as follows: ; Wherein, is the optimization weight of the jth view image corresponding to the ith audience, is the position moving angular velocity of the ith audience, with a dimension of rad / s (radian / second), is the environmental light intensity corresponding to the jth view image, with a dimension of lux (lx), is the head rotation angle of the ith audience, with a dimension of rad (radian), is the optimal observation angle, with a dimension of rad (radian), , , , are constants greater than 0, dimensionless, used to adjust the contribution of each influencing factor to the optimization weight.
[0009] Further preferably, when the interactive instruction synchronously triggers the adjustment of the 3D visual parameters in S4, an interactive-visual linkage response parameter calculation formula is adopted, and the interactive-visual linkage response parameter calculation formula is as follows: ; wherein, is the 3D visual parameter adjustment response value corresponding to the kth interactive instruction, dimensionless, is the priority of the kth interactive instruction, and the value range is 1-10, dimensionless. The larger the value is, the higher the priority is, is the trigger duration of the kth interactive instruction, and the unit is second (s), is the time variation rate of the multi-dimensional perspective optimization weight, and the unit is 1 / second (1 / s), , , are proportional coefficients, and are constants greater than 0, dimensionless, used to adjust the influence of the interactive instruction related factors and the weight variation rate on the response value.
[0010] Further preferably, when the AI multi-perspective image adaptation engine reversely optimizes the interactive region rendering parameters in S4, the interactive-visual linkage response parameter adjusts the pixel brightness, contrast, color saturation of the interactive region, and the texture resolution and edge sharpening degree of the 3D model. The adjustment amplitude is positively correlated with the value of .
[0011] Further preferably, the spatial coordinates shared by the naked-eye 3D digital human collaboration module in S5 include three-dimensional coordinates in the world coordinate system, and the unit is meter. The visual parameters include the field of view angle, the depth of field range, and the perspective projection matrix. The unit of the field of view angle is radian, and the unit of the depth of field range is meter. When the AI multi-perspective image adaptation engine adjusts the parallax of the digital human model, the parallax reference value of the digital human model is determined based on the depth information of the 3D model in the scene, and then dynamically offset adjustment is performed according to the audience position coordinates.
[0012] An AI intelligent image adaptation naked-eye 3D visual display system based on the AI intelligent image adaptation naked-eye 3D visual display method, comprising a scene data acquisition module, a naked-eye 3D screen parameter acquisition module, an AI multi-view image adaptation engine module, a camera array module, an interactive device module, a naked-eye 3D digital human collaboration module, a naked-eye 3D display module, and a feedback transmission module. The scene data acquisition module is used to acquire the viewing area range, the initial audience position, the environmental lighting information, and the interactive device deployment position. The naked-eye 3D screen parameter acquisition module is used to acquire the screen resolution, the display area size, and the pixel arrangement mode. The AI multi-view image adaptation engine module is used to receive the data transmitted by the scene data acquisition module and the naked-eye 3D screen parameter acquisition module, perform perspective decomposition and dynamic reconstruction on the 3D model or video based on deep learning, generate a multi-view image sequence, update the perspective image sequence according to the audience position coordinates transmitted by the camera array module, reversely optimize the interactive area rendering parameters, and adjust the parallax and occlusion relationship of the digital human model. The camera array module is used to dynamically track the audience position, acquire the audience position coordinates, and transmit them to the AI multi-view image adaptation engine module. The interactive device module comprises a touch device, a gesture recognition device, and a line-of-sight tracking device, and is used to generate interactive instructions and transmit them to the AI multi-view image adaptation engine module. The naked-eye 3D digital human collaboration module is used to realize the sharing of the spatial coordinates and visual parameters of the digital human model and the scene, control the digital human model to enter the interactive area in a naked-eye 3D posture, and display the perspective image sequence transmitted by the AI multi-view image adaptation engine module. The feedback transmission module is used to acquire the audience position feedback signal, the interactive instruction feedback signal, and the display effect feedback signal, and transmit them to the AI multi-view image adaptation engine module to form a closed-loop transmission link.
[0013] Further preferably, when the AI multi-view image adaptation engine module adjusts the parallax of the digital human model, a digital human-scene collaborative parallax adjustment calculation formula is adopted, which is as follows: ; wherein, is the parallax adjustment value of the mth digital human model in the nth scene area, with the dimension of meters (m), is the three-dimensional depth value of the mth digital human model, with the dimension of meters (m), is the background depth value of the nth scene area, with the dimension of meters (m), 、 、 is a correction coefficient, and all are constants greater than 0, dimensionless, used to correct the influence of the digital human depth, the scene background depth, and the perspective optimization weight on the parallax adjustment value, For the interactive-visual linkage response parameter, dimensionless, For the multi-dimensional perspective optimization weight, dimensionless.
[0014] The technical effects achieved by the above embodiments include: The creative technical point of the present application is to construct an integrated scheme of AI multi-perspective adaptation, interactive closed loop, and digital human collaboration, to realize multi-user real-time perspective adaptation, interactive instruction and 3D visual parameter synchronous adjustment, and digital human and scene parameter sharing through multi-step collaboration. The problems of multi-user perspective distortion, interactive-visual disconnection, and digital human fusion difference in the background art are solved, and multi-user distortion-free experience is realized, interactive immersion is strengthened, and digital human and scene are naturally fused. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 For the present application, an AI intelligent image adaptation naked eye 3D visual display method flow chart is provided; Figure 2 For the present application, an AI intelligent image adaptation naked eye 3D visual display system connection block diagram is provided. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.
[0017] In the existing naked eye 3D display technology, it is difficult to obtain clear and distortion-free effect when multiple users watch at the same time, traditional single-perspective optimization cannot cover different audience positions, interactive instructions and 3D visual parameter adjustment are not synchronized, digital human model and scene fusion degree are low, and there are inherent defects in perspective, interaction and fusion.
[0018] Based on this, please refer to Figure 1 and Figure 2 The present embodiment provides an AI intelligent image adaptation naked eye 3D visual display method, which comprises the following steps: S1: collecting naked eye 3D screen parameters and scene data, wherein the naked eye 3D screen parameters include screen resolution, display area size, pixel arrangement mode, and the scene data includes viewing area range, initial audience position, environmental lighting information, and interactive device deployment position; S2: inputting the naked eye 3D screen parameters and scene data into an AI multi-perspective image adaptation engine, wherein the AI multi-perspective image adaptation engine performs perspective decomposition and dynamic reconstruction on 3D models or videos based on deep learning, and automatically generates a multi-perspective image sequence covering the viewing area; S3: Camera dynamically tracks audience position, real-time acquires audience position coordinates and transmits to AI multi-view image adaptation engine, which updates view image sequence according to audience position coordinates in real time; S4: Interactive instruction synchronously triggers 3D visual parameter adjustment, the AI multi-view image adaptation engine reversely optimizes interactive area rendering parameters; S5: Naked-eye 3D digital human cooperation module makes digital human model share space coordinates and visual parameters with scene, and when interacting, the digital human model enters the interactive area with naked-eye 3D posture, and the AI multi-view image adaptation engine adjusts the parallax and occlusion relationship of the digital human model in real time; S6: Form a closed-loop process of scene data, AI adaptation, naked-eye 3D display, signal feedback to AI dynamic adjustment, the signal feedback includes audience position feedback signal, interactive instruction feedback signal, display effect feedback signal.
[0019] It is worth mentioning that: S1 is the basic step of data collection, and the naked eye 3D screen parameters include screen resolution, which refers to the number of pixels in the horizontal and vertical directions of the screen, such as 1920x1080 pixels, which determines the level of detail of image display; display area size refers to the physical length and width of the screen that can display images, such as 10m x 5m for outdoor large screens, which is used to calculate the viewing angle coverage of the viewing area; pixel arrangement refers to the distribution structure of pixels on the screen, such as RGB stripe arrangement, which affects the parallax generation effect of naked eye 3D images. In the scene data, the viewing area range is the physical space range where the audience may exist, such as a circular area with a radius of 5m centered on the screen in the exhibition hall, which is used to determine the coverage of the multi-view image sequence; the initial audience position is the three-dimensional coordinates of the audience at the beginning of the demonstration, which is obtained through the initial collection of the camera array and serves as the initial reference for updating the viewing angle; environmental lighting information includes lighting intensity, direction, etc., which affects the brightness and contrast adjustment of the image; the deployment position of the interactive device is the installation coordinates of touch devices, gesture recognition devices, etc., which are used to determine the position of the interactive area and ensure that the interactive instructions match the visual parameter adjustment of the corresponding area. S2 is the core step of AI multi-view image generation, and the deep learning model used by the AI multi-view image adaptation engine is an encoder-decoder architecture. The encoder extracts the depth features of the 3D model or video through a convolutional neural network, including information such as the spatial position, contour, and texture of the object, such as the surface concave-convex features and color distribution of the 3D product model; the decoder reconstructs the depth features into images of different viewing angles based on the deconvolution operation, combined with the input naked eye 3D screen parameters and scene data, such as generating an image for every 1° viewing angle interval within the viewing area, ensuring that the generated multi-view image sequence can cover all possible audience positions, and the resolution of each viewing angle image is consistent with the screen resolution, avoiding distortion caused by pixel stretching. S3 is the dynamic updating step of the viewing angle, and the camera array is distributed, such as arranging 8 cameras evenly around the screen, with a 10% overlap in the detection range of adjacent cameras to ensure no blind area. The camera obtains the audience position coordinates in real time through computer vision algorithms, including the eye three-dimensional coordinates (x, y, z) and the head rotation angle. The coordinate data is sent to the AI multi-view image adaptation engine in real time through Ethernet, and the engine calculates the viewing angle offset based on the coordinate changes, such as when the audience moves from the front of the screen to the left side 30° position, the engine calls the corresponding left side 30° viewing angle image from the multi-view image sequence, or generates a transition viewing angle image through an interpolation algorithm, ensuring smooth image updating when the audience moves without stuttering or distortion.S4 is the interactive-visual linkage step. Interactive commands are generated by the interactive device. For example, when a viewer touches the product model area on the touch screen, a touch command is generated. The gesture recognition device captures the viewer's rotation gesture and generates a rotation command. These commands are transmitted to the AI multi-view image adaptation engine via the data bus. The engine synchronously triggers the adjustment of 3D visual parameters, such as increasing the pixel brightness and contrast of the touch area. The rotation command triggers the 3D model to rotate in the direction of the gesture. At the same time, the engine reversely optimizes the rendering parameters of the interactive area, such as increasing the texture resolution of the interactive area from 512×512 to 1024×1024, so that the viewer can see the interactive details more clearly when operating. S5 is the digital human collaborative integration step. The naked-eye 3D digital human collaborative module establishes a parameter sharing mechanism between the digital human model and the scene. The shared spatial coordinates are three-dimensional coordinates in the world coordinate system. For example, if the coordinates of the 3D cultural relic model in the scene are (2,3,4)m, the coordinates of the digital human model are set to (2.5,3,4)m to ensure that the digital human is positioned reasonably next to the cultural relic model. Visual parameters include field of view, depth of field, and perspective projection matrix. These parameters are consistent with the scene parameters. For example, if the scene's field of view is 60°, the field of view of the digital human is also set to 60° to avoid visual disproportion. Disharmony; During interaction, the digital human enters the interactive area in a naked-eye 3D posture. For example, when the audience triggers the instruction to explain the cultural relic, the digital human moves from the edge of the scene to the side of the cultural relic. The AI multi-view image adaptation engine adjusts the parallax of the digital human according to the audience's position coordinates. For example, when the audience is on the right side of the screen, the parallax of the digital human on the right side decreases and the parallax on the left side increases, so that the digital human appears to be in the same space as the cultural relic. At the same time, it handles occlusion. When the digital human's arm occludes part of the cultural relic, the engine calculates the occlusion area in real time, hides the pixels of the occluded cultural relic, and displays the pixels of the digital human's arm, which conforms to the occlusion logic of real physical space. S6 is a closed-loop optimization process. Scene data is continuously collected and updated through the scene data acquisition module, such as updating ambient lighting information and audience position data every 0.5 seconds. AI adaptation adjusts images and parameters based on the updated data. The naked-eye 3D display module displays the adjusted image. The feedback transmission module collects display effect feedback signals, such as the actual brightness detected by the screen brightness sensor, audience position feedback signals, and interaction command feedback signals. These signals are sent back to the AI multi-view image adaptation engine. The engine further dynamically adjusts based on the feedback. For example, when the display effect feedback signal shows that the brightness is too low, the engine increases the overall brightness of the image, forming a continuous optimization loop to ensure a stable viewing experience for multiple users.
[0020] The technical effects achieved by the above embodiments include: solving the problems of multi-user perspective distortion, interaction-visual disconnect, and poor digital human integration, realizing distortion-free naked-eye 3D viewing for multiple users, smooth interaction and visual linkage, natural integration of digital human and scene, and enhancing immersion.
[0021] The prior art scene data collection is incomplete, the environmental lighting information only considers the lighting intensity, does not involve the direction and spectral distribution, the initial audience position collection method is single, and the interactive device deployment position does not have specific parameters, resulting in insufficient subsequent AI adaptation accuracy.
[0022] Based on this, the environmental lighting information in S1 of the AI intelligent image adaptation naked eye 3D visual display method includes lighting intensity, lighting direction, and lighting spectral distribution, the initial audience position is collected by a camera array, and the interactive device deployment position includes touch device coordinates, gesture recognition device detection range, and line-of-sight tracking device effective monitoring angle.
[0023] It is worth mentioning that: the collection of environmental lighting information is realized by a lighting sensor array, for example, 4 lighting sensors are arranged around the viewing area, each sensor simultaneously collects lighting intensity and lighting direction, and the lighting spectral distribution is calculated through the orientation and light intensity distribution of the sensor, for example, 45° incident from above the screen, the lighting spectral distribution is detected by a spectrometer module, for example, the proportions of red light, green light, and blue light in the visible light range, and these information is used to adjust the white balance and color gamut of the image by the AI multi-view image adaptation engine, for example, when the lighting direction is incident from the left side, the engine increases the brightness of the right side of the image to avoid the dark area on the left side; when the proportion of red light in the lighting spectral distribution is high, the engine reduces the red channel gain of the image to ensure accurate color restoration. The initial audience position is collected by a camera array for the first time, and the collection time is 10 seconds before the start of the display. The camera array fuses the images collected by each camera through a multi-view stitching algorithm to determine the eye three-dimensional coordinates and head rotation angle of each audience, for example, 5 audiences are identified, and their coordinates (1, 2, 3) m, (-1, 2, 3) m, etc. are recorded as initial data for subsequent view updates. The collection of the interactive device deployment position is realized by device calibration, the touch device coordinates are determined by measuring the coordinates of the upper left corner and the lower right corner of the touch screen in the world coordinate system, such as (0, 0, 0) m and (2, 0, 0) m, to determine the range of the touch area; the detection range of the gesture recognition device is obtained from the device parameter manual and tested in the field, for example, the detection range of the Kinect gesture recognition device is 0.5-5 m, the horizontal view angle is 60°, and the vertical view angle is 45°, to ensure the accuracy of the collected detection range data; the effective monitoring angle of the line-of-sight tracking device is determined through device calibration experiments, for example, the effective monitoring angle of the Tobii line-of-sight tracking device is ± 30° horizontally and ± 20° vertically, and these data are used by the AI multi-view image adaptation engine to determine the effective triggering area of the interactive instruction to avoid false triggering of the interactive instruction.
[0024] The above embodiments achieve the technical effects including: improving the scene data collection dimension, improving the AI adaptation accuracy, ensuring the accuracy of image color, view update, and interactive triggering, and optimizing the multi-user experience.
[0025] The deep learning model structure of the AI multi-view image adaptation engine in the prior art is not clear, the depth feature extraction and reconstruction lack pertinence, the generated multi-view image sequence has insufficient coverage range, and the resolution mismatch leads to poor image quality.
[0026] Therefore, in the AI intelligent image adaptation naked-eye 3D visual display method, the deep learning model of the AI multi-view image adaptation engine in S2 adopts an encoder-decoder architecture. The encoder extracts the depth features of the 3D model or video, and the decoder dynamically reconstructs the extracted depth features according to the naked-eye 3D screen parameters and scene data. The generated multi-view image sequence covers all possible audience observation angles in the viewing area, and the resolution of each view image matches the naked-eye 3D screen parameters.
[0027] It is worth mentioning that: the encoder adopts a 3-layer CNN structure, the input is the point cloud data of the 3D model or the frame image of the video, the first layer of convolution kernel size is 3x3, the number is 64, the step is 1, and the activation function is ReLU, which is used to extract low-level features such as edges and colors; the second layer of convolution kernel size is 3x3, the number is 128, the step is 1, and the activation function is ReLU, which is used to extract middle-level features such as local texture and shape; the third layer of convolution kernel size is 5x5, the number is 256, the step is 2, and the activation function is ReLU, which is used to extract high-level features such as overall contour and spatial position. The decoder adopts a 3-layer deconvolution structure, the first layer of deconvolution kernel size is 5x5, the number is 128, the step is 2, and the activation function is ReLU, which is used to upsample the high-level features; the second layer of deconvolution kernel size is 3x3, the number is 64, the step is 1, and the activation function is ReLU, which is used to generate middle-level reconstruction features; the third layer of deconvolution kernel size is 3x3, the number is 3, the step is 1, and the activation function is Sigmoid, which is used to output the view image. During the reconstruction process, the decoder combines the naked-eye 3D screen parameters, such as a screen resolution of 3840x2160, and sets the output image resolution to 3840x2160. The decoder also combines the viewing area range in the scene data, such as a horizontal viewing angle range of -60° to 60° and a vertical viewing angle range of -30° to 30°, and generates view images at a horizontal interval of 1° and a vertical interval of 1°, ensuring that the multi-view image sequence covers all possible audience observation angles, and the resolution of each view image completely matches the screen resolution, avoiding image distortion caused by stretching or compression.
[0028] The above embodiments achieve the following technical effects: the model structure of the AI multi-view image adaptation engine is clear, the depth feature extraction and reconstruction are highly targeted, the multi-view image coverage is comprehensive, the resolution is matched, and the image quality is improved.
[0029] The unreasonable camera deployment in the prior art leads to tracking blind area, the inconsistency between the frequency of audience position coordinate collection and the frequency of view angle update leads to image lag, and the incomplete position coordinate parameters affect the view angle calculation accuracy. Based on this, in the AI intelligent image adaptive naked eye 3D visual display method, the camera array in S3 adopts a distributed deployment manner, the detection ranges of adjacent cameras have overlapping areas, the frequency of real-time collection of audience position coordinates by the camera array is consistent with the frequency of view angle image sequence update by the AI multi-view image adaptive engine, and the audience position coordinates include audience eye three-dimensional coordinates and head rotation angle.
[0030] It is worth mentioning that: the camera array adopts a ring-shaped distributed deployment, takes the naked eye 3D screen as the center, and uniformly arranges 6 cameras on the circumference 1-2 m in front of the screen, the horizontal field of view angle of each camera is 60°, the vertical field of view angle is 40°, the detection ranges of adjacent cameras overlap by 10° in the horizontal direction and by 5° in the vertical direction, so that no tracking blind area is ensured in the viewing area, for example, all the audiences within the range of 5 m in front of the screen can be captured by at least two cameras at the same time. The frequency of real-time collection of audience position coordinates by the camera array is determined through experiments, for example, it is set to 30 Hz, the frequency of view angle image sequence update by the AI multi-view image adaptive engine is also set to 30 Hz, and the two are time-synchronized through a clock synchronization protocol, so as to avoid image lag caused by inconsistent frequencies, for example, when the audience moves quickly, the coordinate collection and image update are synchronized, so as to ensure that the image can follow the change of the audience view angle in real time. The collection of audience position coordinates is realized through a depth sensor of the camera and an image recognition algorithm, the depth sensor obtains the three-dimensional coordinates (x, y, z) of the audience's eyes, the unit is meter, the x-axis is along the horizontal direction of the screen, the y-axis is along the vertical direction of the screen, and the z-axis is perpendicular to the screen plane; the image recognition algorithm calculates the head rotation angle through face key point detection, including horizontal rotation angle and vertical rotation, the unit is radian, and these complete coordinate parameters are transmitted to the AI multi-view image adaptive engine to accurately calculate the view angle offset and ensure the accuracy of view angle update.
[0031] The above embodiment achieves the technical effects including: eliminating the camera tracking blind area, avoiding image lag, improving the view angle calculation accuracy, and ensuring that the audience can see clear and distortion-free images in real time when moving.
[0032] In the prior art, when the AI multi-view image adaptive engine updates the view angle image, only the audience position is considered, and factors such as moving speed, light intensity, and head rotation angle are not considered, which leads to unreasonable view angle optimization weight and poor viewing effect for some audiences.
[0033] Based on this, in the AI-based intelligent image adaptation naked-eye 3D visual display method, when the AI multi-view image adaptation engine in S3 updates the view image sequence according to the viewer's position coordinates, it adopts a multi-dimensional view optimization weight calculation formula, which is as follows: ;
[0034] in, Let the optimized weights be the image from the j-th viewpoint corresponding to the i-th viewer. Let be the angular velocity of the position movement of the i-th spectator, expressed in radians per second (rad / s). Let be the ambient light intensity corresponding to the j-th viewpoint image, with the dimension lux (lx). Let be the head rotation angle of the i-th audience member, measured in radians (rad). The optimal observation angle is measured in radians (rad). , , , These are adjustment coefficients, all of which are dimensionless constants greater than 0, used to balance the contribution of various influencing factors to the optimization weights.
[0035] It is worth mentioning that the theoretical basis for this formula is the synergistic influence of multiple factors on the priority of perspective optimization, including the angular velocity of the audience's position movement. This reflects the speed of audience movement; the faster the movement, the more urgent the need for a change in perspective. Therefore, it adopts... The reciprocal of, makes The larger the value, the larger the value of that component, and the greater its weight. Higher; ambient light intensity Image sharpness is affected by light intensity; the optimal image sharpness is achieved when the light intensity is between 500-1000 lx. Within this range, the larger the value of this item, the higher its weight; head rotation angle With the optimal observation angle The smaller the deviation, the better the viewing experience for the audience; therefore, an exponential function is used. The smaller the deviation, the larger the value of that component, and the higher its weight. The formula derivation process is as follows: First, the core factors affecting the angle optimization weights are identified as angular velocity, light intensity, and angle deviation. Then, a monotonic relationship between each factor and its weight is established: angular velocity is positively correlated with the weight, light intensity is positively correlated with the weight within a suitable range, and angle deviation is negatively correlated with the weight. Next, mathematical expressions for each factor are designed: angular velocity is expressed as a reciprocal, light intensity as a linear form, and angle deviation as an exponentially decaying form. Finally, the expressions for each factor are weighted and combined, with the denominator added... To avoid the denominator being zero, we arrive at the final formula. Adjustment coefficient. , 、 、 Through experimental calibration, such as in the exhibition hall scene, multiple tests are determined 、 、 、 , ensure that each factor contributes to a balanced contribution. In practical applications, for example, the first audience angular velocity , the second perspective image light intensity , the angle of the audience head rotation , the optimal observation angle , the formula is calculated as , this high weight indicates that the corresponding perspective image of the audience needs to be updated first.
[0036] The technical effects achieved by the above embodiments include: realizing multi-dimensional factor coordinated optimization of perspective weight, ensuring that the audience with fast movement, appropriate lighting and small angle deviation can obtain optimized perspective first, and improving the overall multi-user viewing experience.
[0037] In the prior art, the visual parameter adjustment triggered by the interaction instruction does not combine the perspective optimization weight and the attribute of the instruction itself, resulting in unreasonable response parameters and poor interaction and visual linkage effect.
[0038] Therefore, when the interaction instruction synchronously triggers 3D visual parameter adjustment in S4 of the AI intelligent image adaptive naked eye 3D visual display method, an interaction-visual linkage response parameter calculation formula is used, which is: ;
[0039] wherein is the 3D visual parameter adjustment response value corresponding to the kth interaction instruction, dimensionless, is the priority of the kth interaction instruction, the value range is 1-10, dimensionless, the larger the value, the higher the priority, is the trigger duration of the kth interaction instruction, the dimension is second (s), is the time variation rate of the multi-dimensional perspective optimization weight, the dimension is 1 / second (1 / s), 、 、 is a proportional coefficient, and all are constants greater than 0, dimensionless, used to adjust the influence degree of the interaction instruction related factors and the weight variation rate on the response value.
[0040] It is worth mentioning that: the design theory basis of the formula is that the interaction-visual linkage needs to combine the perspective optimization state and the attribute of the instruction, and the multi-dimensional perspective optimization weight Reflects the current perspective optimization state of the audience, the higher the weight, the more sensitive the audience is to visual changes, so it is used as the base coefficient; priority of interactive instructions Reflects the importance of the instruction, for example, the product detail viewing instruction priority is set to 10, and the screen zooming instruction priority is set to 5, the higher the priority, the greater the response value should be; instruction trigger duration Reflects the effectiveness of the instruction, the longer the duration, the stronger the audience's demand for the interaction, the greater the response value should be; weight time change rate Reflects the trend of the perspective optimization state, when the change rate is positive, the weight increases, the audience's perspective optimization effect improves, and the response needs to be enhanced to match the optimization trend. Formula derivation process: First, determine the factors that affect the response parameter as the perspective optimization weight, instruction priority, instruction duration, and weight change rate; Then establish a positive correlation between each factor and the response value; Next, combine the perspective optimization weight with the instruction attributes, and add a correction term for the weight change rate; Finally, adjust the contribution of each factor through a proportional coefficient to get the final formula. Proportional coefficient 、 、 Determined through experiments, for example, in a commercial advertising scene, 、 、 In practical applications, for example, the 3rd interactive instruction priority , trigger duration , corresponding to the audience's , weight time change rate , into the formula , this response value is used to determine the adjustment amplitude of the visual parameter, such as increasing the brightness by 40% and the contrast by 16%.
[0041] The technical effects achieved by the above embodiments include: realizing the coordinated response of interactive instructions and perspective optimization state, ensuring reasonable response parameters, smooth interaction and visual linkage, and improving the sense of immersion of interaction.
[0042] In the prior art, when the AI multi-perspective image adaptation engine reversely optimizes the rendering parameters of the interactive region, the adjustment basis is not clear, the adjustment amplitude and the matching degree of the interactive instruction are low, and the visual effect of the interactive region is poor.
[0043] Therefore, in the AI intelligent image adaptation naked-eye 3D visual display method, when the AI multi-perspective image adaptation engine reversely optimizes the rendering parameters of the interactive region in S4, the interactive-visual linkage response parameter is used to adjust the pixel brightness, contrast, color saturation of the interactive region, and the texture resolution and edge sharpening degree of the 3D model, and the adjustment amplitude is positively correlated with the value of .
[0044] It is worth mentioning that: the rendering parameter adjustment adopts a linear mapping method, which maps the value range (0-10) of the interaction-visual linkage response parameter to the adjustment range of each rendering parameter. When the pixel brightness adjustment range is , for example , the brightness is increased by 20%; when the contrast adjustment range is , the contrast is increased by 16%; when the color saturation adjustment range is , , the saturation is increased by 12%. The texture resolution adjustment of the 3D model is determined according to , when , the texture resolution remains the original resolution (such as 512x512); when 3 , the texture resolution is increased to 1.5 times the original resolution (such as 768x768); when , the texture resolution is increased to 2 times the original resolution (such as 1024x1024). The edge sharpening degree is adjusted by the intensity parameter of the sharpening filter, and the intensity parameter value is equal to , for example , the intensity parameter is 0.4, which ensures clear edges without noise. During the adjustment process, the AI multi-view image adaptation engine applies the adjusted parameters to the interactive area through the rendering pipeline, for example, using the OpenGL rendering framework, the brightness and contrast parameters are transmitted into the fragment shader to adjust the pixel color in real time; the texture resolution and edge sharpening parameters are transmitted into the vertex shader to update the texture sampling and edge processing logic of the 3D model. The adjusted rendering effect is verified by the display effect feedback signal collected by the feedback transmission module, and if the brightness is too high to cause overexposure, the engine automatically reduces the adjustment range to ensure the best visual effect.
[0045] The technical effects achieved by the above embodiments include: clearly defining the rendering parameter adjustment basis and range, ensuring that the visual effect of the adjusted interactive area matches the interaction instruction, and improving the visual clarity and immersion during interaction.
[0046] In the prior art, the parameters shared by the naked-eye 3D digital human collaboration module are incomplete, and the digital human parallax adjustment lacks a reference and dynamic offset, resulting in low digital human and scene fusion.
[0047] Based on this, the space coordinates shared by the naked-eye 3D digital human collaboration module in S5 of the AI-based intelligent image adaptation naked-eye 3D visual display method include three-dimensional coordinates in the world coordinate system, with a dimension of meters (m), and the visual parameters include the field of view angle, the depth range, and the perspective projection matrix, the field of view angle has a dimension of radians (rad), and the depth range has a dimension of meters (m). When the AI multi-view image adaptation engine adjusts the parallax of the digital human model, the depth information of the 3D model in the scene is used to determine the parallax reference value of the digital human model, and then the dynamic offset adjustment is performed according to the audience position coordinates.
[0048] It is worth mentioning that: the space coordinates are shared by using the world coordinate system calibration, the coordinates of the 3D model in the scene are set by modeling software (such as Blender), for example, the coordinates of the base of the 3D cultural relic model are (0, 0, 0) m, and the coordinates of the top are (0, 0, 2) m; the coordinates of the digital human model are set by the naked-eye 3D digital human collaboration module, for example, the coordinates of the standing position of the digital human are (1, 0, 0) m, to ensure that the digital human and the cultural relic model are in the same space. In the sharing of visual parameters, the field of view angle is determined by the naked-eye 3D screen parameters, for example, the horizontal field of view angle of the screen is 60° = 1.047 rad, and the field of view angle of the digital human is also set to 1.047 rad; the depth range is determined according to the viewing area range, for example, the z-axis range of the viewing area is 2-8 m, and the depth range is set to 2-8 m; the perspective projection matrix is calculated according to the screen resolution and the field of view angle, and the OpenGL perspective projection matrix formula is used to ensure that the perspective effect of the digital human and the scene is consistent. The parallax adjustment of the digital human is divided into two steps, the first step is to determine the parallax reference value, which is based on the depth information of the 3D model in the scene, for example, the average depth of the cultural relic model is 1 m, and the parallax reference value of the digital human is set to 1 m (the parallax is inversely proportional to the depth, the greater the depth, the smaller the parallax); the second step is dynamic offset adjustment, which calculates the offset amount according to the z value (perpendicular to the screen direction) of the audience position coordinates, when the z value of the audience is 3 m, the offset amount is 0.2 m, and finally the parallax of the digital human is 1 m + 0.2 m = 1.2 m, to ensure that the parallax of the digital human matches the scene when the audience is at different z positions. The occlusion relationship processing adopts the depth buffer algorithm, the AI multi-view image adaptation engine calculates the depth values of the digital human and the 3D model in the scene in real time, the object with a smaller depth value occludes the object with a larger depth value, for example, the depth value of the digital human arm is 0.8 m, and the depth value of the corresponding area of the cultural relic model is 1 m, the engine displays the digital human arm and hides the occluded cultural relic area.
[0049] The above embodiments achieve the technical effects including: realizing complete parameter sharing of the digital human and the scene, accurate parallax adjustment of the digital human, natural fusion with the scene, eliminating the sense of suspension, and improving the visual coherence.
[0050] The prior art naked eye 3D visual display system has unclear module division, unclear module function, and lacks closed-loop data interaction, resulting in unstable system operation.
[0051] Therefore, the AI intelligent image adaptation naked eye 3D visual display system includes a scene data acquisition module, a naked eye 3D screen parameter acquisition module, an AI multi-view image adaptation engine module, a camera array module, an interactive device module, a naked eye 3D digital human collaboration module, a naked eye 3D display module, and a feedback transmission module. The scene data acquisition module is used to acquire the viewing area range, the initial audience position, the environmental lighting information, and the interactive device deployment position. The naked eye 3D screen parameter acquisition module is used to acquire the screen resolution, the display area size, and the pixel arrangement mode. The AI multi-view image adaptation engine module is used to receive the data transmitted by the scene data acquisition module and the naked eye 3D screen parameter acquisition module, perform perspective decomposition and dynamic reconstruction on the 3D model or video based on deep learning, generate a multi-view image sequence, update the perspective image sequence according to the audience position coordinates transmitted by the camera array module, reversely optimize the interactive area rendering parameters, and adjust the parallax and occlusion relationship of the digital human model. The camera array module is used to dynamically track the audience position, acquire the audience position coordinates, and transmit them to the AI multi-view image adaptation engine module. The interactive device module includes a touch device, a gesture recognition device, and a line-of-sight tracking device, and is used to generate interactive instructions and transmit them to the AI multi-view image adaptation engine module. The naked eye 3D digital human collaboration module is used to realize the sharing of the spatial coordinates and visual parameters of the digital human model and the scene, control the digital human model to enter the interactive area in a naked eye 3D posture, and display the perspective image sequence transmitted by the AI multi-view image adaptation engine module. The feedback transmission module is used to acquire the audience position feedback signal, the interactive instruction feedback signal, and the display effect feedback signal, and transmit them to the AI multi-view image adaptation engine module to form a closed-loop transmission link.
[0052] It is worth mentioning that: the scene data acquisition module adopts STM32F407 microcontroller as the core, connects laser ranging sensor, high-definition camera, light sensor array, GPS module, and data is transmitted to the AI multi-view image adaptation engine module through RS485 bus. The naked eye 3D screen parameter acquisition module reads the EDID information of the screen through the HDMI interface, obtains the screen resolution and pixel arrangement, measures the physical size of the screen through the laser range finder to obtain the display area size, and transmits the data to the engine module through the USB interface. The AI multi-view image adaptation engine module adopts NVIDIA Jetson AGXXavier development board, carries Ubuntu 20.04 operating system, realizes deep learning model based on TensorFlow 2.8 framework, receives data of each module through Ethernet, performs view angle decomposition, reconstruction, update and parameter optimization, and the processing result is transmitted to the naked eye 3D display module through HDMI2.1 interface. The camera array module is composed of 6 USB3.0 high-definition cameras, and the OpenCV library is used to realize the audience position tracking, and the data is transmitted to the engine module through gigabit Ethernet. In the interactive device module, the touch device adopts a capacitive touch screen (resolution 4K), the gesture recognition device adopts Kinect for Windows v2, and the gaze tracking device adopts Tobii Pro Fusion. The devices are connected to the engine module through USB3.0 or Ethernet, and the generated interactive instructions are transmitted in real time. The naked eye 3D digital human cooperation module is developed by using Unity3D engine, loads digital human model (such as Metahuman model), shares parameters with the engine module through TCP / IP protocol, and controls the posture and position of the digital human. The naked eye 3D display module adopts LG 55-inch naked eye 3D LCD screen, supports 4K resolution, refresh rate 60Hz, receives image data of the engine module and displays. The feedback transmission module includes brightness sensor, infrared sensor and key feedback module, and the data is transmitted to the engine module through I2C bus to form a closed loop. The clock of each module is kept consistent through time synchronization protocol, ensuring the real-time data interaction.
[0053] The technical effects achieved by the above embodiment include: the system module is clear in division, the function is clear, the data interaction forms a closed loop, the system stable operation is ensured, and the multi-user naked eye 3D immersion experience is realized.
[0054] In the prior art, when the AI multi-view image adaptation engine module adjusts the digital human parallax, the interactive response parameters, the digital human and the scene depth are not combined, resulting in inaccurate parallax adjustment and low digital human and scene fusion degree.
[0055] Based on this, when the AI multi-view image adaptation engine module in the AI intelligent image adaptation naked-eye 3D visual display system adjusts the parallax of the digital human model, a digital human-scene collaborative parallax adjustment calculation formula is adopted, and the digital human-scene collaborative parallax adjustment calculation formula is: ;
[0056] Wherein, is the parallax adjustment value of the mth digital human model in the nth scene area, with the dimension of meters (m), is the three-dimensional depth value of the mth digital human model, with the dimension of meters (m), is the background depth value of the nth scene area, with the dimension of meters (m), , , is a correction coefficient, and all are constants greater than 0, dimensionless, used to correct the influence of the depth of the digital human itself, the depth of the scene background, and the perspective optimization weight on the parallax adjustment value, is an interactive-visual linkage response parameter, dimensionless, is a multi-dimensional perspective optimization weight, dimensionless.
[0057] It is worth mentioning that: the design theory basis of this formula is that the parallax of the digital human needs to be coordinated with the interactive state, the depth of the digital human itself, the depth of the scene, and the perspective optimization state. The interactive-visual linkage response parameter reflects the interaction intensity, and the stronger the interaction, the more obvious the parallax adjustment needs to be; the depth of the digital human itself is the basis of the parallax, and the greater the depth, the smaller the parallax; the depth of the scene background determines the reference benchmark of the parallax of the digital human, and the deeper the background, the more the parallax of the digital human should be adjusted; the perspective optimization weight reflects the perspective state of the audience, and the more the weight deviates from 0.5, the more the parallax needs to be fine-tuned to match the perspective. The formula derivation process is as follows: first, determine the factors affecting the parallax adjustment as the interactive response, the depth of the digital human, the depth of the scene, and the perspective weight; then establish the relationship between each factor and the parallax, the interactive response, the depth of the digital human, the depth of the scene, and the perspective weight deviation, which are positively correlated with the parallax; then take the interactive response as the basic coefficient, combine it with the depth of the digital human and the scene, and add the correction term of the perspective weight deviation; finally, adjust the contribution of each factor through the correction coefficient to obtain the final formula. The correction coefficients, , , are calibrated through experiments, for example, in an education scene, , , In actual application, for example, the depth of the second digital human model , the background depth of the fourth scene area , and the interactive response parameter , view angle optimization weight , by substituting the formula The parallax adjustment value ensures that the parallax of the digital human in the scene area matches the scene and has no floating feeling.
[0058] The technical effects achieved by the above embodiments include: realizing the parallax of the digital human and the multi-factor collaborative adjustment, ensuring the accuracy of the parallax, the natural fusion of the digital human and the scene, and improving the visual coherence.
[0059] The above is only a preferred embodiment of the present application, and does not limit the form of the present application in other forms. Any person skilled in the art can use the disclosed technical content to make changes or modifications as equivalent embodiments applied to other fields, but any simple modification, equivalent change and modification made on the basis of the technical essence of the present application to the above embodiments without departing from the technical solution content of the present application still belongs to the protection scope of the technical solution of the present application.
Claims
1. A method for naked-eye 3D visual display based on AI intelligent image adaptation, characterized in that, include: S1: Collect naked-eye 3D screen parameters and scene data. The naked-eye 3D screen parameters include screen resolution, display area size, and pixel arrangement. The scene data includes viewing area range, initial audience position, ambient lighting information, and interactive device deployment position. S2: Input naked-eye 3D screen parameters and scene data into AI multi-view image adaptation engine. The AI multi-view image adaptation engine performs perspective decomposition and dynamic reconstruction of 3D models or videos based on deep learning, and automatically generates a multi-view image sequence covering the viewing area. S3: The camera dynamically tracks the audience's position, obtains the audience's position coordinates in real time, and transmits them to the AI multi-view image adaptation engine. The AI multi-view image adaptation engine updates the view image sequence in real time according to the audience's position coordinates. S4: Interactive commands synchronously trigger 3D visual parameter adjustments, and the AI multi-view image adaptation engine reversely optimizes the rendering parameters of the interactive area. S5: The naked-eye 3D digital human collaboration module enables the digital human model to share spatial coordinates and visual parameters with the scene. During interaction, the digital human model enters the interaction area in a naked-eye 3D posture. The AI multi-view image adaptation engine adjusts the parallax and occlusion relationship of the digital human model in real time. S6: Form a closed-loop process from scene data, AI adaptation, naked-eye 3D display, signal feedback to AI dynamic adjustment. The signal feedback includes audience position feedback signal, interaction command feedback signal, and display effect feedback signal.
2. The AI-based intelligent image adaptation naked-eye 3D visual display method according to claim 1, characterized in that, The ambient lighting information in S1 includes light intensity, light direction, and light spectrum distribution. The initial audience position is initially acquired through a camera array. The deployment position of the interactive device includes the coordinates of the touch device, the detection range of the gesture recognition device, and the effective monitoring angle of the eye tracking device.
3. The AI-based intelligent image adaptation naked-eye 3D visual display method according to claim 1, characterized in that, The deep learning model of the AI multi-view image adaptation engine in S2 adopts an encoder-decoder architecture. The encoder extracts the depth features of the 3D model or video, and the decoder dynamically reconstructs the extracted depth features according to the naked-eye 3D screen parameters and scene data. The generated multi-view image sequence covers all possible viewing angles of the audience within the viewing area, and the resolution of each view image matches the naked-eye 3D screen parameters.
4. The AI-based intelligent image adaptation naked-eye 3D visual display method according to claim 1, characterized in that, The camera array in S3 adopts a distributed deployment method, and the detection range of adjacent cameras overlaps. The frequency at which the camera array collects the audience's position coordinates in real time is consistent with the frequency at which the AI multi-view image adaptation engine updates the view image sequence. The audience's position coordinates include the three-dimensional coordinates of the audience's eyes and the head rotation angle.
5. The AI-based intelligent image adaptation naked-eye 3D visual display method according to claim 1, characterized in that, When the AI multi-view image adaptation engine in S3 updates the view image sequence based on the viewer's position coordinates, it uses a multi-dimensional view optimization weight calculation formula, which is as follows: ; in, Let the optimized weights be the image from the j-th viewpoint corresponding to the i-th viewer. Let be the angular velocity of the position movement of the i-th spectator, expressed in radians per second (rad / s). Let be the ambient light intensity corresponding to the j-th viewpoint image, with the dimension lux (lx). Let be the head rotation angle of the i-th audience member, measured in radians (rad). The optimal observation angle is measured in radians (rad). , , , These are adjustment coefficients, all of which are dimensionless constants greater than 0, used to balance the contribution of various influencing factors to the optimization weights.
6. The AI-based intelligent image adaptation naked-eye 3D visual display method according to claim 5, characterized in that, When the interactive command in S4 synchronously triggers the adjustment of 3D visual parameters, the interactive-visual linkage response parameter calculation formula is used. The interactive-visual linkage response parameter calculation formula is as follows: ; in, Adjust the response value for the 3D visual parameters corresponding to the k-th interaction command; this value is dimensionless. This represents the priority of the k-th interactive instruction, ranging from 1 to 10. It is dimensionless, with larger values indicating higher priority. Let be the duration of the triggering of the k-th interactive command, measured in seconds (s). The time change rate of the weights is optimized from a multi-dimensional perspective, with the dimension 1 / second (1 / s). , , These are proportionality coefficients, all of which are constants greater than 0, dimensionless, used to adjust the influence of interactive command-related factors and weight change rates on the response value.
7. The AI-based intelligent image adaptation naked-eye 3D visual display method according to claim 6, characterized in that, When the AI multi-view image adaptation engine in S4 reverse-optimizes the rendering parameters of the interactive area, it does so based on the interaction-visual linkage response parameters. Adjust the pixel brightness, contrast, and color saturation of the interactive area, as well as the texture resolution and edge sharpness of the 3D model, with varying degrees of adjustment. The values are positively correlated.
8. The AI-based intelligent image adaptation naked-eye 3D visual display method according to claim 1, characterized in that, The shared spatial coordinates of the naked-eye 3D digital human collaborative module in S5 include three-dimensional coordinates in the world coordinate system, with the dimension in meters. The visual parameters include field of view, depth range, and perspective projection matrix. The field of view is in radians, and the depth range is in meters. When the AI multi-view image adaptation engine adjusts the parallax of the digital human model, it determines the parallax reference value of the digital human model based on the depth information of the 3D model in the scene, and then makes dynamic offset adjustments based on the viewer's position coordinates.
9. An AI-based intelligent image adaptation naked-eye 3D visual display system, applied to the AI-based intelligent image adaptation naked-eye 3D visual display method as described in any one of claims 1-8, characterized in that, The system includes a scene data acquisition module, a naked-eye 3D screen parameter acquisition module, an AI multi-view image adaptation engine module, a camera array module, an interactive device module, a naked-eye 3D digital human collaboration module, a naked-eye 3D display module, and a feedback transmission module. The scene data acquisition module collects information on the viewing area, initial audience position, ambient lighting, and interactive device deployment location. The naked-eye 3D screen parameter acquisition module collects information on screen resolution, display area size, and pixel arrangement. The AI multi-view image adaptation engine module receives data from the scene data acquisition module and the naked-eye 3D screen parameter acquisition module, performs perspective decomposition and dynamic reconstruction on the 3D model or video based on deep learning, generates a multi-view image sequence, updates the perspective image sequence based on the audience position coordinates transmitted by the camera array module, and inversely optimizes the rendering parameters of the interactive area. The system adjusts the parallax and occlusion relationships of the digital human model; the camera array module dynamically tracks the audience's position, collects the audience's position coordinates, and transmits them to the AI multi-view image adaptation engine module; the interactive device module includes a touch device, a gesture recognition device, and a gaze tracking device, used to generate interactive commands and transmit them to the AI multi-view image adaptation engine module; the naked-eye 3D digital human collaboration module is used to realize the sharing of spatial coordinates and visual parameters between the digital human model and the scene, and controls the digital human model to enter the interactive area in a naked-eye 3D posture; the naked-eye 3D display module is used to receive and display the viewpoint image sequence transmitted by the AI multi-view image adaptation engine module; the feedback transmission module is used to collect audience position feedback signals, interactive command feedback signals, and display effect feedback signals, and transmit them to the AI multi-view image adaptation engine module, forming a closed-loop transmission link.
10. The AI-based intelligent image adaptation naked-eye 3D visual display system according to claim 9, characterized in that, When adjusting the parallax of the digital human model, the AI multi-view image adaptation engine module uses the digital human-scene collaborative parallax adjustment calculation formula, which is as follows: ; in, Let be the disparity adjustment value of the m-th digital human model in the n-th scene region, with the dimension in meters (m). Let be the three-dimensional depth value of the m-th digital human model, in meters (m). Let be the background depth value of the nth scene region, measured in meters (m). , , These are correction coefficients, all of which are dimensionless constants greater than 0. They are used to correct the impact of the digital human's own depth, scene background depth, and viewpoint optimization weights on the parallax adjustment value. The parameters for interactive-visual linkage response are dimensionless. Optimize weights from multiple perspectives; dimensionless.