A naked-eye 3D display method for intelligent robot applications
Through the closed-loop optimization of multi-sensing fusion and multi-modal interaction, the dynamic adaptability and interaction reliability problems of naked-eye 3D display technology in robot interaction scenarios are solved, and efficient and stable naked-eye 3D display and human-computer interaction are achieved.
Patent Information
- Application Number
- CN202510716722.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing naked-eye 3D display technology is insufficient dynamic adaptability in robot interactive scenarios, has low interaction error tolerance, low multi-modal coordination efficiency, and is difficult to adapt to complex environment changes and user dynamic displacement in real time, resulting in poor display stability and frequent false triggers.
Multi-sensing fusion is used to build a high-precision environment three-dimensional model, combine the user tracking module to dynamically generate multi-view images and eliminate light field interference, implement instruction recognition and coupling rate matching through environment perception and multi-modal interaction module, set up a secondary instruction verification mechanism, and realize real-time update of interactive data and system iterative optimization through closed-loop optimization module.
It improves the dynamic adaptability of naked-eye 3D display and the reliability of human-computer interaction, reduces the false recognition rate, and enhances the fault tolerance and efficiency of interaction.
Smart Images

Figure CN120263959B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent robots, and in particular relates to naked-eye 3D display technology, and specifically discloses a naked-eye 3D display method applied to intelligent robots. Background Art
[0002] With the widespread application of intelligent robots in industrial collaboration, medical surgery, and home service scenarios, traditional 2D display interfaces can no longer meet the needs of complex tasks for stereoscopic visual feedback.
[0003] Existing naked-eye 3D display technology can provide three-dimensional visualization support. Through grating, lenticular lens, or light field design, the direction of light propagation is controlled, allowing the left and right eyes to capture images from different perspectives. Light barrier technology uses tiny optical elements on the screen surface to separate the left and right eye images. Leviticus technology refracts light through a lens array to achieve multi-angle imaging. Light field technology controls the light propagation path to enhance the sense of three-dimensional reality. Multi-view technology synthesizes images from multiple perspectives to expand the viewing range. However, in robot interaction scenarios, it has the following significant drawbacks:
[0004] Insufficient dynamic adaptability: Traditional naked-eye 3D displays rely on fixed viewing angles or pre-calibrated models, making it difficult to adapt to complex environmental changes in real time, such as light interference and dynamic user movement. This results in poor display stability and is prone to image distortion or viewing angle deviation.
[0005] Low interaction fault tolerance: Most systems use a single instruction recognition mechanism, lacking coupling rate matching and secondary verification logic. False triggering and missed recognition are prominent problems, and reliability is significantly reduced, especially in scenarios with noise interference or multiple concurrent instructions.
[0006] Inefficient multimodal collaboration: Environmental perception and interaction modules often operate in isolation, and command responses rely on a single modality, such as voice or touch. Insufficient cross-modal data fusion leads to delayed or biased intent understanding.
[0007] Therefore, a dynamic adaptive, interactively reliable and efficient response method is needed to solve the above problems. Summary of the Invention
[0008] In view of this, the present invention proposes a naked-eye 3D display method for intelligent robot applications, which constructs a high-precision three-dimensional environmental model through multi-sensor fusion, dynamically generates multi-perspective images and eliminates light field interference in combination with a user tracking module, utilizes environmental perception and multimodal interaction modules to realize command recognition and coupling rate matching, and sets a secondary command verification mechanism to improve command accuracy by determining the overlap rate. Finally, real-time updating of interaction data and system iterative optimization are realized through a closed-loop optimization module. This method effectively improves the dynamic adaptability and human-computer interaction reliability of naked-eye 3D display through a closed-loop architecture of environmental perception-user tracking-multimodal interaction-double verification-data optimization.
[0009] The object of the present invention can be achieved by the following technical solution: A naked-eye 3D display method for intelligent robot application, characterized by comprising a multi-sensor fusion module, a user tracking module, an environment perception module, a multimodal interaction module and a closed-loop optimization module, specifically comprising the following steps:
[0010] S1, 3D scene model construction: The robot uses a multi-sensor fusion module to collect environmental point cloud data in real time, and combines it with algorithms to build a high-precision 3D scene model;
[0011] S2. User Position Tracking: In the 3D scene model, the user tracking module is used to obtain user position data, the 3D content generation unit is used to dynamically generate multi-view image sequences, and the light field control unit is used to eliminate environmental interference;
[0012] S3. Command data collection: The environment perception module identifies the user's command and collects the command data;
[0013] S4. Command data analysis: Analyze the collected command data through the multimodal interaction module, extract key information and compare it with the database to obtain the coupling rate. When the coupling rate is greater than the set value, output the result once;
[0014] S5. Secondary instruction analysis: If the user issues a secondary instruction, the secondary instruction is compared with the primary instruction to obtain the overlap rate. When the overlap rate is greater than or equal to the set value, the secondary instruction and the primary instruction are confirmed to be the same instruction, the primary result is excluded and re-output. If the overlap rate is less than the set value, the secondary verification is triggered and step S4 is repeated;
[0015] S6, closed-loop feedback optimization: The closed-loop optimization module collects interaction data from S3 to S5 and updates the instruction database in real time.
[0016] Combining all the above technical solutions, the present invention has the following positive effects:
[0017] 1. This invention uses multi-sensor fusion and user tracking modules to dynamically generate 3D content and eliminate light field interference in combination with real-time environmental data, thereby improving the stability and accuracy of naked-eye 3D display in complex scenes.
[0018] 2. The present invention adopts multimodal instruction analysis to perform coupling rate matching and secondary instruction coincidence rate verification mechanism, which effectively reduces the misrecognition rate and enhances the fault tolerance and accuracy of human-computer interaction.
[0019] 3. The present invention integrates environmental perception and multimodal input of voice / motion commands to achieve natural, low-latency intention understanding and feedback, thereby improving interaction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Attachment Figure 1 It is a system diagram of the present invention.
[0022] Attachment Figure 2 It is a step diagram of the present invention.
[0023] Attachment Figure 3 Flowchart of the present invention. DETAILED DESCRIPTION
[0024] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0025] See also Figure 1 As shown, the present invention proposes a naked-eye 3D display method for intelligent robot applications, which is characterized by comprising a multi-sensor fusion module, a user tracking module, an environmental perception module, a multimodal interaction module and a closed-loop optimization module, wherein the user tracking module includes a 3D content generation unit and a light field control unit.
[0026] like Figure 2 As shown, the specific implementation steps of the present invention include the following steps:
[0027] S1, 3D scene model construction: The robot uses a multi-sensor fusion module to collect environmental point cloud data in real time, and combines it with algorithms to build a high-precision 3D scene model;
[0028] It should be noted that the robot body is equipped with a multi-sensor fusion system, which includes multiple cameras, depth sensors and lidars. It uses an extended Kalman filter algorithm to fuse the data of multiple cameras, depth sensors and lidars, eliminate the noise of a single sensor, and combine the algorithm to identify and classify objects in the scene to generate a labeled 3D point cloud model.
[0029] S2. User Position Tracking: In the 3D scene model, the user tracking module is used to obtain user position data, the 3D content generation unit is used to dynamically generate multi-view image sequences, and the light field control unit is used to eliminate environmental interference;
[0030] It should be noted that the 3D content generation unit is specifically:
[0031] The perspective change is predicted based on the user's position and the robot's motion state, the display viewpoint range and viewpoint density are calculated, and a multi-view image sequence is dynamically generated through the 3D content generation unit.
[0032] Input the robot's motion state, user position change, and environmental obstacle distribution. The robot's motion state includes its velocity, acceleration, and pitch angle data. Output the horizontal angle α and vertical angle β of the viewpoint range, where α∈[-45°, +45°] and β∈[-20°, +20°].
[0033] Calculate the binocular parallax based on the relative motion relationship between the user position and the robot:
[0034] ;
[0035] Where d is the binocular parallax; B is the binocular camera baseline distance, which represents the horizontal distance between the optical centers of the left and right cameras. The larger the baseline, the higher the parallax calculation accuracy, but this needs to be balanced with the robot's structural size limitations.
[0036] f is the focal length of the camera, which determines the viewing angle and spatial resolution of the imaging system;
[0037] Z is the real-time distance between the user and the robot, directly measured by the depth sensor.
[0038] This formula is based on the principle of stereo vision triangulation. It infers the user depth through the parallax of the left and right images. When the user moves, that is, Z changes, the image rendering parallax d is dynamically adjusted to ensure the consistency of the stereoscopic sense when the user moves.
[0039] The viewpoint density adaptation formula is as follows:
[0040] ;
[0041] Where ρ is the viewpoint density per unit area; N is the number of viewpoints; is the angle between the user's line of sight and the screen normal.
[0042] The viewpoint density per unit area is dynamically adjusted according to the user's distance. When the user gets closer, that is, when Z decreases, the viewpoint density is automatically increased to improve the stereoscopic accuracy.
[0043] It should be noted that when dynamically generating multi-view image sequences, high-resolution rendering is used for the user's gaze area, and low-resolution rendering is used for the non-gaze area to reduce computing power consumption. The rendered image is also super-resolution and denoising processed to compensate for the texture details in the low-resolution area and improve the display quality of low-light or complex texture scenes.
[0044] It should be specifically stated that the light field control unit is specifically:
[0045] The intelligent robot is equipped with a light field display, which has a built-in cylindrical lens array and a dynamically adjustable backlight module. The light field control unit dynamically controls the refraction angle and backlight direction of the adjustable cylindrical lens array to reduce image crosstalk in multi-user scenarios.
[0046] First, the refraction angle and spacing of the adjustable cylindrical lens array are dynamically adjusted based on the user's distance and ambient light intensity to optimize the light propagation path.
[0047] The specific spacing of the cylindrical lens is:
[0048] ;
[0049] Where L is the cylindrical lens spacing;
[0050] K is the refractive index compensation coefficient and K∈[0.8, 1.2], which needs to be calibrated through optical calibration experiments;
[0051] D is the user distance between 0.5-3m, which directly affects the scaling ratio of the lens spacing;
[0052] θ: The incident angle of ambient light is between 0° and 60°, which needs to be measured in real time by a light field sensor to compensate for optical path deviation;
[0053] Dynamically adjust lens spacing to accommodate different user distances and ambient light angles.
[0054] Then, the backlight zone control technology is used to adjust the brightness and color temperature of different areas to suppress ambient light interference and image crosstalk in multi-user scenarios;
[0055] When the ambient light intensity E≥1000lux, the backlight partition control is activated, the screen is divided into N×M grids, and the brightness of each grid is independently adjusted to the target value.
[0056] The backlight brightness is as follows:
[0057] ;
[0058] Where I is the backlight brightness;
[0059] E is the ambient light intensity, which is collected in real time by a photosensor;
[0060] Δx, Δy are the horizontal / vertical offsets between the user’s line of sight and the center of the screen;
[0061] k is a calibration coefficient between 100-200, which is related to the luminous efficiency of the backlight module.
[0062] The formula achieves anti-glare adjustment by making the brightness inversely proportional to the ambient light intensity. At the same time, it dynamically allocates backlight resources according to the line of sight offset to reduce multi-user crosstalk. For example, the larger the offset, the higher the brightness of the corresponding area.
[0063] S3. Command data collection: The environment perception module identifies the user's command and collects the command data;
[0064] Infrared pyroelectric sensors can be used to detect user movement trajectories, such as gesture operations; light-sensitive sensors can be used to monitor changes in ambient light intensity to trigger the acquisition system; and a voice recognition module can be configured to capture voiceprint features and voice commands in real time. This embodiment will not be explained in detail.
[0065] S4. Instruction data analysis: Figure 3 As shown, the multimodal interaction module analyzes the collected instruction data, extracts key information and compares it with the database to obtain the coupling rate. When the coupling rate is greater than the set value, the result is output once;
[0066] It should be noted that the coupling rate formula is:
[0067] ;
[0068] CR(A, X) is the coupling ratio, which indicates the coupling ratio between the key information A extracted by an instruction and a certain data X in the database. The range is between 0 and 1, and the higher the value, the stronger the match.
[0069] The similarity index The measurement result of the i-th similarity needs to be normalized to between 0 and 1 to quantify the degree of match between A and X in a certain feature dimension.
[0070] It should be noted that different similarity algorithms require different data. For example, for images, SSIM structural similarity and feature vector-based cosine similarity are used; for speech, dynamic time warping and MFCC cosine similarity are used; and for text, Jaccard similarity and TF-IDF cosine similarity are used.
[0071] In order to ensure that all similarity values range from 0 to 1, normalization is required. Since the original value ranges of different similarity indicators are different, they cannot be directly compared. Moreover, if normalization is not performed, the large range of indicators will dominate the results.
[0072] Common normalization methods include cosine similarity, Euclidean distance, and dynamic time warping, specifically:
[0073] Cosine similarity, the original range is -1, 1, normalized to:
[0074] .
[0075] Euclidean distance, original distance d≥0, normalized to:
[0076] .
[0077] Dynamic Time Warping (DTW) converts the minimum path cost into similarity after the speech sequence is aligned:
[0078] ;
[0079] in is the attenuation coefficient.
[0080] Where Wi represents the weight of the similarity index in the i-th order, indicating its importance. The weight is set to 1, but the formula allows for free adjustment.
[0081] It should be explained that the method for determining the weights can be allocated according to business needs, and the machine determines the contribution of each feature through training data optimization and principal component analysis.
[0082] The number of indicators n is the total number of similarity indicators used, which can be expanded according to needs. When n=1, it is a single indicator and only one method is used. When n≥2, multi-dimensional features such as color, texture and semantics are integrated.
[0083] If a certain weight = 1 and the others are 0, the formula degenerates into a single indicator calculation; if the weights are evenly distributed, weight = 1 / n, then all indicators are equally important.
[0084] It should be noted that if data A is mixed type data, such as containing images and text, the image coupling rate and text coupling rate can be calculated separately, and then fused through weighted or product fusion;
[0085] ;
[0086] For example, when extracting picture A, similar pictures are retrieved from the database.
[0087] The similarity metrics are as follows:
[0088] sim1: SSIM, structural similarity, weight is 0.5;
[0089] sim2: color histogram intersection, weight 0.3;
[0090] sim3: CNN feature cosine similarity, weight is 0.2;
[0091] Coupling ratio CR(A, X) = 0.5*sim1+0.3*sim2+0.2*sim3;
[0092] It should be noted that when the coupling ratio CR(A, X) ≥ the set value x, a result is output once, where the range of x is between 80% and 100%, determined according to the specific situation. If the coupling ratios of multiple data are greater than the set value, the result with the highest coupling ratio is selected for output.
[0093] S5. Secondary instruction analysis: Figure 3 As shown, if the user issues a secondary instruction, the secondary instruction is compared with the primary instruction to obtain the overlap rate. When the overlap rate is greater than or equal to the set value, it is confirmed that the secondary instruction and the primary instruction are the same instruction, the primary result is excluded and re-output. If the overlap rate is less than the set value, the secondary verification is triggered and step S4 is repeated;
[0094] It should be noted that the formula for the overlap rate is:
[0095] ;
[0096] Among them, O(A, B) is the overlap rate, which indicates the overlap rate between the key information B extracted by the secondary instruction and the key information A extracted by the primary instruction. The higher the value, the higher the overlap. The numerator represents the weighted sum of the multi-dimensional feature matching degree, and the denominator is the weight normalization to ensure that the result is within the range of 0-1.
[0097] in Indicates the matching degree of the kth feature dimension, with a value range between 0 and 1. It should be noted that regardless of the data type, it is mapped to a unified feature space, such as vector, sequence, and graph structure.
[0098] It should be explained that the matching method is selected according to the feature type, and all results are normalized to 0-1. The multi-dimensional matching degree calculation is as follows:
[0099] Semantic vector, MFCC mean selection vector feature matching, specifically:
[0100] .
[0101] Speech waveform and gesture trajectory selection sequence feature matching, specifically:
[0102] ;
[0103] in is the attenuation coefficient.
[0104] The text dependency tree and gesture skeleton graph select structured feature matching, specifically:
[0105] ;
[0106] in Represents the dynamic weight of the kth dimension. The weight here may not be 1. The weight is automatically adjusted according to the feature importance and data quality. If the variance of a dimension feature in the data is large, it indicates that its discrimination is high and is given a higher weight. Specifically:
[0107] ;
[0108] in is the variance of dimension k.
[0109] It should be explained that the weights are manually preset according to business requirements. For example, if the weight of semantic features is greater than that of temporal features and structural features, ω is set to [0.5, 0.3, 0.2].
[0110] Where C(A, B) represents the data type consistency correction factor, which ranges from 0 to 1 and quantifies the compatibility of cross-modal data.
[0111] If A and B are of the same type, such as both are voice, then C=1.
[0112] If the types are different but can be matched across modalities, such as image and text, then C=0.8, which is a manually set threshold.
[0113] If the types are completely incompatible, such as voice and gesture coordinates, C=0.
[0114] Example: Voice and text matching:
[0115] After speech-to-text conversion, the edit distance Match1 and the semantic vector cosine similarity Match2 are calculated.
[0116] Weight distribution: ω1=0.7, indicating that text matching is prioritized, ω2=0.3.
[0117] Consistency correction: C=0.8.
[0118] Gestures and Gesture Matching:
[0119] Skeleton sequence DTW distance Match1, hand direction vector cosine similarity Match2.
[0120] Weight distribution: ω1=0.6, indicating that timing alignment is important, ω2=0.4.
[0121] Consistency correction: C=1.
[0122] It should be specifically noted that when the overlap rate O(A, B) ≥ the set value y, the secondary instruction and the primary instruction are confirmed to be the same instruction, the primary output result is excluded, and the result of the second coupling rate is re-output; if the overlap rate O(A, B) < the set value y, the secondary verification is triggered, the S4 step is repeated, and the coupling rate comparison output is recalculated.
[0123] S6, closed-loop feedback mechanism: The closed-loop optimization module collects the interaction data from S3 to S5 and updates the instruction database in real time.
[0124] Incremental Learning Engine:
[0125] Deploy an online learning algorithm, triggering model weight updates every time 200ms window data is received;
[0126] Dynamically adjust data collection dimensions through feature importance analysis and eliminate redundant features with a contribution of less than 5%.
[0127] Command database update mechanism:
[0128] Double buffering technology is used to achieve lock-free updates: the main database continuously writes incremental data, and the mirror database performs full snapshots every hour;
[0129] ACID characteristics are implemented based on the transaction log (WAL) to ensure data rollback capabilities in the event of an abnormal power outage.
[0130] Dynamic Optimization Strategy:
[0131] Build a Bayesian inference network to evaluate the execution effect of instructions and automatically adjust the parameter weights of the S3-S5 stages;
[0132] Connect to the MES system through the API gateway and synchronize the optimized instruction set to the production scheduling module.
Claims
1. A naked-eye 3D display method for intelligent robot applications, characterized in that: It includes a multi-sensor fusion module, a user tracking module, an environmental perception module, a multimodal interaction module, and a closed-loop optimization module. Specifically, it includes the following steps: S1, 3D scene model construction: The robot uses a multi-sensor fusion module to collect environmental point cloud data in real time, and combines it with algorithms to build a high-precision 3D scene model; S2. User Position Tracking: In the 3D scene model, the user tracking module is used to obtain user position data, the 3D content generation unit is used to dynamically generate multi-view image sequences, and the light field control unit is used to eliminate environmental interference; S3. Command data collection: The environment perception module identifies the user's command and collects the command data; S4. Instruction data analysis: The collected instruction data is analyzed through the multimodal interaction module, key information is extracted and compared with the database, and the coupling ratio between the key information A extracted from the instruction and a certain data X in the database is calculated. The coupling ratio formula is as follows: ; Where CR(A, X) is the coupling ratio; is the similarity index; Wi represents the weight of the similarity index in the i-th one; the number of indicators n is the total number of similarity indicators used; When the coupling ratio CR(A, X) ≥ the set value x, the result is output once; S5. Secondary instruction analysis: If the user issues a secondary instruction, the secondary instruction is compared with the primary instruction to obtain the overlap rate. When the overlap rate is greater than or equal to the set value, the secondary instruction and the primary instruction are confirmed to be the same instruction, the primary result is excluded and re-output. If the overlap rate is less than the set value, secondary verification is triggered and step S4 is repeated. The formula for calculating the overlap rate of key information B extracted from the secondary instruction and key information A extracted from the primary instruction is as follows: ; Where O(A, B) is the coincidence rate; Indicates the matching degree of the kth feature dimension; C(A, B) represents the data type consistency correction factor; Represents the dynamic weight of the k-th dimension; When the overlap rate O(A, B) ≥ the set value y, confirm that the secondary instruction and the primary instruction are the same instruction, exclude the primary output result, and re-output the result of the second coupling rate; if the overlap rate O(A, B) < the set value y, trigger the secondary verification, repeat step S4, and recalculate the coupling rate comparison output; S6, closed-loop feedback optimization: The closed-loop optimization module collects interaction data from S3 to S5 and updates the instruction database in real time.
2. The naked-eye 3D display method for an intelligent robot according to claim 1, wherein: The user tracking module is specifically: Predicting the change in viewing angle based on the user's position and the robot's motion state, calculating the display viewpoint range and viewpoint density, and dynamically generating a multi-view image sequence through the 3D content generation unit; The refraction angle and backlight direction of the adjustable cylindrical lens array are dynamically controlled by the light field control unit.
3. The naked-eye 3D display method for an intelligent robot according to claim 2, wherein: The 3D content generation unit is specifically: Input the robot's motion state, user position change, and environmental obstacle distribution. The robot's motion state includes its velocity, acceleration, and pitch angle data. Output the horizontal angle α and vertical angle β of the viewpoint range, where α∈[-45°, +45°] and β∈[-20°, +20°]. Calculate the binocular parallax based on the relative motion relationship between the user position and the robot: ; Where d is the binocular parallax; B is the binocular camera baseline distance; f is the camera focal length; Z is the real-time distance between the user and the robot; The viewpoint density adaptation formula is as follows: ; Where ρ is the viewpoint density per unit area; N is the number of viewpoints; is the angle between the user's line of sight and the screen normal.
4. The naked-eye 3D display method for an intelligent robot according to claim 2, wherein: The light field control unit is specifically: The refraction angle and spacing of the adjustable lenticular lens array are dynamically adjusted based on the user's distance and ambient light intensity. The specific spacing of the lenticular lenses is: ; Where L is the cylindrical lens spacing; K is the refractive index compensation coefficient and K∈[0.8, 1.2]; D is the user distance; θ is the ambient light incident angle; The backlight partition control technology adjusts the brightness and color temperature of different areas. When the ambient light intensity E ≥ 1000 lux, the backlight partition control is activated, dividing the screen into N × M grids, and each grid independently adjusts the brightness to the target value; The backlight brightness is as follows: ; Where I is the backlight brightness; E is the ambient light intensity; Δx, Δy are the horizontal / vertical offsets between the user's line of sight and the center of the screen; and k is the calibration coefficient.
Citation Information
Patent Citations
Naked eye 3D augmented reality interaction display system and display method thereof
CN106131536A
System for collecting feedback instruction by applying database information
CN117971913A