AI interactive projection method and system based on three-dimensional laser radar
Through the combination of three-dimensional lidar and directional microphone arrays, the problem of limitations in recognition and small coverage of traditional lidar and infrared cameras in large-space projection interaction is solved, precise interaction and personalized projection of multiple people are achieved, and the system's decision-making accuracy and resource allocation efficiency are improved.
Patent Information
- Application Number
- CN202510416672.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-22
AI Technical Summary
In the large-space projection interaction, traditional lidar and infrared cameras have problems such as limitations in identification, small equipment coverage, complex calibration, difficult data fusion and great environmental impact, and are difficult to apply to multi-population scenarios.
Three-dimensional lidar combined with directional microphone arrays are used to fuse three-dimensional point cloud data and voice data through multi-modal models to perform multi-dimensional weighted confidence calculations, combined with deep neural networks and trajectory prediction models to achieve accurate interactive command decisions and personalized projections of the population.
It improves the accuracy of multi-group mobile trajectory recognition, reduces system computing pressure, avoids invalid or incorrect instructions, provides a personalized interactive experience, and enhances the reasonable allocation and coverage area of projection resources.
Smart Images

Figure CN120354345A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of stage projection interaction, and more specifically, relates to an AI interactive projection method and system based on a three-dimensional lidar. Background Art
[0002] In existing large-space projection interaction, traditional lidar or infrared cameras are mainly used to identify and track objects in the space. However, both of these methods have their respective defects: 1. Traditional lidar can only track the trajectories of moving objects in a plane. In actual use, only the accuracy of data on the position attribute of moving objects can be achieved, and there are significant limitations in distinguishing between people and objects, making it not suitable for application scenarios with a large number of people; 2. Although cameras and infrared cameras have functions such as being able to identify various objects, the coverage area of a single device is small. When multiple devices are deployed, there are problems such as complex calibration and difficult data fusion, and they are also easily affected by environmental factors such as light and occlusion.
[0003] Therefore, it is necessary to propose an AI interactive projection method and system based on a three-dimensional lidar to solve the above technical problems. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an AI interactive projection method and system based on a three-dimensional lidar, which can perform multi-dimensional weighted confidence calculation on the interaction instructions of a crowd, improve the accuracy of decision-making, reduce the computing pressure of the system, avoid invalid or incorrect instructions, and reasonably allocate resources to the spatial group; at the same time, it can also provide personalized interactive experiences for different users.
[0005] An AI interactive projection method based on a three-dimensional lidar of the present invention includes steps S1 - S4.
[0006] Step S1: Acquisition of human body three-dimensional point cloud data. Specifically, the three-dimensional lidar emits laser beams and receives reflection signals to generate high-precision three-dimensional point cloud data, capturing the position and shape of the human body. Extract key points of the wrist and fingertips of the hand bones to capture the dynamic gestures of the human body. Acquisition and recognition of voice data. Specifically, deploy a directional microphone array to collect command voices and suppress on-site noise.
[0007] Step S2: Combine the three-dimensional point cloud data and voice data obtained in step S1 to form a data set, and input it into the AI unit. The AI unit obtains the total confidence by weighting the confidence of each modality, and makes a decision on whether to generate an interactive projection based on the fusion of multiple modalities. If the total confidence reaches the set threshold, the AI unit generates the corresponding interactive projection.
[0008] Step S3: Achieve the fusion of multi-modal information through a multi-modal model to generate the projection required by the user.
[0009] Step S4: Transfer the projection generated by the AI unit to the location where the user is through a projection device.
[0010] As a further improvement of the present invention, step S2 further includes step S21. Specifically, step S21 is to collect the movement trajectories and gesture change frequencies of each user, divide the user groups through a clustering algorithm. The user groups include active exploration types and conservative observation types, where the interaction frequency of the active exploration type user group is higher than that of the conservative observation type user group. Reduce the number of elements required for interactive projection by the active exploration type user group. Increase the number of elements required for interactive projection by the conservative observation type.
[0011] As a further improvement of the present invention, in step S2, classify each data in the data set, and the classification includes gestures, sounds, and relative positions. Evaluate the influence of each category on the decision of generating a projection. Set a total confidence evaluation formula where CL is the total confidence, W i is the weight of the i-th category, C i is the confidence of the i-th category, and n is the total number of categories participating in the calculation. The AI unit extracts the features of each category module and generates corresponding feature vectors through independent encoders, and obtains the confidence and weight of each category through dynamic calculation. If the total confidence does not reach the set threshold, the AI unit does not generate the interactive projection for this user.
[0012] As a further improvement of the present invention, step S2 further includes step S22. Specifically, step S22 is to dynamically adjust the projection content according to the user's body proportion and the projection environment time. By extracting the positions and hierarchical information of all elements on the projection interface, construct a hierarchical relationship diagram between the interface elements, optimize the layout of the elements on the interface according to the user's body proportion, and adjust the element ratio in the projection interface layout to adapt to the user's body proportion through the Canny edge detection algorithm. Adjust the colors, shapes, and gradients of the interface elements according to the projection environment time to better integrate into the projection interface, calculate the color distribution consistency score between the elements and the background, and determine the adjustment range.
[0013] As a further improvement of the present invention, it further includes step S5. According to the historical positions of the collected three-dimensional point cloud data, track the position changes of the user in the space, and through data analysis, infer the future movement trajectory of the user, so that the projection interface can move following the user's position changes. Determine whether the current movement trajectory of the user is within the preset stay range. If the current movement trajectory exceeds the preset stay range, stop the interactive projection for this user. If the movement trajectory is within the preset stay range, continue the interactive projection for this user.
[0014] An AI interactive projection system based on a 3D lidar, which includes an information unit, an AI unit, and a projection unit.
[0015] The information unit includes a 3D lidar and a directional microphone array, which are used to collect human body 3D point cloud data and command voices. The information unit forms a data set with the obtained 3D point cloud data and voice data.
[0016] The AI unit includes a multi-modal model, a single-modal model, a deep neural network model, and a trajectory prediction model. The multi-modal model is used to calculate the total confidence of whether the user makes a projection decision currently and generate a projection. The single-modal model adjusts the corresponding interactive projection according to the user's personalized information and the current environmental state. The deep neural network model is used to divide user groups through a clustering algorithm and adaptively adjust the number of elements in the interactive projections of different user groups. The trajectory prediction model is used to track and predict the movement trajectory of the user in space and select whether to continue the interactive projection according to the movement trajectory.
[0017] The projection unit includes a projection device and an acoustic-optical device. The projection device is used to output a corresponding interactive projection to the position where the user is currently located according to the AI unit. The acoustic-optical device is used to output lights and audio.
[0018] As a further improvement of the present invention, in the deep neural network model, user groups are divided through a clustering algorithm. The user groups include active exploration types and conservative observation types, where the interaction frequency of the active exploration type user group is higher than that of the conservative observation type user group. Reduce the number of elements required for the interactive projection of the active exploration type user group. Increase the number of elements required for the interactive projection of the conservative observation type.
[0019] As a further improvement of the present invention, in the multi-modal model, the features of each category module are generated into corresponding feature vectors through independent encoders. The multi-modal model can recognize gestures such as waving, clenching, and drawing circles through the feature vectors of the hand key points extracted, and calculate the corresponding confidence according to the clarity of the gesture actions. The multi-modal model can adjust and calculate the corresponding confidence according to the signal-to-noise ratio of the voice.
[0020] As a further improvement of the present invention, in the single-modal model, the projection content is dynamically adjusted according to the user's body proportion and the projection environment time. By extracting the position and hierarchical information of all elements on the projection interface, a hierarchical relationship graph between interface elements is constructed, and the layout of elements on the interface is optimized according to the user's body proportion. Through the Canny edge detection algorithm, the element proportion within the projection interface layout is adjusted to adapt to the user's body proportion. The colors, shapes, and gradients of the interface elements are adjusted according to the projection environment time to better integrate into the projection interface. The color distribution consistency score between the elements and the background is calculated to determine the adjustment amplitude. According to the overall visual effect of the interface and the coordination between the elements and the background, the parameters of the gradient effect are adjusted until the gradient effect is consistent with the background and the overall interface design. The parameters of the gradient effect include color transition, direction, and diffusion range.
[0021] As a further improvement of the present invention, in the trajectory prediction model, according to the historical positions of the collected three-dimensional point cloud data, the position changes of the user in the space are tracked, and through data analysis, the future movement trajectory of the user is inferred, so that the projection interface can move following the user's position change. It is judged whether the current movement trajectory of the user is within the preset stay range. If the current movement trajectory exceeds the preset stay range, the interactive projection for this user is stopped. If the movement trajectory is within the preset stay range, the interactive projection for this user continues.
[0022] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0023] 1. By setting up a three-dimensional lidar, a directional microphone array, and an AI unit, it is possible to improve the accuracy of identifying and tracking the movement trajectories of multiple groups of people, increase the coverage area of the device, and reduce the influence of environments such as light and occlusion.
[0024] 2. By setting up a multi-modal model, it can recognize gestures such as waving, clenching, and drawing circles through the feature vectors of the hand key points extracted. The corresponding confidence levels are calculated according to the clarity of the gesture movements and the signal-to-noise ratio of the voice. The total confidence level is obtained by weighting the confidence levels of each modality. A decision on whether to generate an interactive projection is made through the fusion of multiple modalities. If the total confidence level reaches the set threshold, the AI unit generates the corresponding interactive projection; a multi-dimensional weighted confidence level calculation is performed on the interaction instructions of the crowd, improving the accuracy of the decision-making, reducing the computing pressure of the system, avoiding invalid or incorrect instructions, and reasonably allocating resources within the spatial group.
[0025] 3. By setting up a single-modal model, the projection content is dynamically adjusted according to the user's body proportion and the projection environment time; the colors, shapes, and gradients of the interface elements are adjusted to better integrate into the projection interface; a personalized interactive experience is provided for different users. Description of the Drawings
[0026] Figure 1 Schematic flowchart of the method according to an embodiment of the present invention;
[0027] Figure 2 Schematic flowchart of step S2 of the method according to an embodiment of the present invention;
[0028] Figure 3 Schematic diagram of the system according to an embodiment of the present invention. Detailed implementation manners
[0029] Specific embodiment 1: Please refer to Figure 1 - Figure 2 An AI interactive projection method based on a three-dimensional lidar. It includes steps S1 - S5.
[0030] Step S1: Acquisition of human body three-dimensional point cloud data. Specifically, the three-dimensional lidar emits laser beams and receives reflected signals to generate high-precision three-dimensional point cloud data, capturing the position and shape of the human body. Key points of the hand bones, wrists and fingertips are extracted to capture the dynamic gestures of the human body. Voice data acquisition and recognition. Specifically, a directional microphone array is deployed to collect command voices and suppress on-site noise.
[0031] It should be noted that the acquisition of three-dimensional point cloud data and voice data through a three-dimensional lidar and a directional microphone array is an existing conventional technical means, and its specific working principle will not be elaborated here.
[0032] Step S2: The three-dimensional point cloud data and voice data obtained in step S1 are combined into a data set and input into the AI unit. The AI unit obtains the total confidence by weighting the confidence of each modality, and makes a decision on whether to generate an interactive projection through multi-modal fusion. If the total confidence reaches the set threshold, the AI unit generates the corresponding interactive projection.
[0033] In step S2, the data in the data set are classified, and the classification includes gestures, sounds and relative positions. The influence of each category on the decision of whether to generate a projection is evaluated. A total confidence evaluation formula is set where CL is the total confidence, W i is the weight of the i-th category, C i is the confidence of the i-th category, and n is the total number of categories participating in the calculation. The AI unit extracts the features of each category module and generates corresponding feature vectors through independent encoders, and obtains the confidence and weight of each category through dynamic calculation. If the total confidence does not reach the set threshold, the AI unit does not generate the interactive projection of this user.
[0034] In this embodiment, the multi-modal model can recognize gestures such as waving, clenching and drawing circles through the feature vectors of the hand key points extracted, and calculate the corresponding confidence C i. The multimodal model can adjust and calculate the corresponding confidence level C according to the signal-to-noise ratio of the voice. i . For example, the clearer the waving action, the greater its confidence level, and the lower the signal-to-noise ratio of the sound, the smaller its confidence level. At the same time, the consistency between the action and the sound is evaluated. The higher the consistency, the higher the confidence levels of both, and vice versa. For example, when the user says "move left", if the gesture points to the right, the confidence levels of the voice and gesture modalities need to be dynamically adjusted and decreased. Among them, the confidence weight W i can be set artificially in advance or dynamically adjusted by the multimodal model. Perform multi-dimensional weighted confidence level calculations on the interaction instructions of the crowd, improve the accuracy of decision-making, reduce the computing pressure of the system, avoid invalid or incorrect instructions, and reasonably allocate resources within the spatial group.
[0035] Preferably, in this embodiment, steps S21 - S22 are further included in step S2.
[0036] Step S21: Specifically, step S21 is to collect the movement trajectories and gesture change frequencies of each user, and divide the user group through a clustering algorithm. The user group includes active exploration types and conservative observation types. Among them, the interaction frequency of the active exploration type user group is higher than that of the conservative observation type user group. Reduce the number of elements required for the interactive projection of the active exploration type user group. Increase the number of elements required for the conservative observation type.
[0037] Since the active exploration type user group has a relatively high desire for active interaction, even if the number of elements in the projection is slightly reduced, its attractiveness can still be maintained. For the conservative observation type user group, increase the number of elements in the projection to improve the attractiveness of the interactive projection and enhance the overall user experience.
[0038] Step S22: Specifically, step S22 is to dynamically adjust the projection content according to the user's body proportion and the projection environment time. By extracting the position and hierarchical information of all elements on the projection interface, constructing a hierarchical relationship diagram between the interface elements, optimizing the layout of the elements on the interface according to the user's body proportion, and using the Canny edge detection algorithm to adjust the element ratio in the projection interface layout to adapt to the user's body proportion. Adjust the color, shape, and gradient of the interface elements according to the projection environment time to better integrate into the projection interface, calculate the color distribution consistency score between the elements and the background, and determine the adjustment range.
[0039] Dynamically adjust the projection content according to the user's body proportion and the projection environment time; adjust the color, shape, and gradient of the interface elements to better integrate into the projection interface; provide personalized interactive experiences for different users.
[0040] Step S3: Achieve the fusion of multimodal information through the multimodal model to generate the projection required by the user.
[0041] Step S4: Transfer the projection generated by the AI unit to the location where the user is through a projection device.
[0042] Step S5: Track the position change of the user in the space according to the historical position of the collected three-dimensional point cloud data, and through data analysis, infer the future movement trajectory of the user, so that the projection interface can move following the user's position change. Determine whether the current movement trajectory of the user is within the preset stay range. If the current movement trajectory exceeds the preset stay range, stop the interactive projection for this user. If the movement trajectory is within the preset stay range, continue the interactive projection for this user.
[0043] Please refer to Figure 3 , an AI interactive projection system based on a three-dimensional lidar, which includes an information unit, an AI unit, and a projection unit.
[0044] The information unit includes a three-dimensional lidar and a directional microphone array, which are used to collect human three-dimensional point cloud data and command voices. The information unit forms a data set with the obtained three-dimensional point cloud data and voice data.
[0045] The AI unit includes a multi-modal model, a single-modal model, a deep neural network model, and a trajectory prediction model. The multi-modal model is used to calculate the total confidence of whether the user makes a projection decision currently and generate a projection. The single-modal model adjusts the corresponding interactive projection according to the user's personalized information and the current environmental state. The deep neural network model is used to divide user groups through a clustering algorithm and adaptively adjust the number of elements in the interactive projections of different user groups. The trajectory prediction model is used to track and predict the movement trajectory of the user in the space and select whether to continue the interactive projection according to the movement trajectory.
[0046] The projection unit includes a projection device and an acoustic-optic device. The projection device is used to output a corresponding interactive projection to the position where the user is currently located according to the AI unit. The acoustic-optic device is used to output lights and audio.
[0047] Preferably, in the deep neural network model, user groups are divided through a clustering algorithm. The user groups include active exploration types and conservative observation types, and the interaction frequency of the active exploration type user group is higher than that of the conservative observation type user group. Reduce the number of elements required for the interactive projection of the active exploration type user group. Increase the number of elements required for the interactive projection of the conservative observation type.
[0048] Preferably, in the multi-modal model, the features of each category module are generated into corresponding feature vectors through independent encoders. The multi-modal model can recognize gestures such as waving, clenching, and drawing circles through the feature vectors of the hand key points extracted, and calculate the corresponding confidence according to the clarity of the gesture actions. The multi-modal model can adjust and calculate the corresponding confidence according to the signal-to-noise ratio of the voice.
[0049] Preferably, in the single-modal model, the projection content is dynamically adjusted according to the user's body proportion and the projection environment time. By extracting the position and hierarchical information of all elements on the projection interface, a hierarchical relationship diagram between the interface elements is constructed, and the layout of the elements on the interface is optimized according to the user's body proportion. Through the Canny edge detection algorithm, the element ratio within the projection interface layout is adjusted to adapt to the user's body proportion. The colors, shapes, and gradients of the interface elements are adjusted according to the projection environment time to better integrate into the projection interface. The color distribution consistency score between the elements and the background is calculated to determine the adjustment amplitude. According to the overall visual effect of the interface and the coordination between the elements and the background, the parameters of the gradient effect are adjusted until the gradient effect is consistent with the background and the overall interface design. The parameters of the gradient effect include color transition, direction, and diffusion range.
[0050] Preferably, in the trajectory prediction model, according to the historical positions of the collected three-dimensional point cloud data, the position changes of the user in the space are tracked, and through data analysis, the future movement trajectory of the user is inferred, so that the projection interface can move following the user's position change. It is judged whether the current movement trajectory of the user is within the preset stay range. If the current movement trajectory exceeds the preset stay range, the interactive projection for this user is stopped. If the movement trajectory is within the preset stay range, the interactive projection for this user continues.
[0051] Specific Embodiment 2: An AI interactive projection system based on a three-dimensional lidar. The same parts as in Specific Embodiment 1 are not described again. The differences are as follows: It further includes a mobile phone terminal. The user connects the mobile phone terminal to the projection device to control the projection device and the sound and light device through the mobile phone terminal.
[0052] The user uses the mobile phone terminal applet to scan the QR code of the on-site device to connect the mobile phone terminal to the on-site device, calls the applet API to call the three-dimensional lidar to collect the user's position data. The mobile phone terminal and the AI unit transmit the positioning and sensor data in real time and receive navigation instructions at the same time. Based on the XR-FRAME capability of the WeChat applet, the recognition and tracking of images are realized, and the rendering of the interactive projection is completed in combination with the AR rendering engine.
[0053] For example, 1. Control the follow spotlights. The sound and light device includes follow spotlight devices. The mobile phone terminal is connected to the on-site follow light devices through the Internet of Things (IoT). The follow spotlights can be controlled in real time through the mobile phone terminal instructions, and the on-site light colors can be switched on the mobile phone terminal;
[0054] 2. AR navigation function. The user selects the point to be reached on the applet, and the projection device automatically projects the corresponding path on the ground to guide the user to the target point;
[0055] 3. Poem recitation function. Users can recite poems on the mobile phone, and the on-site screen will change synchronously according to the user's voice. The system uses speech recognition technology to capture and analyze the user's recitation content in real time, and adjusts the screen content in real time according to information such as the volume, rhythm, and intonation of the recitation to enhance the immersion.
[0056] 4. Element interaction function. Users drag blessing elements such as sky lanterns and fireworks into the air through the mobile phone, and trigger the projection device to display corresponding effects in the real space; users click on the moon icon on the mobile phone, and through the data synchronization mechanism, the moon rises synchronously on the projection side; when users step into the projection "pond", the radar detects the approach of the users, and the pond projection starts to show rippling effects, simulating real water surface fluctuations to increase the sense of interaction; using the XR-FRAME ability of the WeChat mini-program to identify the fish in the projection, when capturing, use the instant messaging function to make the fish disappear in the projection; when the user successfully identifies a certain fish, click the capture button, and the fish will disappear in the projection, and a capture success reward will be displayed on the mobile phone interface.
[0057] By setting the mobile phone, users can directly connect the mobile phone to the projection device. Using the mobile phone can reduce the data collection frequency and accuracy of the 3D lidar, further reduce the operation pressure of the system. Users can directly input commands to obtain the special effects of the interactive projection, realize the reasonable allocation of resources to the users within the space group, reduce the operation cost, and improve the operation efficiency of the system.
Claims
1. An AI interactive projection method based on a 3D lidar, characterized in that: It includes step S1 - step S4; Step S1: Acquisition of human body three - dimensional point cloud data; Specifically, a three - dimensional lidar emits laser beams and receives reflected signals to generate high - precision three - dimensional point cloud data, capturing the position and shape of the human body; Extract key points of the wrist and fingertips of the hand bones to capture the dynamic gestures of the human body; Voice data acquisition and recognition; Specifically, deploy a directional microphone array to collect command voices and suppress on - site noise; Step S2: Compose the three - dimensional point cloud data and voice data obtained in step S1 into a data set and input it into the AI unit. The AI unit obtains the total confidence by weighting the confidence of each modality and makes a decision on whether to generate an interactive projection through multi - modality fusion; If the total confidence reaches the set threshold, the AI unit generates the corresponding interactive projection; Step S3: Achieve the fusion of multi - modality information through a multi - modality model to generate the projection required by the user; Step S4: Transfer the projection generated by the AI unit to the location where the user is through a projection device.
2. The AI interactive projection method based on 3D lidar according to claim 1, wherein: Step S2 also includes step S21; Step S21 is specifically as follows: Collect the movement trajectories and gesture change frequencies of each user, and divide the user group through a clustering algorithm. The user group includes active exploration types and conservative observation types, where the interaction frequency of the active exploration type user group is higher than that of the conservative observation type user group; Reduce the number of elements of the interactive projection required by the active exploration type user group; Increase the number of elements of the interactive projection required by the conservative observation type.
3. An AI interactive projection method based on a 3D lidar according to claim 1, characterized in that: In step S2, each data in the data set is classified, including gestures, sounds, and relative positions; the influence of each category on the decision of generating a projection is evaluated; a total confidence evaluation formula is set where CL is the total confidence, and W i is the weight of the i-th category, and C i is the confidence of the i-th category, and n is the total number of categories participating in the calculation; the AI unit extracts the features of each category module and generates corresponding feature vectors through independent encoders, and obtains the confidence and weight of each category through dynamic calculation; if the total confidence does not reach the set threshold, the AI unit does not generate the interactive projection of this user.
4. The AI interactive projection method based on 3D lidar according to claim 2, wherein: Step S2 also includes step S22; Step S22 is specifically as follows: Dynamically adjust the projection content according to the user's body proportion and the projection environment time; By extracting the position and hierarchical information of all elements on the projection interface, construct a hierarchical relationship diagram between interface elements, optimize the layout of elements on the interface according to the user's body proportion, and adjust the element ratio within the projection interface layout to adapt to the user's body proportion through the Canny edge detection algorithm; Adjust the color, shape, and gradient of the interface elements according to the projection environment time to better integrate into the projection interface, calculate the color distribution consistency score between the elements and the background, and determine the adjustment amplitude.
5. A method for AI interactive projection based on a three-dimensional lidar according to claim 1, characterized in that: It also includes step S5; According to the historical positions of the collected three - dimensional point cloud data, track the position changes of the user in space, and through data analysis, infer the user's future movement trajectory, so that the projection interface can move following the user's position changes; Determine whether the current movement trajectory of the user is within the preset stay range. If the current movement trajectory exceeds the preset stay range, stop the interactive projection for this user; If the movement trajectory is within the preset stay range, continue the interactive projection for this user.
6. An AI interactive projection system based on a 3D lidar, characterized in that: It includes an information unit, an AI unit, and a projection unit; The information unit includes a three - dimensional lidar and a directional microphone array, which are used to collect human body three - dimensional point cloud data and command voices; The information unit composes the obtained three - dimensional point cloud data and voice data into a data set; The AI unit includes a multi - modality model, a single - modality model, a deep neural network model, and a trajectory prediction model; The multi - modality model is used to calculate the total confidence of whether the user makes a projection decision currently and generate a projection; The unimodal model adjusts the corresponding interaction projection according to the user's personalized information and the current environmental state; the deep neural network model is used to divide user groups through a clustering algorithm and adaptively adjust the number of elements in the interaction projections of different user groups; the trajectory prediction model is used to track and predict the movement trajectory of users in space and select whether to continue the interaction projection based on the movement trajectory. The projection unit includes a projection device and an acoustic-optical device; the projection device is used to output the corresponding interaction projection to the user's current location according to the AI unit; the acoustic-optical device is used to output lights and audio.
7. An AI interactive projection system based on a three-dimensional lidar according to claim 6, characterized in that: In the deep neural network model, user groups are divided through a clustering algorithm. User groups include active exploration types and conservative observation types. Among them, the interaction frequency of the active exploration type user group is higher than that of the conservative observation type user group; reduce the number of elements required for the interaction projection of the active exploration type user group; increase the number of elements required for the interaction projection of the conservative observation type.
8. An AI interactive projection system based on a three-dimensional lidar according to claim 6, characterized in that: In the multimodal model, the features of each category module are used to generate corresponding feature vectors through independent encoders; the multimodal model can recognize gestures such as waving, clenching, and drawing circles based on the feature vectors of the hand key points extracted, and calculate the corresponding confidence according to the clarity of the gesture actions; the multimodal model can adjust and calculate the corresponding confidence according to the signal-to-noise ratio of the voice.
9. An AI interactive projection system based on a three-dimensional lidar according to claim 6, characterized in that: In the unimodal model, the projection content is dynamically adjusted according to the user's body proportion and the projection environmental time; by extracting the position and hierarchical information of all elements on the projection interface, a hierarchical relationship diagram between interface elements is constructed, and the layout of elements on the interface is optimized according to the user's body proportion. Through the Canny edge detection algorithm, the element ratio within the projection interface layout is adjusted to adapt to the user's body proportion; the colors, shapes, and gradients of the interface elements are adjusted according to the projection environmental time to better integrate into the projection interface, and the color distribution consistency score between the elements and the background is calculated to determine the adjustment amplitude; according to the overall visual effect of the interface and the coordination between the elements and the background, the parameters of the gradient effect are adjusted until the gradient effect is consistent with the background and the overall interface design. The parameters of the gradient effect include color transition, direction, and diffusion range.
10. An AI interactive projection system based on a 3D lidar according to claim 6, characterized in that: In the trajectory prediction model, based on the historical positions of the collected three-dimensional point cloud data, the position changes of users in space are tracked, and through data analysis, the future movement trajectories of users are inferred, enabling the projection interface to move following the user's position changes; determine whether the current user's movement trajectory is within the preset stay range. If the current movement trajectory exceeds the preset stay range, stop the interaction projection for this user; if the movement trajectory is within the preset stay range, continue the interaction projection for this user.