Stereographic projection robot system, control method thereof, storage medium and program product

By designing a stereo projection robot system, using multimodal interaction and environmental feature recognition technology, the problem that traditional systems cannot achieve stereo projection and adaptive interaction is solved, and high-definition three-dimensional projection and convenient multimodal interaction are achieved.

CN120219587APending Publication Date: 2025-06-27NANJING SMARTVISION ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510334757.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional projection robot systems cannot realize multimodal interaction, adaptive complex scenes and stereo projections, resulting in unstable clarity of projection content and the inability to generate scene-based stereo projection content in real time.

Method used

A stereoscopic projection robot system is designed, including a robot main mechanical subsystem, a multimodal interaction subsystem, a drive engine subsystem and a stereoscopic projection subsystem. The system obtains multimodal input information, identifies environmental features, and generates a solid projection screen based on the projection content, and adjusts the projection display parameters.

Benefits of technology

It realizes multimodal interaction mode, adapts to complex scenes, and can generate realistic three-dimensional images, improving the practicality and adaptability of the robot system, making communication between users and robots more natural and convenient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219587A_ABST
    Figure CN120219587A_ABST
Patent Text Reader

Abstract

The invention relates to a stereoscopic projection robot system, a control method thereof, a storage medium and a program product. The system comprises a robot main body mechanical subsystem which is used for carrying out initialization setting on a multi-mode interaction subsystem, a driving engine subsystem and a stereo projection subsystem; the multi-modal interaction subsystem is used for acquiring surrounding environment information and multi-modal input information of a user, and the multi-modal input information comprises at least one of display screen operation information, voice instructions, gesture actions and expression information; the driving engine subsystem is used for identifying environment characteristics of the surrounding environment information and determining projection content according to the multi-modal input information; and the stereoscopic projection subsystem is integrated in the robot main body in the robot main body mechanical subsystem and is used for generating a stereoscopic projection picture according to the projection content and adjusting projection display parameters according to the environment characteristics. By adopting the system, a multi-mode interaction mode can be realized, the system is adaptive to complex scenes, and stereoscopic projection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and particularly to a stereoscopic projection robot system, a control method thereof, a storage medium, and a program product. Background Art

[0002] In the traditional projection robot vision display mode, the main components are LCD (Liquid Crystal Display) or OLED (Organic Light Emitting Diode) straight panels, and it relies on voice or touch commands, with a single interaction mode. It can only project preset content and cannot adaptively adjust, resulting in unstable clarity of the projected content. Moreover, most of the projected content is 2D images and cannot generate scene-based stereoscopic projection content in real time. Summary of the Invention

[0003] Based on this, in order to solve the above technical problems, it is necessary to provide a stereoscopic projection robot system, a control method thereof, a storage medium, and a program product that can achieve multi-modal interaction modes, adapt to complex scenarios, and realize stereoscopic projection.

[0004] In a first aspect, the present application provides a stereoscopic projection robot system, which includes:

[0005] A robot main body mechanical subsystem for initializing the multi-modal interaction subsystem, the drive engine subsystem, and the stereoscopic projection subsystem;

[0006] A multi-modal interaction subsystem for obtaining surrounding environment information and the user's multi-modal input information, where the multi-modal input information includes at least one of display screen operation information, voice commands, gesture actions, and facial expression information;

[0007] A drive engine subsystem for identifying the environmental characteristics of the surrounding environment information and determining the projection content according to the multi-modal input information;

[0008] A stereoscopic projection subsystem integrated in the robot main body of the robot main body mechanical subsystem for generating a stereoscopic projection image according to the projection content and adjusting the projection display parameters according to the environmental characteristics.

[0009] In one embodiment, the stereoscopic projection subsystem includes:

[0010] A microdisplay chip for receiving the projection content, converting the projection content into a projectable optical image, and sending it to the optical component;

[0011] A light source component for emitting light of a preset color;

[0012] A spectroscopic device for decomposing light of a preset color into three-color light and projecting it onto an optical component;

[0013] An optical component for modulating the three-color light through a projectable optical image to generate a stereoscopic projection image and project it.

[0014] In one embodiment, the stereoscopic projection subsystem further includes:

[0015] An adaptive adjustment component for invoking a neural network model to identify the content of the stereoscopic projection image and adjusting the projection display parameters according to the image content and environmental characteristics.

[0016] In one embodiment, there are two stereoscopic projection subsystems, which are respectively used to generate a left-eye projection image and a right-eye projection image according to the projection content, and match and fuse the left-eye projection image and the right-eye projection image to generate a stereoscopic projection image.

[0017] In one embodiment, the robot main body mechanical subsystem is further used to construct a three-dimensional environment map and perform three-dimensional scanning and recognition of objects, identify obstacle information according to the three-dimensional environment map, and perform obstacle avoidance according to the obstacle information; perform object recognition and identification according to the three-dimensional object scanning technology;

[0018] The stereoscopic projection subsystem is further used to realize the switching between the planar and stereoscopic perspectives according to the three-dimensional environment map, so as to realize the switching between the planar projection image and the stereoscopic projection image.

[0019] In one embodiment, the stereoscopic projection subsystem includes a spatial light modulator.

[0020] In one embodiment, the multi-modal interaction subsystem is further used to convert the voice command into text information when obtaining the user's voice command, and send the text information to the drive engine subsystem.

[0021] In a second aspect, the present application further provides a control method for a stereoscopic projection robot system, including:

[0022] Controlling the robot main body mechanical subsystem to perform initialization settings on the multi-modal interaction subsystem, the drive engine subsystem, and the stereoscopic projection subsystem;

[0023] Controlling the multi-modal interaction subsystem to obtain the surrounding environment information and the user's multi-modal input information, where the multi-modal input information includes at least one of display screen operation information, voice commands, gesture actions, and facial expression information;

[0024] Controlling the drive engine subsystem to identify the environmental characteristics of the surrounding environment information and determine the projection content according to the multi-modal input information;

[0025] Control the stereoscopic projection subsystem to generate a stereoscopic projection screen according to the projection content, and adjust the projection display parameters according to the environmental characteristics.

[0026] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0027] Control the robot main body mechanical subsystem to perform initialization settings on the multimodal interaction subsystem, the driving engine subsystem, and the stereoscopic projection subsystem;

[0028] Control the multimodal interaction subsystem to obtain the surrounding environment information and the user's multimodal input information, where the multimodal input information includes at least one of display screen operation information, voice commands, gesture actions, and facial expression information;

[0029] Control the driving engine subsystem to identify the environmental characteristics of the surrounding environment information, and determine the projection content according to the multimodal input information;

[0030] Control the stereoscopic projection subsystem to generate a stereoscopic projection screen according to the projection content, and adjust the projection display parameters according to the environmental characteristics.

[0031] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0032] Control the robot main body mechanical subsystem to perform initialization settings on the multimodal interaction subsystem, the driving engine subsystem, and the stereoscopic projection subsystem;

[0033] Control the multimodal interaction subsystem to obtain the surrounding environment information and the user's multimodal input information, where the multimodal input information includes at least one of display screen operation information, voice commands, gesture actions, and facial expression information;

[0034] Control the driving engine subsystem to identify the environmental characteristics of the surrounding environment information, and determine the projection content according to the multimodal input information;

[0035] Control the stereoscopic projection subsystem to generate a stereoscopic projection screen according to the projection content, and adjust the projection display parameters according to the environmental characteristics.

[0036] In a fifth aspect, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0037] Control the robot main body mechanical subsystem to perform initialization settings on the multimodal interaction subsystem, the driving engine subsystem, and the stereoscopic projection subsystem;

[0038] The control multi-modal interaction subsystem acquires the surrounding environment information and the user's multi-modal input information, where the multi-modal input information includes at least one of display screen operation information, voice commands, gesture actions, and facial expression information;

[0039] The control driving engine subsystem identifies the environmental features of the surrounding environment information and determines the projection content according to the multi-modal input information;

[0040] The control stereoscopic projection subsystem generates a stereoscopic projection image according to the projection content and adjusts the projection display parameters according to the environmental features.

[0041] The above-mentioned stereoscopic projection robot system, its control method, storage medium, and program product achieve multiple interaction methods by obtaining multi-modal input information, making the communication between the user and the robot more natural and convenient; being able to present realistic three-dimensional images; and enabling the robot to automatically adjust the projection content and display method according to different environments and user needs, improving the practicability and adaptability of the robot. Description of the Drawings

[0042] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0043] Figure 1 It is a schematic structural diagram of a stereoscopic projection robot system in an embodiment;

[0044] Figure 2 It is a schematic flowchart of the control method of a stereoscopic projection robot system in an embodiment;

[0045] Figure 3 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments

[0046] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0047] In an exemplary embodiment, as Figure 1 shown, a stereoscopic projection robot system is provided. This system integrates intelligent projection technology, multi-modal interaction technology, and autonomous environment perception, and is applicable to scenarios such as museum guided tours, educational services, surgical navigation, and emotional companionship that require high-precision spatial interaction.

[0048] The stereoscopic projection robot system includes a robot main body mechanical subsystem, a multimodal interaction subsystem, a driving engine subsystem, and a stereoscopic projection subsystem. Among them:

[0049] The robot main body mechanical subsystem is used to perform initialization settings for the multimodal interaction subsystem, the driving engine subsystem, and the stereoscopic projection subsystem;

[0050] The multimodal interaction subsystem is used to obtain surrounding environment information and the user's multimodal input information, and the multimodal input information includes at least one of display screen operation information, voice commands, gesture actions, and expression information;

[0051] The driving engine subsystem is used to identify the environmental characteristics of the surrounding environment information and determine the projection content according to the multimodal input information;

[0052] The stereoscopic projection subsystem is integrated into the robot main body of the robot main body mechanical subsystem, and is used to generate a stereoscopic projection image according to the projection content and adjust the projection display parameters according to the environmental characteristics.

[0053] Among them, multimodal information refers to various forms of information input by the user.

[0054] Optionally, the robot main body mechanical subsystem includes a robot body, a movable robotic arm, a mobile chassis with a driving device, a main control system and a power module arranged inside the robot body, and a sensor sequence for completing actions such as machine power supply, walking, grasping, and obstacle avoidance. The mobile chassis enables the robot body to move flexibly in a complex environment, and the robotic arm can rotate and extend at various angles to complete specific operation tasks, such as stretching and unfolding the projection screen with both arms. The sensor sequence is used for obstacle detection to achieve obstacle avoidance.

[0055] After the power is turned on, the power module inside the robot body supplies power to each component of the robot main body mechanical subsystem. After the main control system is started, it performs initialization settings for the multimodal interaction subsystem, the driving engine subsystem, and the stereoscopic projection subsystem. The driving engine subsystem loads a pre-trained model and data, and this model is a neural network model for multimodal input information interaction. Perform light source calibration and image parameter settings for the stereoscopic projection subsystem. Specifically, based on a pre-trained color coordinate calculation model (such as trained through multiple sets of projection data), input the real-time collected color data into the model, output the difference between the current color coordinate and the target value, and dynamically adjust the intensity ratio of red, green, and blue three-color lights. Image parameters can include brightness, contrast, color depth, color temperature, etc. Detect whether each sensor and input device in the modal interaction subsystem is working properly. Specifically, the main control system reads the status registers of each sensor and device to check whether the status is normal or abnormal.

[0056] The multimodal interaction subsystem includes a multimodal perception subsystem and a human-computer interaction subsystem. The multimodal perception subsystem is used to obtain the multimodal input information of the user, and the human-computer interaction subsystem is used to interact with the user to achieve outputs such as voice. The multimodal perception subsystem includes a camera, a microphone, a voice recognition component, a gesture recognition sensor disposed on the robot body, and a touch display screen disposed on the surface of the robot body. The camera is used to obtain the surrounding environment information in real time and can also be used to recognize the user's facial expression information. Further, the camera can be a vision camera (RGB-D camera). The multimodal perception subsystem realizes the full-dimensional acquisition of environmental information by integrating vision (RGB-D camera), lidar, microphone, touch display screen, etc.

[0057] The microphone is used to listen to the surrounding sounds in real time, such as the user's voice commands. The voice recognition component is used to convert the voice commands into text information and send it to the driving engine subsystem. The gesture recognition sensor continuously monitors the user's gesture actions and can recognize various gesture actions of the user, such as waving, making a fist, etc. The touch display screen waits for the user's touch operation and is used to receive the user's display screen operation information. The display screen operation information can include operation information for adjusting the movie playback progress, selecting projection content, adjusting the function settings of the robot, etc. The human-computer interaction subsystem includes a voice interaction component, and the voice interaction component realizes the voice dialogue and communication between the user and the robot. The user controls the actions and projection display of the robot through voice commands.

[0058] The driving engine subsystem is an AI driving engine subsystem. The driving engine subsystem is used to call the neural network model for multimodal input information interaction, integrate the multimodal input information to determine the projection content. The neural network model for multimodal input information interaction is a model based on deep learning algorithms, which can realize real-time environment adaptation and adaptive learning, perform local parameter adjustment through incremental learning and lightweight network branches, and combine the attention mechanism to dynamically allocate computing resources. It can still maintain performance in few-shot scenarios, improve performance on the premise of few samples, reduce latency, and improve the real-time display effect of the projection. Specifically, the timestamp alignment algorithm (DTW dynamic time warping) is used to solve the timing difference between voice and gesture. The neural network model for multimodal input information interaction is adopted to extract the modal features of each modal input information respectively, dynamically allocate weights through the attention mechanism, synthesize the multimodal output results, and determine the projection content. Further, the driving engine subsystem can intelligently generate corresponding projection content according to different scenarios and user needs, and control the stereoscopic projection system for display. For example, in an educational scenario, when it is recognized that the user asks a question about a historical event, generate the stereoscopic projection content of the relevant historical scenario and dynamically display it to the customer; in an entertainment scenario, generate the corresponding game stereoscopic picture according to the game type selected by the user.

[0059] Optionally, the driving engine subsystem can also recognize gesture actions and perform corresponding operations according to the recognized gesture actions, such as switching the projection screen, pausing or playing the projection content, etc.

[0060] The driving engine subsystem is also used to recognize the environmental characteristics of the surrounding environment information, such as recognizing the size, layout, lighting conditions, etc. of the room.

[0061] The stereoscopic projection subsystem is used for stereoscopic projection screen generation and display: generating a stereoscopic projection screen according to the projection content, and adjusting the projection display parameters according to the environmental characteristics, such as projection brightness, contrast, etc. The projection display parameters are image parameters. During the projection process, the projection brightness, contrast and other parameters are adjusted in real time according to the environmental lighting conditions to ensure the best viewing effect.

[0062] The stereoscopic projection robot system can also achieve interactive response and continuous interaction: when the user operates through the touch display screen, such as adjusting the movie playback progress, the touch display screen transmits the operation information to the main control system, and the main control system controls the stereoscopic projection subsystem to pause or continue playing the movie and adjusts the projection screen accordingly. If the user makes a gesture action, such as waving to switch to the next movie clip, the gesture recognition sensor sends the gesture information to the driving engine subsystem, and the driving engine subsystem performs the corresponding operation and updates the projection content. During the entire interaction process, the voice interaction component remains active, receiving the user's voice commands at any time and making responses to achieve continuous and smooth human-machine interaction.

[0063] Optionally, as Figure 1 shown, the stereoscopic projection robot system further includes a multi-robot cooperation subsystem, and the multi-robot cooperation subsystem realizes the collaborative work of multiple subsystems through high-speed network communication and intelligent networking.

[0064] In this embodiment, the stereoscopic projection robot system has the following technical effects:

[0065] 1) Enhance the human-machine interaction experience: By obtaining multi-modal input information, various interaction methods are realized, such as voice, touch, gesture, expression, etc., making the communication between the user and the robot more natural and convenient;

[0066] 2) Immersive stereoscopic projection display: The stereoscopic projection function can present realistic three-dimensional images, bringing an immersive visual experience to users, and has broad application prospects in multiple fields such as education, entertainment, and commerce. For example, in the education field, complex scientific knowledge and historical scenes can be displayed more vividly, and in the entertainment field, a more immersive game and movie viewing experience can be provided;

[0067] 3) Intelligent environment adaptation: Enables the robot to automatically adjust the projection content and display method according to different environments and user needs, improving the practicality and adaptability of the robot.

[0068] In an exemplary embodiment, the stereoscopic projection subsystem includes:

[0069] A microdisplay chip, configured to receive projection content, convert the projection content into a projectable optical image, and send it to an optical component;

[0070] A light source component, configured to emit light of a preset color;

[0071] A beam splitting device, configured to split the light of the preset color into three-color light and project it onto the optical component;

[0072] An optical component, configured to modulate the three-color light with the projectable optical image to generate a stereoscopic projection image and project it.

[0073] Optionally, the stereoscopic projection subsystem includes a microdisplay chip, a light source component, a beam splitting device, and an optical component. The light source component supports high-power and high-brightness RGB light sources. The optical component includes an LCOS liquid crystal light valve, an optical path system, and a lens assembly.

[0074] The driving engine subsystem is configured to determine projection content, such as video content or expression images. The projection content is transmitted to the microdisplay chip of the stereoscopic projection subsystem. The microdisplay chip decodes, sharpens, color-enhances, etc. the projection content, converts it into a format recognizable by the optical component, i.e., a projectable optical image, and sends it to the optical component. Subsequently, the light source component emits white light, which is split into red, green, and blue three-color lights by the beam splitting device and projected onto the optical component respectively. Finally, after the three-color lights are modulated by the LCOS liquid crystal light valve of the optical component, the light is focused and magnified through the optical path system and the lens assembly and projected onto the screen or the wall. The LCOS liquid crystal light valve controls the arrangement of liquid crystal molecules through an electric field to modulate the reflection characteristics of the three-color lights, thereby achieving the purpose of controlling the three-color lights. The optical path system and the lens assembly are responsible for controlling and correcting the light path to achieve high-quality imaging. When the stereoscopic projection subsystem is provided with a projection screen, the generated stereoscopic projection image is projected onto the projection screen. When the stereoscopic projection subsystem does not involve a projection screen, the generated stereoscopic projection image can be projected onto the wall.

[0075] The above stereoscopic projection subsystem can project a stereoscopic projection image of a robot's facial expression and high-quality and high-definition by converting the projection content into a projectable optical image and modulating the three-color light according to the projectable optical image.

[0076] In an exemplary embodiment, the stereoscopic projection subsystem further includes:

[0077] An adaptive adjustment component, configured to call a neural network model to identify the content of the stereoscopic projection image and adjust the projection display parameters according to the content of the image and the environmental characteristics.

[0078] Optionally, the adaptive adjustment component is used for projection correction based on a deep learning-based adaptive projection calibration algorithm. The deep learning-based adaptive projection calibration algorithm refers to an optimization method that uses a deep learning module to automatically adjust projection display parameters. Specifically, using a deep learning model, the content of the stereoscopic projection screen is identified, and combined with environmental features, the optimal projection display parameters, such as the optimal brightness, contrast, and depth of field, are predicted to improve the projection display effect.

[0079] Traditional projection adjustment methods usually rely on the internal parameters of the camera and manually input data, which may lead to poor correction effects or complex operations in some cases. The deep learning-based adaptive projection calibration algorithm can achieve high-precision and high-efficiency projection correction without relying on prior parameters, and can adjust the projection display parameters in real time according to environmental conditions.

[0080] In an exemplary embodiment, there are two stereoscopic projection subsystems. The two stereoscopic projection subsystems are respectively used to generate a left-eye projection image and a right-eye projection image according to the projection content, and match and fuse the left-eye projection image and the right-eye projection image to generate a stereoscopic projection screen.

[0081] The stereoscopic projection robot system can use various methods through the intelligent network module to generate a stereoscopic projection screen in real time and project it for the user.

[0082] For a single stereoscopic projection robot system, it can be networked and fused with traditional display devices, such as TVs, projectors, etc. through the intelligent network module to generate 3D stereoscopic picture content in real time. Among them, the intelligent network module can be a built-in WIFI / Bluetooth wireless module.

[0083] A single stereoscopic projection robot system can project the projection content onto a special screen to form a 3D stereoscopic effect.

[0084] A single stereoscopic projection robot system can have two stereoscopic projection subsystems built in, which work together to match and fuse the respectively generated left-eye and right-eye images to generate picture content with a 3D stereoscopic effect in real time.

[0085] Optionally, a single stereoscopic projection robot system can set a spatial light modulator in the stereoscopic projection subsystem to directly project a 3D stereoscopic screen.

[0086] When at least two stereoscopic projection robot systems perform projection display, the stereoscopic projection subsystems of at least two stereoscopic projection robot systems can be intelligently networked through a high-speed network protocol, and the left-eye and right-eye images are matched and fused through the naked-eye 3D technology to generate a 3D stereoscopic projection screen in real time and perform projection display.

[0087] In this embodiment, by setting multiple projection methods, it is convenient for the stereoscopic projection robot system to perform projection displays in various service scenarios.

[0088] In an exemplary embodiment, the robot body mechanical subsystem is further configured to construct a three-dimensional environmental map and perform three-dimensional scanning and recognition of items, identify obstacle information based on the three-dimensional environmental map, and perform obstacle avoidance according to the obstacle information; perform item recognition and identification according to the three-dimensional item scanning technology.

[0089] The stereoscopic projection subsystem is further configured to switch between planar and stereoscopic perspectives according to the three-dimensional environmental map, so as to realize the switching between the planar projection screen and the stereoscopic projection screen.

[0090] Optionally, the sensor sequence of the robot body mechanical subsystem includes a structured light device, a millimeter-wave radar, and a ToF camera. The structured light device projects a known encoded pattern (such as stripes, dot matrices, etc.) onto the object surface, and the ToF camera captures the deformed pattern, and calculates the depth information in combination with the triangulation principle. At the same time, the millimeter-wave radar emits millimeter-wave electromagnetic signals of 30 - 300 GHz, receives the reflected echoes and analyzes the time difference, frequency change (Doppler effect), and beam angle to construct the environmental information. And, based on the Time-of-Flight method, the distance is calculated by measuring the round-trip time of the light pulse or modulated light. That is, the structured light provides local high-precision modeling, the millimeter-wave radar is responsible for large-range obstacle detection, and the ToF camera supplements dynamic object tracking to realize the modeling of the three-dimensional environmental map. The robot body mechanical subsystem is also used to perform item recognition and identification using the three-dimensional item scanning technology during the process of constructing the three-dimensional environmental map to accurately identify obstacles and items in the environment.

[0091] Constructing a three-dimensional environmental map can effectively identify information such as the height and slope of obstacles, providing auxiliary information for robot navigation and obstacle avoidance. The three-dimensional map supports multi-perspective switching and two-dimensional and three-dimensional linkage display, and the intelligent projection device can realize seamless switching between planar and stereoscopic perspectives, thereby improving the display effect of the projection.

[0092] In an exemplary embodiment, as Figure 2 shown, a control method for a stereoscopic projection robot system is provided. Taking the application of this method to the stereoscopic projection robot system in Figure 1 as an example for illustration, it includes the following steps 202 to step 208. Among them:

[0093] Step 202, controlling the robot body mechanical subsystem to perform initialization settings on the multi-modal interaction subsystem, the drive engine subsystem, and the stereoscopic projection subsystem.

[0094] Step 204, control the multimodal interaction subsystem to obtain the surrounding environment information and the user's multimodal input information, where the multimodal input information includes at least one of display screen operation information, voice commands, gesture actions, and facial expression information.

[0095] Step 206, control the driving engine subsystem to identify the environmental characteristics of the surrounding environment information, and determine the projection content according to the multimodal input information.

[0096] Step 208, control the stereoscopic projection subsystem to generate a stereoscopic projection image according to the projection content, and adjust the projection display parameters according to the environmental characteristics.

[0097] In the control method of the above stereoscopic projection robot system, by obtaining multimodal input information, multiple interaction methods such as voice, touch, gesture, and facial expression are realized, making the communication between the user and the robot more natural and convenient. By identifying the environmental characteristics of the surrounding environment information, determining the projection content according to the multimodal input information, generating a stereoscopic projection image according to the projection content, and adjusting the projection display parameters according to the environmental characteristics, the projection content and display method can be automatically adjusted according to different environments and user needs, improving the practicability and adaptability of the robot system.

[0098] In an exemplary embodiment, the stereoscopic projection subsystem includes a microdisplay chip, a light source component, a beam splitter device, and an optical component; Step 208, controlling the stereoscopic projection subsystem to generate a stereoscopic projection image according to the projection content includes:

[0099] Control the microdisplay chip to receive the projection content, convert the projection content into a projectable optical image, and send it to the optical component; control the light source component to emit light of a preset color; control the beam splitter device to decompose the light of the preset color into three-color light and project it onto the optical component; control the optical component to modulate the three-color light through the projectable optical image to generate a stereoscopic projection image and project it.

[0100] In an exemplary embodiment, the stereoscopic projection subsystem further includes: an adaptive adjustment component; the method further includes: calling a neural network model to identify the content of the stereoscopic projection image, and adjusting the projection display parameters according to the content of the image and the environmental characteristics.

[0101] In an exemplary embodiment, there are two stereoscopic projection subsystems. Step 208, controlling the stereoscopic projection subsystem to generate a stereoscopic projection image according to the projection content includes: controlling the two stereoscopic projection subsystems to generate a left-eye projection image and a right-eye projection image according to the projection content respectively, and matching and fusing the left-eye projection image and the right-eye projection image to generate a stereoscopic projection image.

[0102] In an exemplary embodiment, the method further includes:

[0103] The mechanical subsystem of the control robot body constructs a three-dimensional environmental map and three-dimensional scanning and recognition of objects, identifies obstacle information based on the three-dimensional environmental map, and performs obstacle avoidance according to the obstacle information; according to the three-dimensional object scanning technology, object recognition and identification are performed.

[0104] The control stereoscopic projection subsystem realizes the switching between the planar and stereoscopic perspectives according to the three-dimensional environmental map, thereby realizing the switching between the planar projection screen and the stereoscopic projection screen.

[0105] In an exemplary embodiment, the method further includes: when the multimodal interaction subsystem is further configured to obtain a voice command of a user, converting the voice command into text information and sending the text information to the driving engine subsystem.

[0106] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages, and these steps or stages do not necessarily have to be executed at the same time, but can be executed at different times, and the execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.

[0107] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 3As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a control method for a stereoscopic projection robot system. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0108] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0109] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0110] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0111] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0112] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0113] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0114] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0115] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A stereoscopic projection robot system, characterized in that: The system comprises: The robot main body mechanical subsystem is used to initialize the multimodal interaction subsystem, the drive engine subsystem, and the stereoscopic projection subsystem; A multimodal interaction subsystem, used to obtain surrounding environment information and multimodal input information of the user, wherein the multimodal input information includes at least one of display screen operation information, voice instructions, gesture actions, and expression information; A driving engine subsystem, configured to identify environmental features of the surrounding environment information and determine projection content according to the multimodal input information; The stereoscopic projection subsystem is integrated into the robot body in the robot body mechanical subsystem, and is used to generate a stereoscopic projection picture according to the projection content and adjust the projection display parameters according to the environmental characteristics.

2. The stereoscopic projection robot system according to claim 1, characterized in that: The stereoscopic projection subsystem comprises: A micro display chip, used for receiving the projection content, converting the projection content into a projectable optical image, and sending it to the optical component; A light source assembly, used for emitting light of a preset color; A light splitting device, used for splitting the light of the preset color into three-color lights and projecting them to the optical component; The optical component is used to modulate the three-color light through the projectable optical image to generate a three-dimensional projection picture for projection.

3. The stereoscopic projection robot system according to claim 2, characterized in that: The stereoscopic projection subsystem also includes: The adaptive adjustment component is used to call the neural network model to identify the picture content of the stereoscopic projection picture and adjust the projection display parameters according to the picture content and the environmental characteristics.

4. The stereoscopic projection robot system according to claim 1, characterized in that: The stereoscopic projection subsystems are provided with two, and the two stereoscopic projection subsystems are respectively used to generate a left-eye projection image and a right-eye projection image according to the projection content, and to match and fuse the left-eye projection image and the right-eye projection image to generate a stereoscopic projection picture.

5. The stereoscopic projection robot system according to claim 1, characterized in that: The robot main body mechanical subsystem is also used to construct a three-dimensional environment map and three-dimensional scanning and identification of objects, identify obstacle information according to the three-dimensional environment map, and avoid obstacles according to the obstacle information; and identify and recognize objects according to the three-dimensional scanning technology of objects; The stereoscopic projection subsystem is also used to switch between plane and stereoscopic viewing angles according to the three-dimensional environment map, thereby switching between plane projection images and stereoscopic projection images.

6. The stereoscopic projection robot system according to claim 1, characterized in that: The stereoscopic projection subsystem includes a spatial light modulator.

7. The stereoscopic projection robot system according to claim 1, characterized in that: The multimodal interaction subsystem is also used to convert the voice instructions into text information when obtaining the user's voice instructions, and send the text information to the driving engine subsystem.

8. A control method for a stereoscopic projection robot system, characterized in that: The method comprises: Controlling the robot main body mechanical subsystem to initialize the multimodal interaction subsystem, the driving engine subsystem and the stereoscopic projection subsystem; Controlling the multimodal interaction subsystem to obtain surrounding environment information and multimodal input information of the user, wherein the multimodal input information includes at least one of display screen operation information, voice command, gesture action, and expression information; Controlling the driving engine subsystem to identify environmental features of the surrounding environment information, and determining projection content according to the multimodal input information; The stereoscopic projection subsystem is controlled to generate a stereoscopic projection picture according to the projection content, and the projection display parameters are adjusted according to the environmental characteristics.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method described in claim 8 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in claim 8 are implemented.