Data recommendation system based on display screen
By combining a transparent OLED display with a multi-sensor module, a thin and transparent display and proactive content recommendation have been achieved for smart terminal devices, solving the problems of bulky devices, low screen-to-body ratio, and limited interaction methods, thus improving the user experience.
Patent Information
- Application Number
- CN202511591129.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-02-03
AI Technical Summary
Existing smart terminal devices suffer from problems such as bulky design, low screen-to-body ratio, inability to display transparently, lack of depth perception capabilities, limited interaction methods, and inability to provide proactive services.
It employs a transparent OLED display and multiple sensor modules working together, combined with the processor's scene perception and intelligent analysis capabilities, to achieve proactive content recommendation based on user status and environmental context, supporting virtual-real fusion interaction and multimodal emotional interaction in augmented reality scenarios.
It achieves device thinness and lightness, transparent display, stable alignment of virtual information with real-world scenes, provides natural and intuitive augmented reality interaction and multimodal emotion perception, and enhances user experience.
Smart Images

Figure CN121455439A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information recommendation, in particular to a data recommendation system based on a display screen. BACKGROUND
[0002] At present, the existing intelligent terminal represented by "girlfriend machine" generally adopts a traditional LCD display scheme. Since it must rely on a backlight module, it leads to a thick and heavy overall device, a low screen ratio, and a physical structure that cannot realize transparent display, resulting in a visual split between the digital interface and the real environment. In terms of function, most of such devices continue the design idea of "tablet magnification", and the functions are concentrated in application layer calling, lacking deep perception ability of user state and surrounding environment, and mainly relying on touch and voice for interaction, which cannot provide active services according to scene semantics.
[0003] Although transparent OLED technology provides a hardware possibility for device transparency, the device simply carrying the screen still has obvious limitations at present: on the one hand, the system cannot accurately understand the environmental semantics and user intent, and it is difficult to realize stable alignment and perspective matching of virtual information and real scene; on the other hand, the interaction dimension is single, lacks natural and intuitive out-of-reach operation ability, and is weak in multi-modal emotional perception and feedback, resulting in lack of emotional connection and personalized response in the human-computer interaction process. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a data recommendation system based on a display screen to overcome at least one of the above-mentioned defects.
[0005] In a first aspect, the embodiments of the present application provide a data recommendation system based on a display screen, which comprises: a transparent OLED display screen; a multi-sensor module comprising a non-vision sensor for non-contact perception of a user's presence state and a vision sensor for collecting environmental images; a processor configured to: identify a user's presence state and an environmental scene type in which the device is located according to data of the non-vision sensor and data of the vision sensor; and actively display recommended content associated with the current scene on the transparent OLED display screen based on the user's presence state and the environmental scene type.
[0006] In an optional embodiment of the present application, the non-vision sensor comprises a millimeter wave radar and a pulse ultra-wideband chip, wherein the millimeter wave radar is arranged in the shell of the system and is used to emit and receive electromagnetic waves to non-contact perceive the chest movement and breathing frequency of the user; the pulse ultra-wideband chip is arranged in the shell of the system and is used to non-contact perceive the presence state and motion trajectory of the user by measuring the time of flight of the pulse signal.
[0007] In an optional embodiment of the present application, the visual sensor comprises a color camera and a depth camera, wherein the color camera is arranged on the shell facing the back of the transparent OLED display screen, and is configured to collect environmental images to identify scene types and objects; and the depth camera is arranged around the color camera, and is configured to obtain depth information of the real scene.
[0008] In an optional embodiment of the present application, the system further comprises a mobile chassis, wherein the mobile chassis is arranged at the bottom of the system, and is integrated with an encoder and an inertial measurement unit, the inertial measurement unit is in communication connection with the processor, and is configured to detect position changes and moving speeds of the system.
[0009] In an optional embodiment of the present application, the processor is further configured to: obtain wheel hub rotation data of the mobile chassis through the encoder, combine three-axis acceleration and angular velocity data obtained by the inertial measurement unit, calculate displacement vectors and rotation angles of the system; determine a current pose of the transparent OLED display screen relative to a real scene coordinate system through coordinate transformation based on the displacement vectors and rotation angles, the current pose comprising position coordinates and an orientation angle; calculate field of view angle changes and perspective projection matrices of the transparent OLED display screen according to the current pose; and adjust display coordinates, size ratios and perspective relationships of recommended content displayed on the transparent OLED display screen based on the perspective projection matrices, so that the recommended content is spatially registered with corresponding target objects in the real scene.
[0010] In an optional embodiment of the present application, the processor is further configured to: determine target anchor positions of virtual content in the real scene based on depth information of each object in the real scene obtained by the depth camera; calculate initial display parameters of the virtual content on the display screen according to the relative spatial relationship between the target anchor positions and the transparent OLED display screen; dynamically correct the initial display parameters in real time in combination with inertial measurement unit data from the mobile chassis; and perform perspective matching fusion display of the virtual content and the real scene according to the corrected display parameters.
[0011] In an optional embodiment of the present application, the visual sensor further comprises a gesture recognition camera, wherein the processor is further configured to: collect image sequences of a user's hand through the gesture recognition camera; extract hand key point spatial coordinates and motion trajectories from the image sequences; match the hand key point spatial coordinates and motion trajectories with a pre-defined gesture instruction library to identify the user's operation intention; and generate control instructions for the fusion display of the virtual content and perform corresponding operations according to the identified operation intention.
[0012] In an optional embodiment of the present application, the processor is further configured to: pre-process the environment image collected by the visual sensor, and extract a scene visual feature vector from the pre-processed environment image; perform similarity calculation on the scene visual feature vector and standard scene features in a preset scene database; and determine, according to the similarity calculation result, a scene type with the highest matching degree to the current environment image as the recognition result of the environment scene type.
[0013] In an optional embodiment of the present application, the processor is further configured to: collect a user voice signal through a microphone array, and extract an acoustic feature parameter of the voice; input the acoustic feature parameter into a voice emotion classification model to obtain a first emotion recognition result; collect a user face image through a visual sensor, and extract a face feature point motion parameter; input the face feature point motion parameter into an expression recognition model to obtain a second emotion recognition result; fuse the first emotion recognition result and the second emotion recognition result to determine a final user emotion state; and according to the final user emotion state, select matched recommended content from a content library for display.
[0014] In an optional embodiment of the present application, the processor is further configured to: continuously collect a user chest movement signal through the millimeter wave radar, and continuously collect a user body movement signal through the pulse ultra-wideband chip; perform filtering processing on the chest movement signal and the body movement signal respectively, extract a breathing frequency from the chest movement signal, and extract a body movement feature from the body movement signal; analyze the breathing frequency and the amplitude of the body movement feature to determine whether the user is in a stationary sleep state; when the breathing frequency meets a preset condition and the body movement feature is lower than a threshold value in a plurality of continuous periods, determine that the user is in a sleep state; trigger a system control instruction to switch the system to a low-power consumption mode or turn off the display function of the transparent OLED display screen.
[0015] The data recommendation system based on a display screen provided by the embodiments of the present application comprises: a transparent OLED display screen; a multi-sensor module comprising a non-visual sensor for contactless sensing of a user presence state, and a visual sensor for collecting an environment image; and a processor configured to: identify a user in-place state and an environment scene type in which a device is located according to data of the non-visual sensor and data of the visual sensor; and based on the user in-place state and the environment scene type, actively display recommended content associated with a current scene on the transparent OLED display screen. Through the present application, active content recommendation based on scene perception is realized, and user experience is improved.
[0016] In order to make the above objectives, features and advantages of the present application more apparent and easy to understand, a preferred embodiment is described below in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be considered as a limitation to the scope. For those of ordinary skill in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0018] Figure 1 Structure diagram of the display screen-based data recommendation system provided by the embodiments of the present application; Figure 2 Structure diagram of the non-vision sensor working in the embodiments of the present application; Figure 3 Flowchart of the display screen-based data recommendation system provided by the embodiments of the present application; Figure 4 Flowchart of the display screen-based data recommendation system provided by the embodiments of the present application; Figure 5 Schematic diagram of the transparent OLED display screen displaying in the embodiments of the present application; Figure 6 Flowchart of the display screen-based data recommendation system provided by the embodiments of the present application; Figure 7 Flowchart of the display screen-based data recommendation system provided by the embodiments of the present application; Figure 8 Flowchart of the display screen-based data recommendation system provided by the embodiments of the present application; Figure 9 Flowchart of the display screen-based data recommendation system provided by the embodiments of the present application; Figure 10 Schematic diagram of the depth-of-field camera working in the embodiments of the present application; Figure 11 Schematic diagram of the depth-of-field camera working in the embodiments of the present application. DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the accompanying drawings in the embodiments of the present application to make a clear and complete description of the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by a person skilled in the art without creative work belongs to the scope of protection of the present application.
[0020] Firstly, the application scenarios applicable to the present application are introduced. The present application can be applied to the technical field of information recommendation.
[0021] It is found through research that the current mainstream mobile intelligent terminal device generally adopts a traditional LCD liquid crystal screen as a display core. Due to the fact that the LCD structure must rely on a backlight module, the overall device is thick and heavy, the screen ratio is limited, and transparent display cannot be achieved, which seriously affects the visual immersion and technological sense. In terms of function implementation, the existing device continues the simple magnification design idea of "tablet computer + support", and the function is concentrated in the application layer call, lacking deep perception ability to the environment and the user, and the interaction mode mainly relies on touch and voice, which cannot provide active intelligent services based on scene understanding. In addition, the non-transparent screen completely blocks the physical space behind when in use, resulting in the separation of digital information and the real environment.
[0022] The currently simply integrated device has obvious limitations: the system cannot accurately understand the environmental semantics, and it is difficult to achieve stable alignment of virtual information and real scenery; the interaction mode lacks naturalness and cannot support intuitive augmented reality interaction; at the same time, the existing device lacks emotional perception and connection ability in the human-computer interaction process, which limits the further improvement of user experience.
[0023] Based on this, the embodiments of the present application provide a data recommendation system based on a display screen, which realizes active content recommendation based on user state and environmental scene in a transparent display mode through the cooperative work of a transparent OLED display screen (Transparent Organic Light-Emitting Display, transparent organic light-emitting display) and a multi-sensor module, combined with the scene perception and intelligent analysis ability of a processor, while supporting virtual-real fusion interaction and multi-modal emotional interaction in an augmented reality scene, thereby improving the intelligent level of the device and the user experience.
[0024] Please refer to Figure 1 , Figure 1The structural diagram of the data recommendation system based on the display screen provided in the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the data recommendation system provided in the embodiments of the present application comprises a transparent OLED display screen 1, a millimeter wave radar 21, a pulse ultra-wideband chip 22, a color camera 31, a depth-of-field camera 32, a mobile chassis 4, an encoder 41 and an inertial measurement unit 42. Figure 1
[0025] The transparent OLED display screen 1 is installed on the front of the device and is directly connected to the processor through a data bus. The transparent OLED display screen 1 adopts a self-luminous structure and does not need a backlight module, thereby realizing thin and light device and high screen ratio, and supporting transparent display mode, so that digital information can be naturally integrated with the real scene, and the problem of traditional device and environment being cut off is solved.
[0026] The display interface uses transparent OLED technology, mainly using the transmittance of the transparent OLED ≥ 45% to avoid the oppression caused by the non-transparent shell of the conventional LCD product.
[0027] AR virtual and real combination, the LCD product with ordinary rear shell cannot pass through the screen to see the object behind, while the transparent OLED uses its own natural transparent advantage to easily realize the virtual and real interaction of air touch.
[0028] The multi-sensor module comprises a non-vision sensor for non-contact sensing of the user's presence state and a vision sensor for collecting environmental images.
[0029] Preferably, real-time data is generated by fusing detection of multiple sensors, and the sensors and peripheral hardware are not limited to millimeter wave radar, UWB chip, RGB camera, ToF depth-of-field camera and microphone array, etc.
[0030] Here, the non-vision sensor comprises the millimeter wave radar 21 and the pulse ultra-wideband chip 22, wherein the millimeter wave radar 21 is arranged in the shell of the system and is used for emitting and receiving electromagnetic waves to non-contact sense the chest movement and breathing frequency of the user; the pulse ultra-wideband chip 22 is arranged in the shell of the system and is used for non-contact sensing of the presence state and movement trajectory of the user by measuring the flight time of the pulse signal.
[0031] In terms of personnel sensing, specifically, please refer to Figure 2 , Figure 2 The structural diagram of the non-vision sensor working in the embodiments of the present application is shown in FIG. 2. As shown in FIG. 2, the millimeter wave radar 21 and the pulse ultra-wideband chip 22 are used for non-contact sensing of the movement trajectory and breathing frequency of the person, judging whether the user is present and whether the user is asleep, and deciding whether to automatically turn off the screen or enter the low-power mode. Figure 2
[0032] Figure 2 The basic principle of pulsed ranging is explained. The core idea is that the sensor emits a signal, which is reflected back after encountering the human body, and by measuring the time it takes for the signal to go from emission to reception, the distance between the person and the device can be calculated.
[0033] Non-vision sensors refer to millimeter wave radar and pulse ultra-wideband chips. The sensor emits an energy pulse into space. For millimeter wave radar, this pulse is an electromagnetic wave; for pulse ultra-wideband chips, this pulse is an extremely narrow radio pulse; the pulse is reflected after encountering the user (human body) on the propagation path, and the sensor receives the reflected pulse.
[0034] The transmission time t refers to the total time the pulse takes to be emitted from the sensor, reflected, and returned to the sensor. This is the basis of all calculations.
[0035] The distance D is the straight-line distance between the sensor and the target user. It is calculated by a classic physical formula: D = (c x t) ÷ 2 (where c is the speed of light). The reason for dividing by 2 is that the pulse travels a distance of "sensor-human-sensor", which is twice the distance D.
[0036] Millimeter wave radar 21 can detect the position and speed of the patient's chest through the combination of frequency-modulated continuous wave (FMCW) detection and multiple-input multiple-output (MIMO) antenna radar systems.
[0037] By emitting 24GHz / 77GHz frequency electromagnetic waves, the electromagnetic waves can penetrate non-metallic materials such as clothes, thin quilts, etc. covering the surface of personnel, capture 0.1-0.5mm level fluctuations in the chest cavity, and the system analyzes the target distance, speed and angle through the reception of the reflected signal, and analyzes the respiratory rate through signal analysis (accuracy ±1 times / minute).
[0038] By capturing 0.1-0.5mm level fluctuation information of the chest cavity, it can be analyzed whether there is someone in the room. If the fluctuation signal of the chest cavity cannot be captured, it means that there is no one.
[0039] Pulse ultra-wideband chip 22 (UWB chip) works in the 3.1-10.6GHz frequency band, and determines the distance by measuring the two-way time of flight (TW-TOF) of the pulse signal. Combined with the Doppler effect, it can capture the tiny body surface displacement caused by human respiration or heartbeat (the accuracy can reach centimeter level), and realize real-time sensing of human presence state and motion trajectory.
[0040] Specifically, the millimeter wave radar 21 and the pulse ultra-wideband chip 22 are built-in in the device shell as non-vision sensors, and are connected with the processor through a special interface. The millimeter wave radar 21 transmits and receives specific frequency electromagnetic waves to penetrate non-metal materials such as clothes and perceive the micro-movement of the user's chest cavity, thereby realizing non-contact breathing frequency detection. The pulse ultra-wideband chip 22 accurately captures the user's motion trajectory by calculating the time of flight of the pulse signal. The cooperative work of the two realizes accurate judgment of the user's presence state and sleep state without relying on optical collection, which not only protects the user's privacy but also provides the system with a silent active service capability.
[0041] The vision sensor includes a color camera and a depth camera. The color camera 31 is arranged on the shell facing the back of the transparent OLED display screen 1, and is used to collect environmental images to identify scene types and objects. The depth camera 32 is arranged around the color camera 31, and is used to obtain depth information of the real scene.
[0042] The color camera 31 and the depth camera 32 are installed side by side on the shell of the device back facing the environment behind the transparent OLED display screen 1 as a vision sensor group, and are connected with the processor through an image collection interface. The color camera 31 is responsible for collecting environmental color images and identifying scene types and objects through computer vision algorithms. The depth camera 32 obtains scene depth information by actively projecting and receiving infrared light signals. The data fusion of the two enables the system to accurately understand the environmental semantics and provide necessary spatial geometric data for virtual-real fusion.
[0043] The mobile chassis 4 is arranged at the bottom of the system and is integrated with an encoder 41 and an inertial measurement unit 42. The inertial measurement unit 42 is in communication connection with the processor, and is used to detect the position change and movement speed of the system.
[0044] The mobile chassis 4 is installed at the bottom of the device and is fixed with the main shell through a mechanical structure. The integrated encoder 41 and inertial measurement unit 42 of the mobile chassis 4 are in communication with the processor through a motion sensing interface. The mobile chassis 4 records the wheel hub rotation data through the encoder 41, and combines the acceleration and angular velocity data collected by the inertial measurement unit 42 to calculate the device displacement and attitude change in real time. This enables the system to dynamically adjust the perspective relationship of the display content according to the device movement, ensures that the virtual information and the real scene remain stable registration, and solves the technical problem of AR content alignment in a mobile environment.
[0045] The inertial measurement unit (IMU, Inertial Measurement Unit) is a device for measuring the three-axis attitude angle (or angular velocity) and acceleration of an object. The IMU integrates the functions of three-axis gyroscope and three-axis accelerometer MEMS inertial sensors. It is used to detect the change of the position where the chassis is located and the change of the movement speed.
[0046] a processor configured to: identify a user presence state and an environmental scene type where the device is located according to data of the non-vision sensor and data of the vision sensor; and actively display recommended content associated with the current scene on the transparent OLED display based on the user presence state and the environmental scene type.
[0047] Here, the processor continuously receives and fuses the respiration rate signal from the millimeter wave radar 21 and the motion trajectory signal from the pulse ultra-wideband chip 22. Through a pre-set algorithm model (such as a state machine or a machine learning classifier), the processor analyzes the continuity, intensity and pattern of these signals. For example, when a continuous, regular respiration signal is detected and the body movement signal is extremely weak, the user is determined to be in a "sleep" state; when irregular respiration and obvious movement signals are detected, the user is determined to be in an "active" state; when no vital sign signal is detected for a long time, it is determined to be an "off-site" state.
[0048] The processor calls the color camera 31 to capture environmental images and performs image preprocessing (such as noise reduction and enhancement). Subsequently, the image is feature-extracted using an embedded computer vision model (for example, a trained convolutional neural network CNN), and the features are compared with a pre-set scene database (containing typical features of living room, kitchen, bedroom, etc. Scene) to calculate the similarity. Finally, the processor takes the highest similarity scene type (such as "kitchen") and the identified specific items (such as "tomatoes", "eggs") as the recognition result.
[0049] The system realizes the stereoscopic and deep understanding of the environment and the user state by fusing non-contact physiological sensing and precise visual recognition technology. This ability lays a solid foundation for the system to realize active intelligent service, and completely breaks away from the limitations of the traditional passive response mode. At the same time, the system protects user privacy to the greatest extent by giving priority to using non-vision sensors to perform sensitive state detection under the premise of ensuring function implementation, embodying the concept of unobtrusive and thoughtful service.
[0050] Then, the processor takes the above-mentioned recognition result (such as "user: active, scene: kitchen, items: tomatoes, eggs") as input and makes decisions through a recommendation strategy engine. The engine is embedded with a rule base or an association model (for example, "IF kitchen AND tomatoes AND eggs THEN recommend scrambled egg with tomatoes recipe"). The processor then retrieves and calls the most matched recommended content (such as recipe video, text and step) from the content library.
[0051] The transparent OLED display screen 1 enters the transparent display mode, and the recommended content is rendered in the form of semi-transparency or small window (such as "popping up" in the corner) in a specific area of the screen. This display mode ensures that the digital content does not completely block the real scene behind the screen (such as pots and pans), and the user can continue to operate in reality while viewing the recommendations.
[0052] The integrated millimeter wave radar 21, pulse ultra-wideband chip 22 (UWB chip), color camera 31 (RGB camera), depth camera 32 (ToF depth camera), and microphone array and other sensor peripheral devices realize "active" content recommendation and interaction based on scene perception through corresponding algorithms.
[0053] The present application realizes the transition from passive response to active service, making the device an intelligent companion that can anticipate needs. Through transparent display technology, the system seamlessly integrates digital information with the real scene, and at the same time, based on the accurate perception of the scene and user state, the system can provide highly relevant services at the most needed time, significantly improving interaction efficiency, practicality, and user stickiness.
[0054] Further, please refer to Figure 3 , Figure 3 The flowchart of the display screen-based data recommendation system provided by the embodiment of the present application, the processor is further configured to: S101, obtain the wheel hub rotation data of the mobile chassis through the encoder, and calculate the displacement vector and rotation angle of the system in combination with the three-axis acceleration and angular velocity data obtained by the inertial measurement unit.
[0055] The encoder 41 directly measures the motor rotation number and angular velocity of the drive wheel hub, and the accelerometer of the inertial measurement unit 42 (IMU, Inertial Measurement Unit) measures the linear acceleration on the X, Y, and Z axes, and the gyroscope measures the rotation angular velocity around the three axes.
[0056] The processor first solves the encoder data, calculates the preliminary displacement and heading angle change of the device in the two-dimensional plane according to the wheel hub diameter and the number of rotations; at the same time, the processor integrates the angular velocity data of the IMU to obtain the attitude angle change of the device; the acceleration data is integrated (the gravity acceleration component needs to be removed) and fused with the displacement data of the encoder, and is corrected through a sensor fusion algorithm (such as Kalman filter) to eliminate the inherent drift error of the IMU data; finally, the processor outputs a more accurate three-dimensional displacement vector (the movement distance and direction of the device in space) and rotation angle (the yaw, pitch, and roll angles of the device).
[0057] By fusing the accurate odometry information of the encoder and the high-frequency pose information of the IMU, the system can accurately perceive its arbitrary movement (including translation and rotation) in real time.
[0058] S102, based on the displacement vector and the rotation angle, the current pose of the transparent OLED display screen relative to the real scene coordinate system is determined through coordinate transformation, and the current pose includes a position coordinate and an orientation angle.
[0059] The input parameters are the displacement vector and the rotation angle calculated by S101.
[0060] The system will establish a real scene coordinate system at the initial moment (for example, taking the position and orientation of the device at the time of starting as the origin), and the processor will transform the device from the pose at the last moment to the new pose relative to the fixed world coordinate system at the current moment through the three-dimensional coordinate transformation matrix (including the rotation matrix and the translation vector). The pose is a parameter containing 6 degrees of freedom, i.e. 3 degrees of freedom of the position coordinate (X, Y, Z) and 3 degrees of freedom of the orientation angle (yaw, pitch, roll).
[0061] The present application converts the physical movement of the device into a computer-processable, accurate position and direction in a virtual coordinate system. This enables the system to clearly know "where is the screen currently in the real world, and which direction is it facing", which is a key prerequisite for virtual content to align with the real scene.
[0062] S103, according to the current pose, the field of view angle change of the transparent OLED display screen and the perspective projection matrix are calculated.
[0063] The input parameters are the current pose determined by S102, and the system has preset the inherent parameters of the camera (such as focal length, principal point coordinates) and the physical parameters of the screen.
[0064] The movement and turning of the device will cause the field of view angle (FOV) of the device "observing" the real world to change. The processor calculates this change according to the current pose (especially the orientation angle); more importantly, the processor calculates a perspective projection matrix in real time according to the current pose (viewpoint position and orientation) and the preset camera parameters. This matrix defines the mathematical rules for how points in a three-dimensional virtual world are projected onto a two-dimensional screen.
[0065] The perspective projection matrix of the present application is the core of simulating the observation of the world by the human eye or the camera. By updating this matrix in real time, the system ensures that the form of virtual objects presented on the screen can conform to the perspective rules under the current viewing angle (such as near large and far small, occlusion relationship).
[0066] S104, based on the perspective projection matrix, adjusting the display coordinates, size ratio and perspective relationship of the recommended content displayed on the transparent OLED display, so that the recommended content is spatially registered with the corresponding target object in the real scene.
[0067] The input parameter is the perspective projection matrix calculated in S103, and at the same time, the system knows the pre-set anchor position (for example, fixed 20 cm above the recognized table) of the virtual content (such as the recipe card) in the real scene coordinate system.
[0068] The processor multiplies the three-dimensional model vertex coordinates of the virtual content with the perspective projection matrix to perform coordinate transformation, and finally obtains the two-dimensional display coordinates of the virtual content on the transparent OLED display 1; at the same time, according to the distance of the viewpoint, the size ratio and the perspective deformation are automatically calculated and adjusted.
[0069] When the device moves, the S101-S103 process continues to run, the perspective projection matrix is updated, and then the display parameters of the virtual content on the screen are driven to change in real time, so that it looks like "stuck" on the specified real location.
[0070] The application solves the core challenge of AR interaction on the transparent screen - the alignment of virtual information and real scenery. No matter how the user moves the device, the virtual content can maintain stable spatial registration with the real scene, thereby providing an immersive and realistic AR experience.
[0071] Further, please refer to Figure 4 , Figure 4 The processor of the flowchart two of the display screen-based data recommendation system provided by the embodiment of the application is further configured to: S201, based on the depth camera, acquiring the depth information of each object in the real scene, and determining the target anchor position of the virtual content in the real scene.
[0072] The depth camera 32 (such as a ToF camera). The camera measures the light flight time by emitting and receiving light pulses, calculates an accurate distance value for each pixel point in the scene, and generates a depth map.
[0073] The processor combines the two-dimensional color image collected by the color camera 31 and the depth map provided by the depth camera 32, and constructs a three-dimensional space model of the environment around the device through SLAM (simultaneous localization and mapping) or similar three-dimensional reconstruction algorithm.
[0074] Subsequently, the processor identifies a specific plane or object in the environment through a computer vision algorithm (e.g., identifies a flat tabletop through depth information). The target anchor position is determined as a specific three-dimensional coordinate point above the identified physical plane (e.g., 20 cm above the center of the tabletop).
[0075] The present application upgrades the system from "seeing" the world to "understanding" the three-dimensional structure of the world by obtaining depth information, so that a reasonable and stable real anchor point can be found for virtual content, solving the fundamental problem of "where to place virtual objects".
[0076] Further, please refer to Figure 5 , Figure 5 The schematic diagram of the transparent OLED display screen provided by the embodiment of the present application for display.
[0077] As Figure 5 shown, the rear scene is captured in real time by the camera, and the depth information is obtained by the ToF camera. The algorithm corrects the position and perspective relationship of the AR image in real time according to the movement of the device (perceived by the encoder and IMU unit on the chassis) and the change of the field of view angle, ensuring that the virtual object is "stably" placed on the real tabletop.
[0078] Figure 5 Several levels of implementing AR display on the transparent OLED screen are decomposed from the structure.
[0079] Background entity object: This is the object in the real world that can be seen through the transparent OLED screen, such as a water cup placed on a table.
[0080] Transparent OLED: This is the hardware core of the present application. It is not an ordinary light-shielding screen, but a display medium that can transmit the real scene behind it like glass.
[0081] Virtual display content: This is digital information or graphics generated by the system and rendered on the transparent OLED screen, such as a floating recipe card.
[0082] The effect ultimately presented to the user is AR virtual-real fusion, which describes the visual integration of virtual display content and background entity objects, forming a unified, augmented reality picture.
[0083] S202, according to the relative spatial relationship between the target anchor position and the transparent OLED display screen, calculate the initial display parameters of the virtual content on the display screen.
[0084] The system knows the pose of the transparent OLED display screen 1 at the current time (from S102 of Figure 3 , i.e., the position and orientation of the screen in three-dimensional space).
[0085] The processor calculates the relative spatial relationship (including distance and angle) between the target anchor position and the display screen pose.
[0086] Based on this relative relationship and an initial perspective projection matrix (which can be calculated according to the camera intrinsic parameters and the initial pose), the processor projects the three-dimensional model of the virtual content from the world coordinate system to the two-dimensional screen coordinate system, thereby calculating the initial coordinates, size and orientation of the virtual content on the screen, etc.
[0087] This step converts the three-dimensional world coordinates into two-dimensional screen coordinates, so that the virtual content can first appear correctly on the corresponding position of the screen, forming a basic correspondence with the real object it is anchored to, laying the foundation for the fusion display.
[0088] S203, real-time dynamic correction of the initial display parameters combined with the inertial measurement unit data from the mobile chassis.
[0089] Due to the delay in image processing of the camera, when the device experiences rapid and slight shaking, relying solely on visual information will cause the virtual content to "jump" or "lag" on the screen.
[0090] The processor uses IMU data (especially the angular velocity of the gyroscope) to predict the slight attitude change of the device in a very short time. Then, it performs feedforward correction on the initial display parameters (especially the position and perspective) calculated in S202 according to this predicted slight change. This is similar to the optical image stabilization technology in cameras or mobile phones, but applied in the field of AR rendering.
[0091] The high-frequency data of the IMU compensates for the lag of the visual processing, allowing the virtual content to "follow" every slight movement of the screen closely, eliminating the shaking and trailing phenomenon. This ensures that the virtual object appears to be truly "nailed" in the real world, greatly improving the stability and realism of the fusion display.
[0092] S204, according to the corrected display parameters, the virtual content is matched with the real scene for perspective fusion display.
[0093] The graphics rendering engine uses these corrected parameters to draw the virtual content (such as 3D models, UI interfaces) into the frame buffer of the transparent OLED display screen 1.
[0094] Since the screen is transparent, the user can also see the real scene behind the screen. The virtual content is superimposed on the real scene with semi-transparency or appropriate opacity, and its perspective relationship and occlusion relationship (processed through depth information) are perfectly matched with the real world, as if it is part of the real environment.
[0095] The results of all previous steps - spatial perception, dynamic tracking, perspective projection, are finally converted into a virtual-real fusion picture that is visible to the naked eye and has no sense of strangeness. This enables users to intuitively interact with virtual information suspended in real objects, fully exploiting the huge potential of transparent OLED in the AR field and solving the core technical challenge of "unnatural and misaligned AR interaction on transparent screens".
[0096] Specifically, the visual sensor also includes a gesture recognition camera, which is usually set on the top or side of the device. This camera is specially optimized for close-range, high-frame-rate hand capture and continuously captures video streams to generate image sequences of the user's hand, ensuring that its field of view can completely cover the space in front and behind the transparent OLED display 1 for interaction.
[0097] The rear scene is captured in real time by the camera, and the depth information is obtained using the ToF camera. Algorithmically, the position and perspective relationship of the AR image are corrected in real time according to the movement of the device (sensed by the encoder 41 and inertial measurement unit 42 on the chassis hub) and the change in viewing angle, and clicking and gesture sliding operations are performed through the gesture recognition camera.
[0098] The user does not need to directly touch the screen. Through the gesture recognition camera, the user can "click" the virtual buttons suspended behind the screen on the real objects, or browse the information stream superimposed on the transparent screen through gesture sliding. This interaction method is very futuristic and avoids leaving fingerprints on the transparent screen.
[0099] Among them, please refer to Figure 6 , Figure 6 The processor of the flowchart three of the display screen-based data recommendation system provided by the embodiment of the present application is further configured to: S301, acquire image sequences of the user's hand through the gesture recognition camera.
[0100] The processor calls the image acquisition driver to obtain the original RGB or near-infrared image sequence, providing a data source for subsequent analysis.
[0101] S302, extract hand key point spatial coordinates and motion trajectories from the image sequence.
[0102] The processor runs a hand key point detection model (an optical recognition algorithm based on deep learning). The model identifies and outputs the two-dimensional pixel coordinates of dozens of key points of the hand (such as fingertips, joints, wrists, etc.) in the image sequence. Then, by tracking the positions of the same key points in consecutive frames and combining the timestamps, the processor calculates the motion trajectory (including the moving direction, speed, and acceleration) of each key point; the blurred "hand image" is converted into accurate and quantitative "hand motion data", so that the system can accurately perceive the posture and motion process of the hand, providing a fundamental basis for subsequent intent recognition.
[0103] S303, match the hand key point spatial coordinates and motion trajectory with a predefined gesture instruction library to identify the user's operation intent.
[0104] A predefined gesture instruction library is built in, which stores standard gesture templates corresponding to various instructions (for example, "straighten the index finger and move it forward quickly" corresponds to the "click" instruction; "open the palm and move it left and right" corresponds to the "page turning" instruction). The processor performs real-time pattern matching and classification of real-time gesture data with the templates in the instruction library. When the matching degree exceeds the set threshold, it is determined that the user has issued a specific operation intent (such as "the user intends to click the floating'start' button").
[0105] Instead of simply "seeing" that the hand is moving, the system accurately "understands" what command the user wants to execute. This allows users to interact with virtual content in the most natural and intuitive way, completely eliminating the need for physical controllers or direct touch of the screen.
[0106] S304, according to the identified operation intent, generate control instructions for the fused display virtual content and perform corresponding operations.
[0107] The intent is converted into standard control instructions that the operating system or application can understand (for example: generate a "click" event and position its coordinates at the center of the virtual button; or generate a "slide" event with direction and speed parameters). Then, the system executes the instruction, triggering the corresponding operation at the application layer, such as confirming the selection, browsing the information stream, playing or pausing the video.
[0108] Further, please refer to Figure 7 , Figure 7 For the fourth flowchart of the display screen-based data recommendation system provided by the embodiments of the present application, the processor is further configured to: S401, pre-process the environment image collected by the visual sensor, and extract the scene visual feature vector from the pre-processed environment image.
[0109] The environment image is a raw RGB image captured by the color camera 31 in real time, which contains the complete visual information of the scene behind the device.
[0110] The processor performs a series of optimization operations on the raw image, including: scaling the image to a fixed size required by the model (such as 224x224 pixels); using a filtering algorithm to reduce noise introduced during image acquisition; adjusting image parameters to make key features more prominent.
[0111] The pre-processed image is input into a pre-trained deep learning model (e.g., a convolutional neural network, CNN). The model acts as a "feature extractor", with the layer before the final fully connected layer outputting a high-dimensional numerical array, which is the scene visual feature vector. This vector is an abstract mathematical representation of the original image, condensing the visual patterns (such as texture, shape, spatial layout) in the image that are most relevant to scene classification.
[0112] In this step, through feature extraction, the system can discard irrelevant details in the image (such as changes in lighting, small debris), and grasp the essential visual elements that can distinguish between different scenes such as living room, kitchen, bedroom, etc.
[0113] S402, calculate the similarity between the scene visual feature vector and the standard scene features in the pre-set scene database.
[0114] The pre-set scene database is a database built before shipment or during the training phase, which stores a variety of standard scene features. Each feature is a feature vector of the same dimension as the output of S401, and has a label (such as "kitchen", "living room", "bedroom"). These standard features are representative vectors obtained by training the same feature extraction model with a large number of labeled scene images.
[0115] The processor compares the real-time extracted scene visual feature vector with each standard scene feature in the database. The comparison is achieved by calculating the similarity (such as cosine similarity or the inverse of Euclidean distance) between them. The calculation result is a numerical value, and the higher the value, the more similar the current scene is to a certain standard scene.
[0116] S403, according to the similarity calculation result, determine the scene type with the highest matching degree as the recognition result of the environment scene type.
[0117] The processor executes a decision logic to find the maximum value among all similarity scores. Then, it is determined whether this maximum value exceeds a pre-set confidence threshold (to avoid making incorrect judgments in uncertain scenarios).
[0118] If the threshold is exceeded, the label of the standard scene corresponding to the maximum value (such as "kitchen") is determined as the final recognition result of the environmental scene type. If the threshold is not exceeded, it can be determined as "unknown scene" or the last judgment is maintained.
[0119] Further, please refer to Figure 8 , Figure 8 For the flowchart of the display screen-based data recommendation system provided by the embodiments of the present application, the processor is further configured to: S501, collect user voice signals through a microphone array, and extract acoustic feature parameters of the voice.
[0120] The user voice signal is the original audio data collected from the microphone array. The microphone array can not only pick up sound, but also focus on the user voice through beamforming technology and suppress environmental noise.
[0121] The processor pre-processes the original voice signal, such as noise reduction, frame division, etc.
[0122] Subsequently, acoustic feature parameters for emotion recognition are extracted from the processed signal. These parameters include: prosodic features such as fundamental frequency (reflecting pitch), energy (reflecting volume); voice quality features such as voice spectrum, formant; rate features such as speech rate, pause frequency.
[0123] Convert sound into quantifiable emotion indicators: this is the basis of affective computing. By extracting these acoustic parameters that are strongly related to emotion, the system can objectively analyze the emotional information carried in the user's voice, providing reliable data input for subsequent emotion classification.
[0124] Fusion of speech emotion recognition (judging user emotion from voice tone), facial expression recognition (judging user joy and sorrow), behavior analysis (whether the user is pacing anxiously).
[0125] Voice emotion is collected through an audio source pickup microphone, and the emotional state is recognized by analyzing features such as voice pitch, speech rate, tone fluctuation, etc.; high pitch represents excitement or anxiety, low pitch represents depression or negative emotion; fast speech indicates agitation or anxiety, slow speech represents calmness or low emotion.
[0126] Facial expression recognition identifies facial expression state through a camera, such as smiling with raised corners of the mouth and wrinkles at the corners of the eyes; angry with furrowed brows and closed lips; sad with drooping corners of the mouth and watery eyes; voice emotion and facial expression recognition are both determined by comparing a large amount of trained expression data and voice data.
[0127] S502, input the acoustic feature parameters into the voice emotion classification model to obtain a first emotion recognition result.
[0128] These feature parameters are input into a pre-trained speech emotion classification model (e.g., a model based on support vector machine SVM or recurrent neural network RNN). The model is trained on a large set of speech data labeled with emotion labels, and can map the input feature vector to a specific emotion category or dimension (e.g., "happy", "sad", "calm", "angry", or "pleasure" and "excitement" scores). The output of the model is the first emotion recognition result.
[0129] S503, collect user facial image through visual sensor and extract facial feature point motion parameters.
[0130] The user facial image comes from the image collected by the color camera 31 in real time.
[0131] The processor runs a face detection algorithm to locate the face region in the image; then, using a face key point detection model, the coordinates of the key feature points related to facial muscle movement (such as the corners of the mouth, the tips of the eyebrows, and the corners of the eyes) on the face are identified, and by analyzing the changes in the coordinates of these feature points between consecutive frames, the facial feature point motion parameters are calculated, such as the upward amplitude of the corners of the mouth, the degree of frowning of the eyebrows, and the size of the eyes.
[0132] Convert the ambiguous "expression" into precise and calculable Action Units data. It enables the system to objectively quantify facial expressions and provides a solid foundation for expression-based emotion recognition, avoiding subjective judgment errors.
[0133] S504, input the facial feature point motion parameters into the expression recognition model to obtain the second emotion recognition result.
[0134] These motion parameters are input into a pre-trained expression recognition model. The model learns the mapping relationship between facial movements and emotion categories (e.g., mouth up + eye wrinkles may correspond to "happy"). The output of the model is the second emotion recognition result (an emotion category or dimension score). Through expression recognition, the system can "understand" the user's facial expression, even when not speaking.
[0135] S505, fuse the first emotion recognition result and the second emotion recognition result to determine the final user emotion state.
[0136] The processor adopts decision-level or feature-level fusion strategy.
[0137] In an optional embodiment, decision-level fusion can assign confidence weights to the two results, then perform weighted averaging to obtain the final conclusion. If the signal quality of one modality is poor (e.g., the face is blocked), its weight is reduced.
[0138] In another optional embodiment, the feature-level fusion can concatenate the raw features or the probability vectors of the model outputs from the two modalities at an earlier stage and input them into a final fusion classifier for decision making.
[0139] Through fusion, the system outputs a more reliable and comprehensive final user emotional state, which single modality recognition may not be accurate due to environmental interference or personal habits. Multi-modal fusion can take advantage of different information sources, complement and correct each other, and thus make a judgment closer to the user's real emotional state.
[0140] S506, according to the final user emotional state, selecting matched recommended content from the content library for display.
[0141] The system maintains a content library, in which the contents are labeled with various tags, including emotional tags (such as "relaxing", "uplifting", "funny").
[0142] The application has a multi-modal AI emotion computing engine, which recognizes and analyzes speech emotion, facial expression, behavior, etc., and automatically generates content suitable for the corresponding scene.
[0143] The virtual image assistant uses a microphone and a camera to collect speech and behavior, and changes its tone, expression and recommended content according to the recognized user emotion, for example, detects that the user's emotion is low, and actively plays soothing music or tells a joke; detects that the user is full of energy, and recommends high-intensity fitness courses.
[0144] The processor executes the preset recommendation strategy according to the recognized emotional state. For example, if the emotional state is "sad", the content with the tags "relaxing" and "inspiring" is retrieved from the content library; if the emotional state is "happy", the content with the tags "active" and "high energy" is retrieved from the content library.
[0145] Further, please refer to Figure 9 , Figure 9 The sixth flowchart of the display screen-based data recommendation system provided by the embodiments of the application, the processor is further configured to: S601, continuously collecting the chest movement signals of the user through the millimeter wave radar, and continuously collecting the body movement signals of the user through the pulse ultra-wideband chip.
[0146] The chest movement signal comes from the millimeter wave radar 21. The radar emits a frequency-modulated continuous wave (FMCW) in the 24GHz / 77GHz band, and receives the signal reflected by the human chest. Since the chest will undergo a small periodic fluctuation (0.1-0.5mm) when breathing, this movement will modulate the phase and frequency of the reflected wave, thereby generating a chest movement signal containing breathing information.
[0147] The body motion signal comes from a pulse ultra-wideband chip 22. The chip calculates the distance accurately by measuring the two-way time of flight (TW-TOF) of a pulse signal between the device and the user's body surface. The user's body movements of larger amplitude (such as turning over, waving hands) will cause significant changes in the distance, which are recorded as body motion signals.
[0148] The two non-vision sensors work together, can penetrate the cover such as thin quilt, in the premise of not invading the user's privacy, 7x24 hours, no feelingly obtains the user's most core physiological and behavior data (respiration and body motion).
[0149] S602, the thoracic motion signal and the body motion signal are filtered respectively, the respiratory frequency is extracted from the thoracic motion signal, and the body motion feature is extracted from the body motion signal.
[0150] The processor carries out digital filtering (such as band-pass filtering) on the original signal, so as to remove high-frequency noise (such as environmental electromagnetic interference) and low-frequency drift (such as the user slowly moving away from the device), and retain the useful frequency band in the signal.
[0151] The thoracic motion signal after filtering is subjected to time-frequency analysis (such as fast Fourier transform FFT), and the main frequency is found, which corresponds to the respiratory frequency (unit: times / minute); usually, the variance, amplitude or energy of the filtered body motion signal in the time window is calculated as the body motion feature. This numerical value quantifies the activity level of the user's body.
[0152] S603, the amplitude of the respiratory frequency and the body motion feature is analyzed, and it is judged whether the user is in a static sleep state.
[0153] The processor executes a preset determination logic. The logic is based on the knowledge of sleep physiology: when in a sleep state, the user's respiratory frequency will usually become slow and regular (for example, the resting respiratory frequency of an adult is in the range of 12-20 times / minute, and when sleeping, it may be in this range or lower).
[0154] At the same time, the amplitude of the body motion feature will be significantly reduced and maintained at a low level, and the system analyzes whether the two parameters meet the mode of “steady breathing” and “body at rest” in real time.
[0155] S604, when the respiratory frequency meets the preset condition and the body motion feature is lower than the threshold value for a plurality of continuous periods, it is determined that the user is in a sleep state.
[0156] In order to avoid misjudgment (such as the user just sitting quietly), the system introduces a persistence and confidence mechanism.
[0157] Preferably, the preset condition can be that the respiratory frequency is in the range of 10-18 times / minute.
[0158] Threshold: the body motion feature value is lower than an empirically set low activity threshold.
[0159] The processor requires the above-mentioned "breathing is smooth and the body is still" state to be continuously maintained for multiple detection periods (for example, for 2-5 minutes). Only when the condition is continuously met, the system finally determines that the user is in a sleep state. This greatly improves the accuracy of the judgment.
[0160] And by introducing the requirement of duration, the short-term still behavior (such as reading a book, resting) is filtered, the misjudgment rate is reduced, it is ensured that the subsequent energy-saving operation is triggered only when it is really needed, the user is avoided from being disturbed, and the precision and thoughtfulness of the service are embodied.
[0161] S605, trigger a system control instruction, switch the system to a low-power consumption mode or turn off the display function of the transparent OLED display screen.
[0162] The processor then generates and executes a system control instruction.
[0163] For example, the instruction can include: turning off the backlight of the transparent OLED display screen 1 or completely powering off, eliminating the potential interference of light on the user's sleep.
[0164] Make the system enter a low-power consumption mode, suspend part of the high-power consumption functions (such as high-intensity calculation, unnecessary sensor scanning), and only keep the basic vital sign monitoring function.
[0165] The present application can automatically turn off the screen and save energy after the user falls asleep, which not only creates a better sleep environment but also prolongs the battery life of the device. The entire process does not require the user to manually set or issue instructions, and the concepts of "active intelligence" and "privacy protection" are taken to the extreme, significantly improving the user experience and practical value of the product.
[0166] In an optional embodiment of the present application, please refer to Figure 10 , Figure 10 is one of the schematic diagrams of the depth-of-field camera provided by the embodiments of the present application.
[0167] Depth of field refers to the range of front and back distances that can be clearly imaged when a camera is shooting. This range is not a point, but a space. Figure 10 In the figure, the shooting distance L refers to the straight-line distance from the camera to the photographed object; the distance L1 of the photographed object emphasizes the specific position of the photographed subject within the depth of field.
[0168] The depth of field δL refers to the entire range of vertical depth that can be clearly imaged, from the near point to the far point.
[0169] The foreground depth δL1 and the background depth δL2 are demarcated by the focus point (usually the subject), and the depth of field range can be divided into foreground depth (before the focus point) and background depth (after the focus point). Generally, the background depth is greater than the foreground depth.
[0170] As shown in the figure, when the shooting distance L is very small, the clear range (depth of field δL) will become very narrow, which is "shallow depth of field", and the background and foreground will be easily blurred. When the shooting distance L is very large, the entire clear range (depth of field δL) from the near point to the far point will become very wide, which is "large depth of field" or "deep depth of field", and the foreground and background can be clearly displayed.
[0171] The depth camera (ToF) can accurately obtain the shooting distance (L) of each pixel point in the scene by actively emitting light pulses and measuring the return time, thereby constructing a depth map of the entire scene. The system can understand the front and back spatial relationship between objects by using this information.
[0172] In an optional embodiment of the present application, please refer to Figure 11 , Figure 11 Figure 2 is a schematic diagram of the working of the depth camera provided by the embodiment of the present application.
[0173] This figure shows how the system corrects the AR content in real time through sensor data to maintain the stability of the virtual-real fusion when the device moves.
[0174] The user holds or carries the device through the mobile chassis 4, moves from position 1 to position 2 to position 3, and observes object A behind the screen.
[0175] ToF depth camera: continuously aligns with the real environment behind the screen to obtain the depth information of object A.
[0176] Transparent OLED display screen: displays virtual content (not shown in the figure, but intended to be integrated with the real scene).
[0177] Object A: an object in the real world, which is also the anchor target of the virtual content.
[0178] α1, α2, α3: represent the field of view (FOV) of the camera (i.e. the screen) and the viewing angle of object A at different positions. As the device moves, this angle changes constantly.
[0179] Encoder and IMU sensing unit: responsible for sensing the movement of the device, providing position change and attitude change (such as rotation angle) data.
[0180] From position 1 to position 3, the device is closer and closer to object A, the viewing angle changes from α1 to α3, the field of view also increases, and the range seen is wider but the object appears larger.
[0181] When the device moves from position 1 to position 2, the encoder and IMU immediately perceive this movement and change in pose.
[0182] The processor calculates the new perspective relationship (i.e. the new field of view angle a2) according to these motion data and in combination with the real-time depth information of object A provided by the ToF camera.
[0183] The processor then dynamically corrects the position, size and perspective distortion of the virtual content to be displayed, ensuring that the virtual object still appears to be "firmly" attached to object A.
[0184] Without this correction process, the virtual object will "drift" or "slide" on the screen and fail to align with the real scene.
[0185] Figure 10 And Figure 11 Together, the two key technologies support the high-quality AR interaction realized by the system: one is to perceive the three-dimensional spatial structure through the depth camera to provide the coordinate basis for the placement of virtual content; the other is to dynamically maintain the alignment of virtual content and real coordinates when the device moves through motion sensors and real-time algorithms, thereby solving the core challenge of realizing natural and stable AR interaction on a transparent screen.
[0186] The application provides seamless switching between immersive and transparent modes, movable AR virtual-real fusion, and active intelligent service experience. Millimeter wave radar is used to realize non-inductive sensing, to provide intelligent services while maximizing the protection of user privacy. The non-visual sensor (millimeter wave radar) and the visual sensor are fused to realize deep understanding of the environment and the user state on the premise of protecting privacy, and to provide silent active intelligent services. The emotional computing capability is combined with the transparent mobile display terminal, so that the device is upgraded from a tool to a "companion" with emotional perception and feedback capability, which highly fits the product positioning of the transparent display "best friend machine".
[0187] The data recommendation system based on the display screen provided by the embodiments of the application integrates a transparent OLED display screen and a multi-sensor module, intelligently identifies the user's on-site state and the environmental scene type by comprehensively analyzing the data of the non-visual sensor and the visual sensor through a processor, and actively pushes the recommended content associated with the current scene on the transparent screen, thereby realizing intelligent service based on scene perception and effectively improving the user experience.
[0188] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described herein.
[0189] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. The described device embodiments are merely illustrative, for example, the division of units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electric, mechanical or in other forms.
[0190] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0191] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit.
[0192] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer readable storage medium executable by a processor. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various media that can store program codes.
[0193] Finally, it should be noted that the above examples are merely specific embodiments of the present application, and are used to illustrate the technical solutions of the present application, but not to limit the same. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that, within the technical scope disclosed by the present application, any person skilled in the art can still modify or easily think of changes to the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some of the technical features; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A display screen based data recommendation system, characterized in that, The system comprises: a transparent OLED display screen; a multi-sensor module comprising a non-vision sensor for contactless sensing of a user's presence state and a vision sensor for collecting environmental images; a processor configured to: identify a user's presence state and an environmental scene type in which the device is located according to data of the non-vision sensor and data of the vision sensor; actively display recommended content associated with the current scene on the transparent OLED display screen based on the user's presence state and the environmental scene type.
2. The system of claim 1, wherein, The non-vision sensor comprises a millimeter wave radar and a pulse ultra-wideband chip, wherein the millimeter wave radar is arranged in a housing of the system and is used to emit and receive electromagnetic waves to contactlessly sense a user's chest movement and breathing frequency; the pulse ultra-wideband chip is arranged in the housing of the system and is used to measure the time of flight of a pulse signal to contactlessly sense a user's presence state and movement trajectory.
3. The system of claim 1, wherein, The vision sensor comprises a color camera and a depth-of-field camera, wherein the color camera is arranged on the housing facing the back of the transparent OLED display screen and is used to collect environmental images to identify a scene type and an object; the depth-of-field camera is arranged around the color camera and is used to obtain depth information of a real scene.
4. The system of claim 1, wherein, Further comprising a mobile chassis, wherein the mobile chassis is arranged at the bottom of the system and is integrated with an encoder and an inertial measurement unit, the inertial measurement unit is in communication connection with the processor and is used to detect the position change and movement speed of the system.
5. The system of claim 4, wherein, The processor is further configured to: obtain wheel hub rotation data of the mobile chassis through the encoder, combine three-axis acceleration and angular velocity data obtained by the inertial measurement unit, and calculate displacement vector and rotation angle of the system; determine the current pose of the transparent OLED display screen relative to the real scene coordinate system based on the displacement vector and the rotation angle through coordinate transformation, the current pose comprising position coordinates and orientation angle; calculate the field of view angle change and perspective projection matrix of the transparent OLED display screen according to the current pose; adjust the display coordinates, size ratio and perspective relationship of the recommended content displayed on the transparent OLED display screen based on the perspective projection matrix, so that the recommended content is spatially registered with the corresponding target object in the real scene.
6. The system of claim 3, wherein, The processor is further configured to: determine the target anchoring position of virtual content in the real scene based on the depth information of each object in the real scene obtained by the depth-of-field camera; calculate the initial display parameters of the virtual content on the display screen according to the relative spatial relationship between the target anchoring position and the transparent OLED display screen; combine the inertial measurement unit data from the mobile chassis to dynamically correct the initial display parameters in real time; perform perspective matching fusion display of the virtual content and the real scene according to the corrected display parameters.
7. The system of claim 6, wherein, The vision sensor further comprises a gesture recognition camera, wherein the processor is further configured to: collect an image sequence of a user's hand through the gesture recognition camera; Extract hand key point spatial coordinates and motion trajectories from the image sequence; Match the hand key point spatial coordinates and motion trajectories with a predefined gesture instruction library to identify the user's operation intention; According to the identified operation intention, generate control instructions for the virtual content of the fusion display and perform corresponding operations.
8. The system of claim 1, wherein, The processor is further configured to: Preprocess the environment image collected by the visual sensor, and extract a scene visual feature vector from the preprocessed environment image; Calculate the similarity of the scene visual feature vector and the standard scene features in the preconfigured scene database; According to the similarity calculation result, determine the scene type with the highest matching degree as the recognition result of the environment scene type.
9. The system of claim 1, wherein, The processor is further configured to: Collect user voice signals through a microphone array and extract acoustic feature parameters of the voice; Input the acoustic feature parameters into a voice emotion classification model to obtain a first emotion recognition result; Collect user face images through a visual sensor and extract facial feature point motion parameters; Input the facial feature point motion parameters into an expression recognition model to obtain a second emotion recognition result; Fuse the first emotion recognition result and the second emotion recognition result to determine the final user emotional state; According to the final user emotional state, select matched recommended content from a content library for display.
10. The system of claim 2, wherein, The processor is further configured to: Continuously collect user chest movement signals through the millimeter wave radar and continuously collect user body movement signals through the pulse ultra-wideband chip; Filter the chest movement signals and body movement signals respectively, extract the respiratory frequency from the chest movement signals, and extract the body movement features from the body movement signals; Analyze the amplitude of the respiratory frequency and the body movement features to determine whether the user is in a stationary sleep state; When the respiratory frequency meets the preset condition and the body movement feature is lower than the threshold for a plurality of consecutive periods, it is determined that the user is in a sleep state; Trigger a system control instruction to switch the system to a low-power consumption mode or turn off the display function of the transparent OLED display screen.
Citation Information
Patent Citations
Advertisement pushing management method and device based on MR glasses and application
CN112181152A
Foreign language learning and homework doing system and method based on education artificial intelligence and integrating interest, real scene and like
CN120496377A