Indoor contextual image system and method
By combining emotion collection equipment and central processing units with presentation equipment, motion and image control instructions are generated, solving the problem of the inability to intelligently perceive user emotions in a home environment, and achieving intelligent adjustment of emotions and improvement of the family atmosphere.
Patent Information
- Application Number
- CN202510649456.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies are unable to intelligently perceive user emotions in a home environment and make corresponding adjustments, resulting in an inability to effectively alleviate loneliness and depression, and an unpleasant family atmosphere.
Emotion collection devices such as smart glasses, smart watches and smart headphones are used to collect user emotional data. Through the central processor and presentation equipment, motion and image control instructions are generated, and automatic guided vehicles and projection devices are used to present context-sensitive images to achieve emotional perception and adjustment.
It realizes intelligent perception of user emotions in the home environment, effectively alleviates loneliness and depression and improves the family atmosphere through rich imaging scenes and flexible interactive experience.
Smart Images

Figure CN120661807A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart home technologies, and in particular to an indoor ambient imaging system and method. Background Art
[0002] In modern life, people living alone at home often experience loneliness due to a lack of sensory stimulation, leading to psychological problems. Depressed individuals may choose to end their lives at home because they find the pain of depression unbearable, and conflicts can create a less-than-pleasant atmosphere in families. However, if external sensory stimulation can mitigate this negative atmosphere and proactively interact with people living alone, loneliness can be alleviated. Depressed individuals often show abnormal signals before committing suicide; providing them with powerful sensory stimulation at this time can avert tragedy. Family conflicts can be unpleasant, and if technology can be used to improve emotions and thus regulate the family atmosphere, family life can be more fulfilling.
[0003] In order to relieve psychological emotions, the invention patent with application number 201911064058.1 "3D digital psychological sandbox implementation method, device, electronic equipment and psychological sandbox" proposes a digital sandbox that is different from the traditional physical sandbox. It integrates virtual reality technology and physical psychological sandbox games to provide an immersive and three-dimensional application scenario.
[0004] However, this device is only a treatment tool for psychologists. It is inconvenient to be part of the home and cannot enter the home. The display screen cannot be integrated with the physical landscape. The interaction process is unnatural. It is only a treatment method and cannot perceive the user's emotions and intelligently change the user's emotions. Summary of the Invention
[0005] In view of this, it is necessary to provide an indoor adaptive imaging system to solve the problem in the existing technology that it is unable to perceive the user's emotions and thus intelligently change the user's emotions.
[0006] In order to solve the above problems, in a first aspect, the present invention provides an indoor ambient imaging system, comprising: An emotion collection device is used to collect emotion data of a target user in a target space and send the emotion data to a central processor. The emotion data includes data collected by smart glasses, data collected by smart watches, and data collected by smart headphones; A central processing unit, in communication with the emotion acquisition device, is configured to receive and generate image control instructions based on the emotion data and a preset material library by performing image fusion, calculate a playback position based on the emotion data, generate motion control instructions based on the playback position, and send the motion control instructions and image control instructions to the presentation device; The presentation device is communicatively connected to the central processing unit and includes a motion device and a projection device. The projection device is arranged on the motion device. The motion device is used to receive and move to a playback position based on motion control instructions. The projection device is used to receive and play ambient image data based on image control instructions.
[0007] In a possible implementation, the motion device includes an automatic guided vehicle, a robotic arm is provided on the automatic guided vehicle, the projection device includes a projector, and a distal end of the robotic arm is fixedly connected to the projector.
[0008] In a possible implementation, the emotion collection device includes smart glasses, a smart watch, and a smart headset; The data collected by the smart glasses include eye tracking data, pupil response data, and visual behavior data of the target user, wherein the eye tracking data includes gaze time and gaze point distribution, the pupil response data includes pupil diameter and blink frequency, and the visual behavior data includes eye movement frequency; The data collected by the smartwatch includes physiological data and psychological data of the target user, wherein the physiological data includes heart rate variability data, skin conductivity data, and skin temperature, and the psychological data includes exercise status data and heart rate data; The data collected by the smart headset includes brain wave data, voice data and environmental noise data of the target user.
[0009] In a possible implementation, the central processing unit is further configured to: Obtaining an emotion prediction value of a target user based on the emotion data; A match is obtained based on the predicted emotion value and the preset recovery value; The target material is obtained based on the matching degree and the preset material library and sent to the rendering device.
[0010] In a possible implementation, the central processing unit is further configured to: obtaining a perceived restorability scale, wherein the perceived restorability scale is used to calculate a recovery value; obtaining an environmental restoration value based on the perceived restoration scale and the emotion data; The environment restoration value is updated to a preset material library.
[0011] In a possible implementation, the preset material library includes: Initialize the material library, cloud material library, personalized material library, and multi-person material library; The perceived restorability scale includes a single-user environment restoration value evaluation dimension table, a psychological sand tray diagnosis and evaluation dimension table, and a social connection value evaluation dimension table. The evaluation dimensions of the single-user environment restoration value evaluation dimension table include attractiveness, sense of escape, coherence, sense of vastness, compatibility, familiarity, and preference. The evaluation dimensions of the psychological sand tray diagnosis and evaluation dimension table include object placement pattern, color selection pattern, interaction frequency, and emotion value change. The evaluation dimensions of the social connection value evaluation dimension table include emotional consistency, interaction participation, and group activation.
[0012] In a second aspect, the present invention further provides an indoor environment-dependent imaging method, the indoor environment-dependent imaging method being based on the indoor environment-dependent imaging system described in any one of the above implementations, the method comprising: Receive and generate image control instructions based on the emotional data and a preset material library, perform position calculation based on the emotional data to obtain the playback position, generate motion control instructions based on the playback position, and send the motion control instructions and the image control instructions to the presentation device so that the presentation device moves to the playback position and plays the image. The emotional data is the data of the target user collected by the emotion collection device in the target space, and the emotional data includes data collected by smart glasses, data collected by smart watches, and data collected by smart headphones.
[0013] In a possible implementation, the method further includes: Obtaining an emotion prediction value of a target user based on the emotion data; A match is obtained based on the predicted emotion value and the preset recovery value; The target material is obtained based on the matching degree and the preset material library and sent to the rendering device.
[0014] In a possible implementation, obtaining the target user's emotion prediction value based on the emotion data includes: Perform weighted summation on the early fusion data, mid-term fusion data, and late fusion data of the emotion data to obtain the emotion prediction value; The matching degree obtained based on the emotion prediction value and the preset recovery value includes: When the emotion prediction value is less than a preset threshold, a matching degree between the emotion prediction value and a preset recovery value is calculated.
[0015] The beneficial effects of the present invention are as follows: the present invention provides an indoor context-sensitive imaging system, including an emotion collection device, a central processing unit, and a presentation device, wherein the emotion collection device is used to collect emotion data of a target user in a target space and send the emotion data to the central processing unit, the central processing unit is in communication with the emotion collection device, and is used to receive and generate motion control instructions and image control instructions based on the emotion data, and send the motion control instructions and image control instructions to the presentation device, while generating motion control instructions and image control instructions to cause the presentation device to move and play images, the presentation device is in communication with the central processing unit, and includes a motion device and a projection device, the projection device is arranged on the motion device, the motion device is used to receive and control the movement of the presentation device based on the motion control instructions, the projection device is used to receive and play context-sensitive image data based on the image control instructions, the presentation device and the emotion collection device are combined to realize the combination of motion control and image control, thereby realizing richer image scenes and more flexible interactive experience. The present invention continuously collects the user's emotion data, forms a strong interactive effect through the combination of motion and image, and then adjusts the image according to the emotion data, thereby realizing intelligent perception of user emotions and intelligently changing user emotions through image playback. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a system architecture diagram of an embodiment of an indoor ambient imaging system provided by the present invention; Figure 2 A schematic diagram showing the structure of a device in accordance with an embodiment of an indoor ambient imaging system provided by the present invention; Figure 3 A single-user workflow diagram of an embodiment of an indoor ambient imaging method provided by the present invention; Figure 4 A flowchart of a single user's emotion assessment method for indoor ambient imaging according to an embodiment of the present invention; Figure 5 A table of evaluation dimensions of a single-user environment restoration value of an embodiment of an indoor environment-specific imaging method provided by the present invention; Figure 6 A diagram of a single-user environment restoration value evaluation model for an embodiment of an indoor environment-specific imaging method provided by the present invention; Figure 7 A flowchart of a single-user psychological sandbox workflow in accordance with an embodiment of an indoor environment-specific imaging method provided by the present invention; Figure 8 A psychological sand table diagnosis and assessment dimension table for an embodiment of an indoor environment-specific imaging method provided by the present invention; Figure 9 A diagram of a psychological sandbox assessment and diagnosis model for an embodiment of an indoor environment-specific imaging method provided by the present invention; Figure 10 A multi-user workflow diagram of an embodiment of an indoor ambient imaging method provided by the present invention; Figure 11 This is a flowchart of multi-user emotion assessment in accordance with an embodiment of an indoor ambient imaging method provided by the present invention; Figure 12 A table of evaluation dimensions of social connection values of multiple users according to an embodiment of the indoor ambient imaging method provided by the present invention; Figure 13 This is a diagram of a multi-user social connection value evaluation model for an embodiment of an indoor ambient imaging method provided by the present invention. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0018] In the description of the embodiments of the present invention, unless otherwise specified, "plurality" means two or more. "And / or" describes the association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone.
[0019] The terms "first," "second," and so on, used in the embodiments of the present invention are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, technical features designated as "first" or "second" may explicitly or implicitly include at least one such feature.
[0020] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0021] Before presenting the embodiments, the following terms are explained.
[0022] Restorative Environment refers to a plant that can help people relieve stress and the various negative emotions that accompany it, reduce mental fatigue, and even promote mental and physical health.
[0023] Sandplay therapy, also known as sandplay therapy, is a psychotherapy method primarily based on play. Originating in Europe and based on the principles of Jungian psychology, it was developed by Dora Kraft. Sandplay therapy involves allowing clients to freely place various miniature models in a specially designed container (the sand tray) filled with fine sand, creating scenarios that the therapist analyzes to understand the client's inner world.
[0024] The present invention provides an indoor ambient imaging system, which is described below.
[0025] Figure 1 This is a system architecture diagram of an embodiment of an indoor ambient imaging system. Figure 1 As shown, the indoor ambient imaging system 100 includes: Emotion collection device 101, used to collect emotional data of the target user in the target space and send the emotional data to the central processor, the emotional data including data collected by smart glasses, data collected by smart watches and data collected by smart headphones; The central processing unit 102 is in communication with the emotion acquisition device, and is configured to receive and generate image control instructions based on the emotion data and a preset material library by performing image fusion, calculate a playback position based on the emotion data, generate motion control instructions based on the playback position, and send the motion control instructions and image control instructions to the presentation device; The presentation device 103 is communicatively connected to the central processing unit and includes a motion device and a projection device. The projection device is arranged on the motion device. The motion device is used to receive and move to a playback position based on motion control instructions. The projection device is used to receive and play ambient image data based on image control instructions.
[0026] It should be noted that the presentation device of the present invention is a robotized device. Specifically, the robotized presentation device first performs scanning and mapping, then establishes a static three-dimensional map within the target space, and then scans in real time during work to complete the construction of the dynamic target, that is, dynamically constructs the three-dimensional map according to the position of the presentation device.
[0027] Compared with the prior art, the present embodiment provides an indoor context-sensitive imaging system, comprising an emotion collection device, a central processing unit, and a presentation device. The emotion collection device is used to collect emotion data of a target user in a target space and send the emotion data to the central processing unit. The central processing unit is in communication with the emotion collection device and is used to receive and generate motion control instructions and image control instructions based on the emotion data, and send the motion control instructions and image control instructions to the presentation device, while generating motion control instructions and image control instructions to cause the presentation device to move and play images. The presentation device is in communication with the central processing unit and comprises a motion device and a projection device. The projection device is arranged on the motion device and is used to receive and control the movement of the presentation device based on the motion control instructions. The projection device is used to receive and play context-sensitive image data based on the image control instructions. The presentation device and the emotion collection device are combined to realize the combination of motion control and image control, thereby achieving richer image scenes and more flexible interactive experience. The present invention continuously collects the user's emotion data, forms a strong interactive effect through the combination of motion and image, and then adjusts the image according to the emotion data, thereby realizing intelligent perception of user emotions and intelligently changing user emotions through image playback.
[0028] In some embodiments of the present invention, Figure 2 As shown, the motion device includes an automatic guided vehicle, a robotic arm is provided on the automatic guided vehicle, the projection device includes a projector, and the end of the robotic arm is fixedly connected to the projector.
[0029] In some embodiments of the present invention, the emotion collection device 101 includes smart glasses, smart watches, and smart headphones; The data collected by the smart glasses include eye tracking data, pupil response data, and visual behavior data of the target user, wherein the eye tracking data includes gaze time and gaze point distribution, the pupil response data includes pupil diameter and blink frequency, and the visual behavior data includes eye movement frequency; The data collected by the smartwatch includes physiological data and psychological data of the target user, wherein the physiological data includes heart rate variability data, skin conductivity data, and skin temperature, and the psychological data includes exercise status data and heart rate data; The data collected by the smart headset includes the target user's brain wave data, voice data and environmental noise data.
[0030] In a specific embodiment of the present invention, smart glasses are mainly used to collect data on three aspects: eye tracking, pupil reaction, and visual behavior: 1. Eye tracking mainly includes the following data: Collect fixation duration. It is generally believed that in a high recovery value environment, the fixation duration will be longer, with a value of A1. Gaze Dispersion is collected. It is generally believed that in a high recovery value environment, the focus is concentrated and the attention is restored, while in a low recovery value environment, the focus is dispersed and the attention is scattered, with a value of A2; 2. Pupil response mainly includes the following data: Collect pupil diameter (Pupil Dilation). It is generally believed that in a high recovery value environment, the pupil is in a contracted state, and in a low recovery value environment, the pupil is in a dilated state, with a value of A3; Collect blink rate. It is generally believed that in a high recovery value environment, the blink rate is lower, with a value of A4. 3. Visual behavior mainly includes the following data: By detecting saccades, it is generally believed that in a high recovery value environment, there will be fewer rapid saccades, with a value of A5; The data collected in real time by smart glasses generates data set D1.
[0031] Smart watches are mainly used to collect physiological and psychological signals: 1. Physiological signals mainly include the following data Collect HRV (heart rate variability) data. It is generally believed that in a high recovery value environment, the HRV value will increase, and the value is B1; Collect EDA (skin conductivity) data. It is generally believed that in a high recovery value environment, the EDA value will decrease to B2; Collect skin temperature. It is generally believed that the skin temperature will remain constant in a high recovery value environment. However, when encountering abnormal conditions such as pressure, the skin temperature will rise and the value will be B3.
[0032] 2. Psychological signals mainly include the following data The accelerometer (IMU) data is collected, mainly used to detect movement or stillness. If the stillness is too long, it may be in a relaxed state, and the value is B4; The number of steps and activity level are collected to mainly assess the amount of exercise. If the amount of exercise is appropriate and the heart rate is stable, the user may be in a high recovery value environment, with a value of B5.
[0033] Physiological and behavioral signals are collected through smart watches to generate data set D2.
[0034] 3. Smart headphones Mainly used to collect brain waves (α / θ waves), voice emotion information and environmental noise 1. Brain waves mainly include the following data: Collecting brainwave data, it is generally believed that in a high-restoration environment, α / θ waves will become stronger, with a value of C1; 2. Voice emotion mainly includes the following data: Collect speech data, extract acoustic feature data and semantic feature data, the value is C2; 3. Environmental noise Collect environmental noise data to evaluate the environmental noise level, with a value of C3; The data collected by the smart headset generates dataset D3.
[0035] The above-mentioned acquisition devices ultimately form three data sets D1 (smart glasses), D2 (smart watches), and D3 (smart headphones). In addition, the above three devices can provide user GPS information for locating the user's position in the indoor three-dimensional real-life map.
[0036] In some embodiments of the present invention, the central processor 102 is further configured to: Obtaining an emotion prediction value of a target user based on the emotion data; A match is obtained based on the predicted emotion value and the preset recovery value; The target material is obtained based on the matching degree and the preset material library and sent to the rendering device.
[0037] In a specific embodiment of the present invention, the preset material library includes: an initial material library, a cloud material library, a personalized material library, and a multi-person material library. The preset material library, namely the context-sensitive image material database, is mainly based on user habits and stored according to user accounts. It optimizes the recovery value of the environment for different users and mainly includes the following categories: (1) Initialize the material library, store it locally, and include image data sets, motion control data sets, and image generation models (using image data and motion control data to generate context-sensitive images). Based on the environmental restoration value theory, it includes context-sensitive image materials that are generally considered to have high restoration values. Image data is mainly divided into the following three categories: static images, dynamic videos, and music data. Static images are divided into overall images and common environmental elements, such as trees, grass, waterfalls, rivers, lakes, etc. There is a motion control data set corresponding to the image materials, and the image data and motion control data are in correspondence; (2) Cloud material library, which is stored in the cloud. The service provider regularly updates the image data, motion control data and image generation model in it. The cloud material library is initialized by downloading; (3) Personalized material library, stored locally, including optimized image data sets, optimized motion control data sets, optimized image generation models and optimization models. Based on the initialization material library and the system's evaluation of the environmental restoration theory, the optimized material library includes image data and motion control data, as well as data formed based on user habits. Based on the user's usage data (emotional response after watching the ambient image), Generative Artificial Intelligence (GenAI) optimizes and adjusts (including adjusting the original data sets and generating new image elements) the image data sets, motion control data sets and image generation models to achieve the best restoration value for the user to the environment.
[0038] (4) Multi-person material library, stored locally, includes multi-person optimized image datasets, multi-person optimized motion control datasets, multi-person image generation models and multi-person optimization models. The multi-person optimized image datasets mainly include multi-person image datasets (mainly group environmental elements) and corresponding multi-person optimized motion control datasets. Images that improve group atmosphere are generated through the multi-person image generation model, and the optimization model is also included.
[0039] The image elements in the multi-person library and the personalized library are mainly different in the following ways: Personalized material library is for individual users, and multi-person material library is for group users; Sense of spatial scale: Personalized image data has a smaller spatial scale, showing a certain degree of privacy, while multi-person image data has a larger spatial scale, showing a space that can accommodate multiple people; Privacy: Personalized image data prefers a low-interference environment, while multi-person image data has lower privacy and a "visual social safety zone"; Social attributes: Personalized image data generally does not contain elements with social attributes, while multi-person image data does contain elements with social attributes. Elements with social attributes include shared facilities, group activity venues, etc. Sensory stimulation and personalized image data emphasize simplicity and quietness, while multi-person image data needs to avoid a sense of crowding.
[0040] In a specific embodiment of the present invention, the central controller mainly performs computing tasks, and generating control instructions mainly includes: (1) Read the emotion data from the emotion collection device, calculate the emotion prediction value and the true emotion value, and calculate the environment restoration value based on the ambient image data (image data set and motion control data set) and emotion data; (2) Based on the emotion value, the material data (image data set and motion control data set) in the material library and the image construction model are read to form the context image. After that, the multi-projection image fusion calculation and robot scheduling calculation tasks are performed to generate image control instructions and motion control instructions, which are sent to the "AGV + multi-degree-of-freedom robotic arm + projector" presentation device; Image control commands mainly include image content control commands and multi-screen edge fusion control commands; (3) Use AI models to diagnose user psychological sandbox interaction data, generate diagnostic results, generate contextual image data, and generate image control instructions and motion control instructions; (4) Read the optimization model, and perform optimization calculations on the model data and generate new materials based on the optimization data using Generative Artificial Intelligence (GenAI), and store the relevant results in a personalized material library.
[0041] In some embodiments of the present invention, the central processing unit is further configured to: obtaining a perceived restorability scale, wherein the perceived restorability scale is used to calculate a recovery value; obtaining an environmental restoration value based on the perceived restoration scale and the emotion data; The environment restoration value is updated to a preset material library.
[0042] In some embodiments of the present invention, the preset material library includes: Initialize the material library, cloud material library, personalized material library, and multi-person material library; The perceived restorability scale includes a single-user environment restoration value evaluation dimension table, a psychological sand tray diagnosis and evaluation dimension table, and a social connection value evaluation dimension table. The evaluation dimensions of the single-user environment restoration value evaluation dimension table include attractiveness, sense of escape, coherence, sense of vastness, compatibility, familiarity, and preference. The evaluation dimensions of the psychological sand tray diagnosis and evaluation dimension table include object placement pattern, color selection pattern, interaction frequency, and emotion value change. The evaluation dimensions of the social connection value evaluation dimension table include emotional consistency, interaction participation, and group activation.
[0043] In a second aspect, the present invention further provides an indoor environment-dependent imaging method, the indoor environment-dependent imaging method being based on the indoor environment-dependent imaging system described in any one of the above implementations, the method comprising: Receive and generate image control instructions based on the emotional data and a preset material library, perform position calculation based on the emotional data to obtain the playback position, generate motion control instructions based on the playback position, and send the motion control instructions and the image control instructions to the presentation device so that the presentation device moves to the playback position and plays the image. The emotional data is the data of the target user collected by the emotion collection device in the target space, and the emotional data includes data collected by smart glasses, data collected by smart watches, and data collected by smart headphones.
[0044] In some embodiments of the present invention, further comprising: Obtaining an emotion prediction value of a target user based on the emotion data; A match is obtained based on the predicted emotion value and the preset recovery value; The target material is obtained based on the matching degree and the preset material library and sent to the rendering device.
[0045] In a specific embodiment of the present invention, obtaining the target user's emotion prediction value based on the emotion data includes: Perform weighted summation on the early fusion data, mid-term fusion data, and late fusion data of the emotion data to obtain the emotion prediction value; The matching degree obtained based on the emotion prediction value and the preset recovery value includes: When the emotion prediction value is less than a preset threshold, a matching degree between the emotion prediction value and a preset recovery value is calculated.
[0046] In a specific embodiment of the present invention, the working modes of the indoor ambient imaging system include single-user mode, psychological sandbox mode and multi-user mode. Figure 3 FIG. 1 is a flow chart of a method for indoor ambient imaging for a single user. When the user is a single user, the method includes the following steps: Step 1: Collect home map information through three-dimensional map collection and store it in the memory.
[0047] Step 2: The emotion collection device (smart glasses, smart watches, smart headphones) starts working and generates data sets D1, D2, and D3 in real time, ensuring that the time series of the three data sets are consistent; Step 3: Data sets D1, D2, and D3 are sent to the central controller via Bluetooth / WIFI communication technology. The central controller processes the data sets and finally obtains the predicted value of user emotions. ,like Figure 4 The figure shows the single user emotion assessment flow chart, which is mainly divided into the following steps: Check the data set, mainly including data integrity check, data time synchronization check, and data missing check; 1. The data integrity check mainly detects whether there is any data set loss. The present invention supports the situation where one collection device has lost data, that is, the present invention requires two or more emotion collection devices, which is mainly divided into the following three situations: Missing datasets (device loss) are mainly predicted through cross-modal prediction using Bayesian inference; D1 dataset is missing (smart glasses are missing): D2 dataset + D3 dataset is used to infer D1 dataset; D2 dataset is missing (smartwatch is missing): use D1 dataset + D3 dataset to infer D2 dataset; D3 dataset is missing (smart headset is missing): use D1 dataset + D2 dataset to infer D3 dataset; After predicting the data set, we get the complete data set: D1, D2, D3; 2. Data time synchronization check mainly refers to: Check the timestamps of datasets D1, D2, and D3. If they are inconsistent, use interpolation to align them. 3. Data missing check mainly refers to: Use NaN to fill in missing data and use interpolation to fill in missing data.
[0048] Preprocess the dataset data and denoise the data in D1, D2, and D3. For example, use the Butterworth filter algorithm for HRV / EEG data, the Gaussian filter algorithm for eye tracking data, and the Mel Spectrum Convolutional Cell (MFCC) algorithm for speech data. Then normalize the three datasets. Data fusion is mainly divided into three steps: early fusion, mid-term fusion and late fusion: Early fusion: All dataset data are fused into one large feature vector: F=[A1;A2;A3;A4;A5;B1;B2;B3;B4;B5;C1;C2;C3] Mid-term fusion: Data from different sensors are fused through different neural networks: LSTM processes A1-A5; CNN processes B1-B5; Transformer processes C1-C3 Late fusion Through the individual emotion prediction values of different sensor data, we can get , , ; The final emotion prediction value is obtained by weighting the three emotion values:
[0049] Step 4: Get the sentiment prediction value After that, if the predicted emotion value is lower than the threshold, it is matched with the recovered value S and the predicted emotion value is calculated. The matching degree with the recovery value S is obtained to obtain the matching degree M( ), the matching degree M=1 represents the environmental recovery value S and the emotion prediction value Complete match, matching degree M=0 represents the environmental recovery value S and the emotion prediction value Completely mismatched, the material generation is completed based on the matching degree until the material meets the matching threshold setting; Step 5: The generated material is sent to the "AGV + multi-degree-of-freedom robotic arm + projector" presentation device. The presentation device automatically calculates the optimal playback position based on the user's location information and presents contextual image information, including static images (soothing natural landscapes), dynamic landscapes (waves, flames), and accompanied by music information.
[0050] Step 6: During the playback of the ambient image, repeat steps 1 and 2 at a certain frequency to obtain the actual emotion value E. The playback process is dynamically adjusted according to the actual emotion value until the user's emotion returns to the normal threshold and the environmental recovery value of the material library is optimized. The specific workflow for optimizing the environmental recovery value of the material library is as follows: like Figure 5 The following table shows the environmental restoration value evaluation dimension table for a single user. After starting to play the ambient image, the central controller automatically generates the Perceived Restorativeness Scale (PRS) (based on the Italian version of the Perceived Restorativeness Scale), also known as the environmental restoration value evaluation dimension table. This table evaluates environmental restoration value using seven dimensions and 29 sub-items. The main contents are as follows: Fascination: This mainly assesses whether the environment can attract attention without subjective effort, and contains 7 items; Being-away: This measure primarily assesses the ability of a person to escape from their daily life and consists of six items. Coherence: This assesses the orderliness of the environment. The order of the environment makes people feel that the physical layout is well-organized. It contains four items. Scope: This mainly assesses the temporal and spatial extension of the environment and includes three items; Compatibility: This mainly assesses the match between the environment and personal goals and contains 6 items; Familiarity: This mainly assesses the person's familiarity with the environment, which is generally related to age and contains 1 item; Preference: It mainly assesses the preference of the person for a certain type of environment and contains 2 items.
[0051] The above five aspects are designed with seven dimensions and 29 quantitative indicators; Based on the data sets D1, D2, and D3 obtained from real-time monitoring, the above 29 quantitative indicators are comprehensively calculated, and finally the recovery values RV1, RV2, RV3, RV4, RV5, RV6, and RV7 of the above seven dimensions are obtained. The environmental recovery value RV is obtained by comprehensively weighting the recovery values of the above seven dimensions.
[0052] like Figure 6 The figure shows a single-user environment recovery value evaluation model. Environments with high recovery values (RV) will enter the personalized material library and provide feedback to optimize the subsequent generation of environment recovery values.
[0053] In the above method, if Figure 7 The figure shows the workflow of the indoor single-user psychological sandbox with contextual images. Based on the predicted value of the user's emotions, the "AGV + multi-degree-of-freedom robotic arm + projector" presentation device plays a personalized psychological sandbox environment, which includes the following three categories: natural environment, indoor image, and imaginary world; The "AGV + multi-DOF robotic arm + projector" system detects human movements through sensors, allowing users to drag models or independently build models in the surrounding image, or interact through extended interactive tools such as smart brushes and smart gloves. Emotion collection devices record user data, identify changes in user emotions, and send them to Generative Artificial Intelligence (GenAI); like Figure 8 The following table shows the dimension table of psychological sand table diagnosis and assessment. User behavior data mainly includes the following dimensions: Item placement pattern: tends to be messy or neat; Color selection mode: selection of light and dark tones; Frequency of interaction: frequent or infrequent; Emotion value changes: based on emotion collection data sets D1, D2, and D3; like Figure 9 The picture shows a psychological sandbox assessment and diagnosis model. Based on the data fed back by AI, the environment is adjusted in a timely manner, and the adjusted environment image information is played.
[0054] In addition to the above process, this system can also provide regular remote diagnosis and treatment for depression and autism patients by psychologists, combined with the psychological sandbox function of the ambient imaging system. At this time, the indoor ambient imaging system simultaneously presents the psychologist's video image and the psychological sandbox interactive interface, and completes the psychological sandbox treatment under the guidance of the psychologist.
[0055] In a specific embodiment of the present invention, Figure 10 The following is a workflow diagram for multi-user mode. When there are multiple users, the context-sensitive imaging method includes the following steps: Step 1: Based on the fusion of multiple data sets such as person P1 (D1, D2, D3) and P2 (D1, D2, D3), the group emotion prediction value is obtained. Based on the first working method, one of the GPS modules in smart glasses, smart watches, and smart headphones is used to obtain multi-user spatial positioning data D4; Step 2: Obtain the group emotion prediction value based on the acquired data, and obtain the group emotion map based on the acquired data (individual physiological data D1, D2, D3 and group spatial data D4). The specific steps are as follows: Temporal alignment of group data; Perform three-dimensional reproduction of multi-person spatial trajectories and fill in the multi-person behavior interaction matrix: Each row and column of the interaction matrix represents a user Ui; Each element Mij in the matrix represents the interaction value between users, calculated based on individual data differences and spatial trajectory data. Individual data differences mainly include eye contact (based on smart glasses data D1), conversation frequency (based on smart headset data D3), emotional synchronization (based on individual emotional values), and collaborative action evaluation (based on smart watch data D2). A group emotion network is constructed based on the interaction matrix and individual data sets to obtain a group emotion map.
[0056] Step 3: Based on the group emotional map, play contextual images that improve the group atmosphere Step 4: If Figure 12 The following table shows the Social Connectedness Scale (SCS) for evaluating the social connectivity of multiple users. The table is used to evaluate the effect of images on a group of users: Based on the UCLA version of the Social Connectedness Scale (SCS), combined with automated detection requirements, the group atmosphere is assessed as follows: The dataset for smart glasses detection (D1) is mainly used to focus on whether the focus is shared and the degree of tension / concentration; The smartwatch detection dataset (D2) is mainly used to evaluate group emotional synchronization (excitement and stress); The dataset for smart headset detection (D3) is mainly used to evaluate group intonation consistency and speaking turns; Based on the above data, we get the following 9 indicators in three dimensions: A. Emotional Consistency: Assessing the degree of emotional consistency among individuals in a group, including individual emotion distribution variance, heart rate variability similarity, and voice emotion similarity; B. Engagement: This measures the degree of engagement of individuals in a group in communication and collaboration, including turn-taking, eye contact (focus), and physical activity synchronization. C. Arousal Level: Assessing the overall level of relaxation, activation, and tension in the group, including average EDA, average pupil dilation, and fluctuations in speech rate and volume (based on speech emotion recognition); According to the above method, the group environmental recovery value is obtained, such as Figure 13 The figure shows a multi-user social connection value evaluation model, which optimizes the multi-person material library based on the recovery value.
[0057] Those skilled in the art will appreciate that all or part of the process steps of the above-described embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, such as a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0058] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.
Claims
1. An indoor ambient imaging system, characterized in that: include: An emotion collection device is used to collect emotion data of a target user in a target space and send the emotion data to a central processor. The emotion data includes data collected by smart glasses, data collected by smart watches, and data collected by smart headphones; A central processing unit, in communication with the emotion acquisition device, is configured to receive and generate image control instructions based on the emotion data and a preset material library by performing image fusion, calculate a playback position based on the emotion data, generate motion control instructions based on the playback position, and send the motion control instructions and image control instructions to the presentation device; The presentation device is communicatively connected to the central processing unit and includes a motion device and a projection device. The projection device is arranged on the motion device. The motion device is used to receive and move to a playback position based on motion control instructions. The projection device is used to receive and play ambient image data based on image control instructions.
2. The indoor ambient imaging system according to claim 1, characterized in that: The motion device includes an automatic guided vehicle, a mechanical arm is provided on the automatic guided vehicle, the projection device includes a projector, and the end of the mechanical arm is fixedly connected to the projector.
3. The indoor ambient imaging system according to claim 1, characterized in that: The emotion collection devices include smart glasses, smart watches and smart headphones; The data collected by the smart glasses include eye tracking data, pupil response data, and visual behavior data of the target user, wherein the eye tracking data includes gaze time and gaze point distribution, the pupil response data includes pupil diameter and blink frequency, and the visual behavior data includes eye movement frequency; The data collected by the smartwatch includes physiological data and psychological data of the target user, wherein the physiological data includes heart rate variability data, skin conductivity data, and skin temperature, and the psychological data includes exercise status data and heart rate data; The data collected by the smart headset includes brain wave data, voice data and environmental noise data of the target user.
4. The indoor ambient imaging system according to claim 1, characterized in that: The central processing unit is also used to: Obtaining an emotion prediction value of a target user based on the emotion data; A match is obtained based on the predicted emotion value and the preset recovery value; The target material is obtained based on the matching degree and the preset material library and sent to the rendering device.
5. The indoor ambient imaging system according to claim 4, characterized in that: The central processing unit is also used to: obtaining a perceived restorability scale, wherein the perceived restorability scale is used to calculate a recovery value; obtaining an environmental restoration value based on the perceived restoration scale and the emotion data; The environment restoration value is updated to a preset material library.
6. The indoor ambient imaging system according to claim 1, characterized in that: The preset material library includes: Initialize the material library, cloud material library, personalized material library, and multi-person material library; The perceived restorability scale includes a single-user environment restoration value evaluation dimension table, a psychological sand tray diagnosis and evaluation dimension table, and a social connection value evaluation dimension table. The evaluation dimensions of the single-user environment restoration value evaluation dimension table include attractiveness, sense of escape, coherence, sense of vastness, compatibility, familiarity, and preference. The evaluation dimensions of the psychological sand tray diagnosis and evaluation dimension table include object placement pattern, color selection pattern, interaction frequency, and emotion value change. The evaluation dimensions of the social connection value evaluation dimension table include emotional consistency, interaction participation, and group activation.
7. The indoor ambient imaging system according to claim 4, characterized in that: The presentation device is also provided with a detection device, which includes a first sensor and a second sensor. The first sensor is used to detect the movement information of the target user in the target space to realize dragging and autonomous modeling in the ambient image. The second sensor is used to obtain the current position of the presentation device and build a three-dimensional map based on the current position. It is also used to obtain the position information in front of the travel direction and adjust the route based on the position information to bypass obstacles.
8. An indoor environment-dependent imaging method, applied to the indoor environment-dependent imaging system according to any one of claims 1 to 6, characterized in that: The method comprises: Receive and generate image control instructions based on the emotional data and a preset material library, perform position calculation based on the emotional data to obtain the playback position, generate motion control instructions based on the playback position, and send the motion control instructions and the image control instructions to the presentation device so that the presentation device moves to the playback position and plays the image. The emotional data is the data of the target user collected by the emotion collection device in the target space, and the emotional data includes data collected by smart glasses, data collected by smart watches, and data collected by smart headphones.
9. The indoor environment-specific imaging method according to claim 8, characterized in that: Also includes: Obtaining an emotion prediction value of a target user based on the emotion data; A match is obtained based on the predicted emotion value and the preset recovery value; The target material is obtained based on the matching degree and the preset material library and sent to the rendering device.
10. The indoor environment-specific imaging method according to claim 9, characterized in that: The step of obtaining a target user's emotion prediction value based on the emotion data includes: Perform weighted summation on the early fusion data, mid-term fusion data, and late fusion data of the emotion data to obtain the emotion prediction value; The matching degree obtained based on the emotion prediction value and the preset recovery value includes: When the emotion prediction value is less than a preset threshold, a matching degree between the emotion prediction value and a preset recovery value is calculated.
Citation Information
Patent Citations
Method and device for realizing 3D digital psychological sand table, and electronic equipment as well as psychological sand table
CN110721387A