Cooperative monitoring and man-machine interaction system and method fusing multi-modal perception
By combining wearable devices and sweeping robots, multimodal perception technology and human behavior network model are used to solve the privacy and energy consumption problems existing in traditional monitoring devices, and accurately identify and monitor human behavior, ensuring user safety.
Patent Information
- Application Number
- CN202510113167.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, traditional home surveillance cameras have privacy problems, while traditional wearable devices are difficult to achieve long-term continuous monitoring due to energy consumption limitations, which cannot meet people's needs for long-term monitoring of daily behavior.
By combining wearable devices and sweeping robots, multimodal perception technology, including sensor data on location, objects, sound and motion, the human behavior network model is used for real-time analysis and identification, and emergency information is sent in a timely manner through the alarm module.
It realizes accurate identification and monitoring of human behavior, saves the energy of wearable devices, improves the accuracy of behavior recognition, and promptly detects abnormal behaviors to ensure user safety.
Smart Images

Figure CN120029457A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of behavior monitoring technology, and in particular to a collaborative monitoring and human-computer interaction system and method integrating multi-modal perception. Background Art
[0002] As people's health awareness increases, monitoring of daily human behavior has gradually become a trend. However, there are some problems with the current common monitoring solutions:
[0003] 1) Traditional home surveillance cameras: Although they can obtain relevant information, they have serious privacy issues, making them difficult to be widely accepted by the public;
[0004] 2) Traditional wearable devices: Due to energy consumption limitations, it is difficult to achieve long-term continuous monitoring and cannot meet people's needs for long-term monitoring of daily behavior.
[0005] Mobile or wearable devices such as smartphones and smart watches are widely used in daily life and can be used to collect data related to daily human behavior. For example, smartphones are equipped with monitoring devices such as cameras, microphones, and accelerometers, which can capture multimodal data for daily human behavior recognition. In addition to collecting vital sign data such as heart rate and blood oxygen level, smart watches can also use these data for human behavior recognition.
[0006] However, when wearable devices need to use complex algorithms (such as deep neural networks) to process large amounts of visual and audio data, insufficient computing power and excessive energy consumption become major obstacles. Therefore, to solve these problems, wearable devices need to work together with devices with more powerful computing resources to complete recognition tasks.
[0007] It is worth noting that sweeping robots play an important role in family life. They not only have the function of cleaning, but also provide a variety of services such as entertainment and event reminders. In addition, more and more sweeping robots have high computing power and can perform complex artificial intelligence reasoning.
[0008] To sum up, combining wearable devices with sweeping robots to monitor people's daily activities may be a feasible and effective solution. Summary of the invention
[0009] In view of the needs and deficiencies of current technological development, the present invention provides a collaborative monitoring and human-computer interaction system and method integrating multimodal perception.
[0010] In the first aspect, the present invention provides a collaborative monitoring and human-computer interaction system integrating multimodal perception, and the technical solution adopted to solve the above technical problems is as follows:
[0011] A collaborative monitoring and human-computer interaction system integrating multimodal perception, comprising:
[0012] Wearable devices interact with the sweeping robot and are equipped with multiple sensors to collect multimodal sensor data covering location, objects, sound, and motion in the environment. They are also used to receive instructions sent by the wearable device and dynamically adjust the working modes of multiple sensors.
[0013] The sweeping robot interacts with the wearable device and has a built-in human behavior network model and alarm module. It is used to analyze the data collected by the wearable device in real time and identify the user's human behavior through the human behavior network model, determine the sensor usage mode corresponding to the identification result based on the preset rule set, and send execution commands to the wearable device;
[0014] The human behavior network model is built into the sweeping robot. It is used to periodically obtain the user's historical daily human behavior data accumulated by the wearable device through the sweeping robot, learn the order constraints, the time pattern of the behavior, and the duration information from the acquired data, and is also used to integrate the multimodal sensor data collected by the wearable device in real time and the information learned from the historical behavior data. By analyzing the integrated data, the user's real-time human behavior can be accurately identified;
[0015] The alarm module is built into the sweeping robot and is used to send an alarm to a preset emergency contact in a timely manner when the human behavior network model identifies the user's abnormal human behavior according to the preset abnormal behavior rules;
[0016] The human-computer interaction module provides a human-computer interaction interface for web pages and mobile APPs, which is used for family members to log in to the human-computer interaction interface and display the user's human behavior information in real time.
[0017] The wearable device includes a main control board, a power module and a shell. The main control board is built with a Raspberry Pi Zero computer as the core, and integrates multiple sensors and a Raspberry Pi camera to collect multimodal data and provide data support for human behavior recognition. The multiple sensors include accelerometers, gyroscopes, microphones and cameras. The accelerometers and gyroscopes are used to detect motion data, the microphones are used to collect sound data, and the cameras are used to collect object and scene image data. The wearable device can fully perceive the user's environment and its own status through the above sensors.
[0018] The wearable device communicates with the sweeping robot via a WiFi module.
[0019] The human behavior network model includes a position recognition sub-model, an object recognition sub-model, a sound event recognition sub-model, a body movement recognition sub-model and an integrated analysis module, wherein:
[0020] The location recognition sub-model is used to receive multi-modal sensor data covering location collected by wearable devices, generate probability distributions of different locations, and accurately identify the specific location of the human body;
[0021] The object recognition sub-model is used to receive multimodal sensor data covering objects collected by the wearable device, detect multiple objects in the image, and generate the probability of each object to achieve accurate recognition of specific objects;
[0022] The sound event recognition sub-model is used to receive multi-modal sensor data covering sound collected by wearable devices to accurately recognize various sound events;
[0023] The body motion recognition sub-model is used to receive multi-modal sensor data covering motion collected by wearable devices to achieve accurate recognition of body motions;
[0024] The integrated analysis module is used to collect the position information output by the position recognition sub-model, the object information recognized by the target recognition sub-model, the sound event information detected by the sound event recognition sub-model, and the body movement information recognized by the body movement recognition sub-model, and use rule-based reasoning or classification algorithms in machine learning to fuse the aforementioned multi-source information, explore the intrinsic connections and relationships between the multi-source information, and output accurate recognition results of real-time human behavior.
[0025] The location recognition model is based on the pre-trained Xception model, which has excellent feature extraction capabilities and can achieve a top five verification accuracy of 0.945 and classify 1000 different categories. On the basis of the basic model, a final dense layer with a softmax activation function is added to form a location recognition model.
[0026] The location recognition model is trained using historically accumulated multimodal sensor data on locations, so that the location recognition model can generate probability distributions of different locations and ultimately achieve accurate recognition of specific locations.
[0027] The object recognition model is based on the CNN-based real-time target detection method YOLO as the basic architecture, which has the ability to detect multiple targets at the same time and can effectively detect objects; based on YOLO, its pre-trained model is used to form an object recognition model;
[0028] The object recognition model is trained using historically accumulated multimodal sensor data on objects, so that the object recognition model can simultaneously detect multiple objects in the image, generate the probability of each object, and achieve accurate recognition of specific objects.
[0029] The sound event recognition model is based on a CNN-based network architecture, which has the potential to process sound feature data to identify events;
[0030] The Mel-frequency cepstral coefficients are used as sound features, and the sound sampling duration is set to 1 second, the sampling rate is set to 32000 Hz, the fast Fourier transform window size is set to 2048, the step size is set to 1024, and the number of Mel frequency bands is set to 64. Under the above parameter settings, feature data with a shape of 64×32 is generated. Based on the generated feature data, a CNN network consisting of three convolutional layers and two dense layers is constructed, and the output of each convolutional layer is processed by batch normalization and ReLU activation function to form a sound event recognition model, in which the convolutional layer is used to extract local features in the sound feature data. The convolution kernel size of each convolution layer is 3×3, and the dimensions of the three convolution layers are 64, 128, and 256 respectively. Each convolution layer is followed by a maximum pooling layer, which is used to downsample the feature map output by the convolution layer. The size of the pooling kernel of the maximum pooling layer is 2×2. The dimensions of the two dense layers are 256 and 6 respectively. The 256 neurons in the first dense layer are used to further integrate and abstract the features extracted by the previous convolution layer and pooling layer. The 6 neurons in the latter dense layer correspond to the 6 different sound event categories that need to be identified, and the probability corresponding to each category is output through the softmax activation function.
[0031] The sound event recognition model is trained using historically accumulated multimodal sensor data on sound. The model parameters are continuously adjusted to enable the model to better fit the training data. After trying different numbers of layers and layer dimensions, the model structure with the least parameters and the best performance is determined, so that the sound event recognition model can accurately identify various types of sound events by processing the generated feature data through a CNN network.
[0032] The body motion recognition model is based on a CNN model with two convolutional layers and two dense layers. The basic architecture has the ability to extract features from acceleration data and perform motion recognition, and can effectively learn body motion features. Using 2 seconds of acceleration data as a sample, the convolutional layer kernel size is set to 2×2, and the dimensions are set to 16 and 32 respectively, based on which a body motion recognition model is formed. In the training process of the body motion recognition model, the Adam algorithm is used to optimize the model, and the dropout method is used to prevent the model from overfitting.
[0033] The body motion recognition model is trained using historically accumulated multimodal sensor data on motion, so that it can process 2 seconds of acceleration data, extract features from samples, learn features through convolutional and dense layers, and ultimately achieve accurate recognition of walking movements.
[0034] In the second aspect, the present invention provides a collaborative monitoring and human-computer interaction method integrating multimodal perception, and the technical solution adopted to solve the above technical problems is as follows:
[0035] A method for implementing a collaborative monitoring and human-computer interaction system integrating multimodal perception is based on the system described in the first aspect, and the specific implementation process includes the following steps:
[0036] S1. Collect multimodal sensor data covering location, objects, sound, and motion through wearable devices;
[0037] S2, the sweeping robot interacts with the wearable device and receives multimodal sensor data transmitted by the wearable device;
[0038] The sweeping robot has a built-in human behavior network model. The human behavior network model periodically obtains the user's historical daily human behavior data accumulated by wearable devices through the sweeping robot, and learns the order constraints, time rules of behavior occurrence, and duration information from these data;
[0039] At the same time, the human behavior network model integrates the multimodal sensor data collected by wearable devices in real time based on the information learned from historical behavior data to analyze and identify the user's human behavior;
[0040] S3, the sweeping robot determines which sensor mode to use based on the recognition result of the human behavior network model and the preset rule set, and sends an execution command to the wearable device to realize dynamic adjustment of the working mode of the wearable device sensor;
[0041] S4. When the human behavior network model identifies abnormal human behavior of the user, an alarm message is sent to a preset emergency contact in a timely manner through the alarm module;
[0042] S5. The human-computer interaction module provides a human-computer interaction interface for web pages and mobile APPs. Family members log in to the human-computer interaction interface to display the user's human behavior information in real time, thus realizing information interaction between people and the system.
[0043] The collaborative monitoring and human-computer interaction system and method integrating multimodal perception of the present invention has the following beneficial effects compared with the prior art:
[0044] 1. The present invention combines wearable devices and household sweeping robots to realize human behavior recognition and monitoring through a human behavior network model, and effectively saves the energy of wearable devices;
[0045] 2. The present invention can monitor human behavior from multiple dimensions by integrating multi-modal sensor data such as position, object, sound and motion, which greatly improves the accuracy of behavior recognition compared with single-modal monitoring; the sweeping robot receives data from the wearable device in real time, and quickly identifies the user behavior through the human behavior network model. When it is identified that the user behavior has changed, the wearable device sensor mode can be immediately determined according to the result to achieve dynamic adjustment, which can not only obtain key information in time, but also effectively save energy and extend the battery life of the wearable device; the user's abnormal human behavior can be discovered in time, and an alarm message can be sent to a preset emergency contact through the alarm module, which provides important support for ensuring the user's safety; the human-computer interaction module provides a human-computer interaction interface, and family members can easily log in to view the user's human behavior. This design makes information acquisition more convenient. No matter where the family members are, they can understand the user's situation at any time through mobile phones or computers;
[0046] 3. The present invention can be applied to the field of smart homes to realize intelligent monitoring and management of daily behaviors of family members; it can be applied to the field of medical health to realize health monitoring of patients or the elderly; it can be applied to the field of smart elderly care to comprehensively guarantee the safety and health of the elderly; it can be applied to the field of smart education to monitor students' classroom behaviors and learning status. For special education students, such as children with autism, the system can help teachers and parents better understand their behavioral characteristics and needs by monitoring their behavioral patterns, and provide more targeted education and rehabilitation training. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Attached Figure 1 It is a module architecture diagram of the first embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the technical solution, the technical problem solved and the technical effect of the present invention more clearly understood, the technical solution of the present invention is clearly and completely described below in conjunction with specific embodiments.
[0049] Embodiment 1:
[0050] Combined with Figure 1 This embodiment proposes a collaborative monitoring and human-computer interaction system integrating multimodal perception, which includes:
[0051] Wearable devices interact with the sweeping robot and are equipped with multiple sensors to collect multimodal sensor data covering location, objects, sound, and motion in the environment. They are also used to receive instructions sent by the wearable device and dynamically adjust the working modes of multiple sensors.
[0052] The sweeping robot interacts with the wearable device and has a built-in human behavior network model and alarm module. It is used to analyze the data collected by the wearable device in real time and identify the user's human behavior through the human behavior network model, determine the sensor usage mode corresponding to the identification result based on the preset rule set, and send execution commands to the wearable device;
[0053] The human behavior network model is built into the sweeping robot. It is used to periodically obtain the user's historical daily human behavior data accumulated by the wearable device through the sweeping robot, learn the order constraints, the time pattern of the behavior, and the duration information from the acquired data, and is also used to integrate the multimodal sensor data collected by the wearable device in real time and the information learned from the historical behavior data. By analyzing the integrated data, the user's real-time human behavior can be accurately identified;
[0054] The alarm module is built into the sweeping robot and is used to send an alarm to a preset emergency contact in a timely manner when the human behavior network model identifies the user's abnormal human behavior according to the preset abnormal behavior rules;
[0055] The human-computer interaction module provides a human-computer interaction interface for web pages and mobile APPs, which is used for family members to log in to the human-computer interaction interface and display the user's human behavior information in real time.
[0056] In this embodiment, the wearable device includes a main control board, a power module and a housing, wherein the main control board is built with a Raspberry Pi Zero computer as the core, and integrates multiple sensors and a Raspberry Pi camera to collect multimodal data and provide data support for human behavior recognition; the multiple sensors include an accelerometer, a gyroscope, a microphone and a camera, wherein the accelerometer and the gyroscope are used to detect motion data, the microphone is used to collect sound data, and the camera is used to collect object and scene image data. The wearable device can fully perceive the user's environment and its own state through the above sensors;
[0057] The wearable device communicates with the sweeping robot through the WiFi module. Specifically, the wearable device regularly sends the user's accumulated historical daily human behavior data to the sweeping robot through the WiFi module. At the same time, the wearable device sends the user's collected real-time human behavior data to the sweeping robot through the WiFi module; the sweeping robot decides which sensor mode to use based on the user's real-time human behavior recognized by the human behavior network model, and sends an execution command to the wearable device.
[0058] In this embodiment, the human behavior network model involved includes a position recognition sub-model, an object recognition sub-model, a sound event recognition sub-model, a body movement recognition sub-model and an integrated analysis module, wherein:
[0059] A location recognition sub-model, which is used to receive multi-modal sensor data covering location aspects collected by a wearable device, generate a probability distribution for different locations, and accurately identify the specific location where the human body is located;
[0060] An object recognition sub-model, which is used to receive multi-modal sensor data covering object aspects collected by a wearable device, detect multiple objects in an image, and generate a probability for each object to achieve accurate recognition of specific objects;
[0061] A sound event recognition sub-model, which is used to receive multi-modal sensor data covering sound aspects collected by a wearable device to achieve accurate recognition of various sound events;
[0062] A body movement recognition sub-model, which is used to receive multi-modal sensor data covering movement aspects collected by a wearable device to achieve accurate recognition of body movements;
[0063] An integrated analysis module, which is used to collect the location information output by the location recognition sub-model, the object information recognized by the object recognition sub-model, the sound event information detected by the sound event recognition sub-model, and the body movement information recognized by the body movement recognition sub-model, and use rule-based reasoning or classification algorithms in machine learning to fuse the foregoing multi-source information, mine the internal connections and relationships between the multi-source information, and output an accurate recognition result of the real-time behavior of the human body.
[0064] Specifically, the involved location recognition model is based on the pre-trained Xception model. This basic model has excellent feature extraction capabilities and can achieve a top-five validation accuracy of 0.945 for classifying 1,000 different categories; on the basis of the basic model, a final dense layer with a softmax activation function is added to form a location recognition model;
[0065] The location recognition model is trained using historical multi-modal sensor data in terms of location, so that the location recognition model can generate a probability distribution for different locations and finally achieve accurate recognition of specific locations; the locations involved include the living room, dining room, kitchen, bedroom, bathroom, corridor, and entrance door.
[0066] Specifically, the involved object recognition model is based on the real-time object detection method YOLO based on CNN. This basic architecture has the ability to detect multiple targets simultaneously and can effectively detect objects; on the basis of YOLO, its pre-trained model is used to form an object recognition model;
[0067] The object recognition model is trained using historically accumulated multimodal sensor data on objects, so that the object recognition model can simultaneously detect multiple objects in an image, generate a probability for each object, and achieve accurate recognition of specific objects; the objects include dining tables, sofas, coffee tables, televisions, bookshelves, stoves, desks, computers, things and snacks.
[0068] Specifically, the sound event recognition model involved is based on a CNN-based network architecture, which has the potential to process sound feature data to identify events;
[0069] The Mel-frequency cepstral coefficients are used as sound features, and the sound sampling duration is set to 1 second, the sampling rate is set to 32000 Hz, the fast Fourier transform window size is set to 2048, the step size is set to 1024, and the number of Mel frequency bands is set to 64. Under the above parameter settings, feature data with a shape of 64×32 is generated. Based on the generated feature data, a CNN network consisting of three convolutional layers and two dense layers is constructed, and the output of each convolutional layer is processed by batch normalization and ReLU activation function to form a sound event recognition model, in which the convolutional layer is used to extract local features in the sound feature data. The convolution kernel size of each convolution layer is 3×3, and the dimensions of the three convolution layers are 64, 128, and 256 respectively. Each convolution layer is followed by a maximum pooling layer, which is used to downsample the feature map output by the convolution layer. The size of the pooling kernel of the maximum pooling layer is 2×2. The dimensions of the two dense layers are 256 and 6 respectively. The 256 neurons in the first dense layer are used to further integrate and abstract the features extracted by the previous convolution layer and pooling layer. The 6 neurons in the latter dense layer correspond to the 6 different sound event categories that need to be identified, and the probability corresponding to each category is output through the softmax activation function.
[0070] The sound event recognition model is trained using historically accumulated multimodal sensor data on sound. The model parameters are continuously adjusted to enable the model to better fit the training data. After trying different numbers of layers and layer dimensions, the model structure with the least parameters and the best performance is determined, so that the sound event recognition model can accurately identify various types of sound events by processing the generated feature data through a CNN network.
[0071] Specifically, the body motion recognition model involved is based on a CNN model with two convolutional layers and two dense layers. This infrastructure has the ability to extract features from acceleration data and perform motion recognition, and can effectively learn body motion features. Using 2 seconds of acceleration data as samples, the convolutional layer kernel size is set to 2×2, and the dimensions are set to 16 and 32 respectively, based on which a body motion recognition model is formed. During the training process of the body motion recognition model, the Adam algorithm is used for model optimization, and the dropout method is used to prevent the model from overfitting.
[0072] The body motion recognition model is trained using historically accumulated multimodal sensor data on motion, so that it can process 2 seconds of acceleration data, extract features from samples, learn features through convolutional and dense layers, and ultimately achieve accurate recognition of walking movements.
[0073] Embodiment 2:
[0074] Based on the system of the first embodiment, this embodiment proposes a collaborative monitoring and human-computer interaction method integrating multimodal perception, which includes the following steps:
[0075] S1. Collect multimodal sensor data covering location, objects, sound, and motion through wearable devices.
[0076] Specifically, first, start the wearable device, and the accelerometer and gyroscope continuously collect acceleration and angular velocity data during the user's movement at a set sampling frequency (for example, 100Hz). These data can reflect the user's body movement state, such as whether he is walking, running, turning, etc. For example, when the user walks, the accelerometer will detect periodic changes in acceleration, and the gyroscope can sense the changes in the angle of body rotation; the microphone records the surrounding sound at a sampling rate of 32000Hz and converts the analog sound signal into digital audio data. It can capture various sound events, such as talking, footsteps, and door closing sounds; the camera turns on the image acquisition function and captures images of objects and scenes around the user at a set resolution (such as 1280x720 pixels) and frame rate (such as 30 frames / second). These image data are used for object recognition and scene analysis. Subsequently, the main control board will initially integrate the data collected by the accelerometer, gyroscope, microphone and camera, package them according to the set data format (such as JSON or custom binary format), and send the packaged multimodal sensor data covering location, objects, sound and motion to the sweeping robot through the wearable device's own WiFi module.
[0077] S2, the sweeping robot interacts with the wearable device through the WiFi module. The sweeping robot receives the multimodal sensor data transmitted by the wearable device and transmits it to the built-in human behavior network model for processing;
[0078] The sweeping robot has a built-in human behavior network model. The human behavior network model obtains the user's historical daily human behavior data accumulated by the wearable device through the sweeping robot regularly (for example, at 2 a.m. every day), and learns the order constraints, the time pattern of the behavior, and the duration information from these data. a) Learn the order constraints, such as users usually go to the bathroom first after getting up, and then go to the kitchen to eat breakfast, b) Learn the time pattern of the behavior, such as eating breakfast in the restaurant around 8 a.m., c) Learn the duration information of the behavior, such as cooking in the kitchen generally lasts 30 minutes to 1 hour; it should be added that these historical data are stored in the local storage module of the wearable device or the cloud server, and the sweeping robot obtains these data through the set communication protocol (such as HTTP or MQTT);
[0079] At the same time, the human behavior network model integrates the multimodal sensor data collected by wearable devices in real time based on the information learned from historical behavior data, and analyzes and identifies the user's human behavior. Specifically: ① The location recognition sub-model receives the multimodal sensor data covering the location collected by the wearable device, and generates a probability distribution of different locations such as the living room, dining room, kitchen, bedroom, bathroom, corridor and entrance door, so as to accurately identify the specific location of the human body; ② The object recognition sub-model receives the multimodal sensor data covering objects collected by the wearable device, detects multiple objects in the image, and identifies objects such as dining tables, sofas, coffee tables, TVs, bookshelves, stoves, desks, computers, food and snacks, and generates a probability of belonging to the category for each detected object. For example, when an object is detected to have the shape of a TV and screen features, it is detected that the object has a certain shape and screen features. The model outputs the probability that the object is a TV as 0.95; ③ The multimodal sensor data covering sound collected by the wearable device is received through the sound event recognition sub-model, and various sound events (such as speaking, footsteps, door opening and closing, cooking, TV, and alarm) are accurately recognized. For example, when the sound is detected to have a high rhythm and pitch change, and the energy is concentrated in a specific frequency range, the model outputs the probability that the sound is a speaking sound as 0.9; ④ The multimodal sensor data covering movement collected by the wearable device is received through the body movement recognition sub-model, and 2 seconds of acceleration data is processed with 2 seconds as a data sample, and the 2 seconds of acceleration data is obtained from the sample. ⑤ Through the integration and analysis module, the position information output by the position recognition sub-model, the object information recognized by the target recognition sub-model, the sound event information detected by the sound event recognition sub-model, and the body movement information recognized by the body movement recognition sub-model are collected, and the rule-based reasoning or machine learning is used to identify the user's walking action. The classification algorithm in learning (such as random forest classifier) will fuse the collected multi-source information. For example, rule-based reasoning can be set: if the location is identified as a kitchen, the object is identified as a stove and the sound of cooking is detected, and the body movement is identified as frequent hand movements, then it is judged that the user is cooking; the classification algorithm in machine learning is trained on a large amount of labeled multimodal data to learn the relationship between different information combinations and human behavior, so as to classify the current multi-source information, explore the intrinsic connections and relationships between multi-source information, and output accurate recognition results of real-time human behavior, such as "the user is sleeping in the bedroom", "the user is watching TV in the living room", etc.
[0080] S3. The sweeping robot decides which sensor mode to use based on the recognition result of the human behavior network model and the preset rule set, and sends a message containing a sensor mode adjustment instruction to the wearable device through the WiFi module (for example, sending an instruction to inform the wearable device to reduce the sampling frequency of the accelerometer and gyroscope from 100 Hz to 20 Hz). After the wearable device receives the instruction sent by the sweeping robot, the main control board adjusts the working mode of the corresponding sensor according to the instruction (for example, the main control board sets the registers of the accelerometer and gyroscope to reduce the sampling frequency to the specified value), thereby realizing dynamic adjustment of the working mode of the sensor of the wearable device.
[0081] For example, if it is recognized that the user is stationary and in the bedroom, this may mean that the user is resting. At this time, the sweeping robot determines that the sampling frequency of the motion sensor (accelerometer and gyroscope) can be reduced to save power.
[0082] The rule set can predefine the correspondence between multiple behavioral scenarios and sensor modes. For example, when a user is running outdoors, the sampling frequency of the motion sensor is increased to capture motion data more accurately; when the user is doing quiet activities indoors, the sensitivity of the microphone is appropriately reduced to reduce unnecessary sound collection.
[0083] S4. When the human behavior network model identifies the user's abnormal human behavior according to the preset abnormal behavior rules, the alarm module promptly sends an alarm message to the preset emergency contact. The alarm message includes the abnormal behavior type (such as "suspected fall", "staying still in the bathroom for a long time"), the user's location (such as "bedroom") and possible related information (such as detected sound events, nearby objects, etc.), so that the emergency contact can quickly understand the situation and take corresponding measures.
[0084] For example, the preset abnormal behavior rule is: if the user remains stationary in the bathroom for more than 30 minutes, or a sudden large acceleration change (which may indicate a fall) is detected, it is determined to be an abnormal behavior.
[0085] S5. The human-computer interaction module provides a human-computer interaction interface for the web page and the mobile APP. The web page can be accessed through a browser, and the mobile APP can be downloaded and installed in the mobile application store;
[0086] After a family member opens a webpage or APP, they enter their account and password through the login interface to log in. After successful login, the human-computer interaction interface will display the user's human behavior information in real time. For example, the human-computer interaction interface will display "The user is watching TV in the living room" in text form, and may be accompanied by relevant icons or simple graphical displays, such as marking the user's location in the floor plan of the living room.
[0087] It should be added that the human-computer interaction interface can also provide a historical behavior record query function, so that family members can view the user's behavior history over the past period of time, such as the distribution of the user's activity time in each room every day in the past week. These historical data can be displayed in the form of charts (such as bar charts and line charts), which is convenient for family members to understand the user's behavior patterns and changing trends.
[0088] In summary, a collaborative monitoring and human-computer interaction system and method integrating multimodal perception is adopted in the present invention, combined with a wearable device and a household sweeping robot, human behavior recognition and monitoring are realized through a human behavior network model, and the energy of the wearable device is effectively saved.
[0089] The above specific examples are used to explain the principles and implementation methods of the present invention in detail. These examples are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made by technicians in this technical field without departing from the principles of the present invention should fall within the scope of patent protection of the present invention.
Claims
1. A collaborative monitoring and human-computer interaction system integrating multimodal perception, characterized in that: It includes: Wearable devices, which interact with the robot vacuum cleaner, are equipped with multiple sensors to collect multimodal sensor data covering location, objects, sound, and motion in the environment; It is also used to receive instructions sent by wearable devices and dynamically adjust the working modes of various sensors; The sweeping robot interacts with the wearable device and has a built-in human behavior network model and alarm module. It is used to analyze the real-time data collected by the wearable device and identify the user's human behavior through the human behavior network model, determine the sensor usage mode corresponding to the identification result based on the preset rule set, and send execution commands to the wearable device; The human behavior network model is built into the sweeping robot. It is used to periodically obtain the user's historical daily human behavior data accumulated by the wearable device through the sweeping robot, learn the order constraints, the time pattern of the behavior, and the duration information from the acquired data, and is also used to integrate the multimodal sensor data collected by the wearable device in real time and the information learned from the historical behavior data. By analyzing the integrated data, the user's real-time human behavior can be accurately identified; The alarm module is built into the sweeping robot and is used to send an alarm to a preset emergency contact in a timely manner when the human behavior network model identifies the user's abnormal human behavior according to the preset abnormal behavior rules; The human-computer interaction module provides a human-computer interaction interface for web pages and mobile APPs, which is used for family members to log in to the human-computer interaction interface and display the user's human behavior information in real time.
2. The collaborative monitoring and human-computer interaction system integrating multimodal perception according to claim 1, characterized in that: The wearable device includes a main control board, a power module and a housing, wherein the main control board is built with a Raspberry Pi Zero computer as the core, and integrates multiple sensors and a Raspberry Pi camera to collect multimodal data and provide data support for human behavior recognition; The wearable device communicates with the sweeping robot via a WiFi module.
3. The collaborative monitoring and human-computer interaction system integrating multimodal perception according to claim 2, characterized in that: Various sensors include accelerometers, gyroscopes, microphones and cameras. Among them, accelerometers and gyroscopes are used to detect motion data, microphones are used to collect sound data, and cameras are used to collect object and scene image data. Wearable devices can use the aforementioned sensors to fully perceive the user's environment and their own status.
4. The collaborative monitoring and human-computer interaction system integrating multimodal perception according to claim 2, characterized in that: The human behavior network model includes a position recognition sub-model, an object recognition sub-model, a sound event recognition sub-model, a body movement recognition sub-model and an integrated analysis module, wherein: The location recognition sub-model is used to receive multi-modal sensor data covering location collected by wearable devices, generate probability distributions of different locations, and accurately identify the specific location of the human body; The object recognition sub-model is used to receive multimodal sensor data covering objects collected by the wearable device, detect multiple objects in the image, and generate the probability of each object to achieve accurate recognition of specific objects; The sound event recognition sub-model is used to receive multi-modal sensor data covering sound collected by wearable devices to accurately recognize various sound events; The body motion recognition sub-model is used to receive multi-modal sensor data covering motion collected by wearable devices to achieve accurate recognition of body motions; The integrated analysis module is used to collect the position information output by the position recognition sub-model, the object information recognized by the target recognition sub-model, the sound event information detected by the sound event recognition sub-model, and the body movement information recognized by the body movement recognition sub-model, and use rule-based reasoning or classification algorithms in machine learning to fuse the aforementioned multi-source information, explore the intrinsic connections and relationships between the multi-source information, and output accurate recognition results of real-time human behavior.
5. The collaborative monitoring and human-computer interaction system integrating multimodal perception according to claim 4 is characterized in that: The location recognition model is based on the pre-trained Xception model, which has excellent feature extraction capabilities and can achieve a top five verification accuracy of 0.945 and classify 1000 different categories. On the basis of the basic model, a final dense layer with a softmax activation function is added to form a location recognition model. The location recognition model is trained using historically accumulated multimodal sensor data on locations, so that the location recognition model can generate probability distributions of different locations and ultimately achieve accurate recognition of specific locations.
6. The collaborative monitoring and human-computer interaction system integrating multimodal perception according to claim 4, characterized in that: The object recognition model is based on the CNN-based real-time target detection method YOLO as the basic architecture, which has the ability to detect multiple targets at the same time and can effectively detect objects; based on YOLO, its pre-trained model is used to form an object recognition model; The object recognition model is trained using historically accumulated multimodal sensor data on objects, so that the object recognition model can simultaneously detect multiple objects in the image, generate the probability of each object, and achieve accurate recognition of specific objects.
7. The collaborative monitoring and human-computer interaction system integrating multimodal perception according to claim 4, characterized in that: The sound event recognition model is based on a CNN-based network architecture, which has the potential to process sound feature data to identify events; The Mel-frequency cepstral coefficients are used as sound features, and the sound sampling duration is set to 1 second, the sampling rate is set to 32000 Hz, the fast Fourier transform window size is set to 2048, the step size is set to 1024, and the number of Mel frequency bands is set to 64. Under the above parameter settings, feature data with a shape of 64×32 is generated. Based on the generated feature data, a CNN network consisting of three convolutional layers and two dense layers is constructed, and the output of each convolutional layer is processed by batch normalization and ReLU activation function to form a sound event recognition model, in which the convolutional layer is used to extract local features in the sound feature data. The convolution kernel size of each convolution layer is 3×3, and the dimensions of the three convolution layers are 64, 128, and 256 respectively. Each convolution layer is followed by a maximum pooling layer, which is used to downsample the feature map output by the convolution layer. The size of the pooling kernel of the maximum pooling layer is 2×2. The dimensions of the two dense layers are 256 and 6 respectively. The 256 neurons in the first dense layer are used to further integrate and abstract the features extracted by the previous convolution layer and pooling layer. The 6 neurons in the latter dense layer correspond to the 6 different sound event categories that need to be identified, and the probability corresponding to each category is output through the softmax activation function. The sound event recognition model is trained using historically accumulated multimodal sensor data on sound. The model parameters are continuously adjusted to enable the model to better fit the training data. After trying different numbers of layers and layer dimensions, the model structure with the least parameters and the best performance is determined, so that the sound event recognition model can accurately identify various types of sound events by processing the generated feature data through a CNN network.
8. The collaborative monitoring and human-computer interaction system integrating multimodal perception according to claim 4, characterized in that: The body motion recognition model is based on a CNN model with two convolutional layers and two dense layers. The basic architecture has the ability to extract features from acceleration data and perform motion recognition, and can effectively learn body motion features. Using 2 seconds of acceleration data as a sample, the convolutional layer kernel size is set to 2×2, and the dimensions are set to 16 and 32 respectively, based on which a body motion recognition model is formed. In the training process of the body motion recognition model, the Adam algorithm is used to optimize the model, and the dropout method is used to prevent the model from overfitting. The body motion recognition model is trained using historically accumulated multimodal sensor data on motion, so that it can process 2 seconds of acceleration data, extract features from samples, learn features through convolutional and dense layers, and ultimately achieve accurate recognition of walking movements.
9. A collaborative monitoring and human-computer interaction method integrating multimodal perception, characterized in that: Based on the system as claimed in claim 1, its implementation includes the following steps: S1. Collect multimodal sensor data covering location, objects, sound, and motion through wearable devices; S2, the sweeping robot interacts with the wearable device and receives multimodal sensor data transmitted by the wearable device; The sweeping robot has a built-in human behavior network model. The human behavior network model periodically obtains the user's historical daily human behavior data accumulated by wearable devices through the sweeping robot, and learns the order constraints, time rules of behavior occurrence, and duration information from these data; At the same time, the human behavior network model integrates the multimodal sensor data collected by wearable devices in real time based on the information learned from historical behavior data to analyze and identify the user's human behavior; S3, the sweeping robot determines which sensor mode to use based on the recognition result of the human behavior network model and the preset rule set, and sends an execution command to the wearable device to realize dynamic adjustment of the sensor working mode of the wearable device; S4. When the human behavior network model identifies the abnormal human behavior of the user according to the preset abnormal behavior rules, an alarm module promptly sends an alarm message to the preset emergency contact; S5. The human-computer interaction module provides a human-computer interaction interface for web pages and mobile APPs. Family members log in to the human-computer interaction interface to display the user's human behavior information in real time, thus realizing information interaction between people and the system.
Citation Information
Cited By
Home abnormal state signal detection method and system based on multi-mode sensing
CN120216965A
Abnormal home state signal detection method and system based on multimodal sensing
CN120216965B
Robot body intelligent control method and system based on multi-mode perception
CN120516727A