Man-machine collaboration degree modeling and evaluating system and method based on multi-modal perception
Through multimodal perception and modeling technology, real-time quantitative evaluation and personalized regulation of human-machine collaboration are achieved, and the problems of insufficient human-machine collaboration measurement, poor natural interaction and limited personalized regulation capabilities in the existing technology are solved, which improves the naturalness and efficiency of human-machine collaboration and is suitable for rehabilitation, service and nursing robot scenarios.
Patent Information
- Application Number
- CN202510437564.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-22
AI Technical Summary
The existing technology has shortcomings in human-machine collaborative quantitative evaluation, natural interaction and personalized regulation, which is difficult to fully reflect the quality of collaboration, the robot behavior adjustment mechanism is imperfect, and the interaction strategy cannot be dynamically optimized according to user status, and the application of multimodal perception and modeling technology limits the naturalness and efficiency of human-machine collaboration.
It provides a human-machine collaborative degree modeling and evaluation system based on multimodal perception, including a multimodal perception module, feature fusion and extraction module, a human-machine collaborative degree modeling module, a personalized interactive strategy feedback module and a task recording and evaluation module. Through multi-source data acquisition, feature fusion, collaborative degree modeling and personalized feedback, real-time quantitative evaluation and dynamic regulation are realized.
It significantly improves the naturalness, efficiency and user experience of human-machine collaboration, can more accurately perceive user status, dynamically adjust robot behavior, support long-term trend analysis and user portrait construction, and is suitable for scenarios such as rehabilitation robots, service robots and nursing robots, improving collaboration efficiency and safety.
Smart Images

Figure CN120354353A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robots and intelligent systems, and particularly to a human - machine collaboration degree modeling and evaluation system and method based on multi - modal perception. Background Art
[0002] With the rapid development of artificial intelligence and robotics technologies, robots have been widely used in human - machine collaboration scenarios, such as rehabilitation training, home care, assisted mobility, and industrial production. In rehabilitation training, robots can assist patients in gait recovery or limb function exercises; in home care, service robots can help the elderly complete daily tasks, such as carrying items or providing life assistance; in industrial production, collaborative robots work with workers to complete assembly or handling tasks. These application scenarios have put forward higher and higher requirements for the naturalness, efficiency, and user experience of human - machine collaboration. However, there are still many deficiencies in the prior art in achieving efficient and natural human - machine collaboration, which limits the wide application of robots in complex scenarios.
[0003] Most existing human - machine collaboration systems rely on preset behaviors and fixed rules for control. For example, in the field of rehabilitation robots, robots usually assist patients in training according to pre - programmed motion trajectories and speeds, lacking the ability to perceive the real - time state of patients and make dynamic adjustments; in the field of service robots, robots often execute tasks through simple voice commands or sensor triggers, and are unable to adjust interaction strategies according to the physiological state or emotional changes of users. This control method based on fixed rules is difficult to meet the personalized needs of users, resulting in low human - machine collaboration efficiency, unnatural interaction experience, and low user acceptance. Research shows that about 60% of users feel uncomfortable when using robots with fixed patterns. Especially during long - term collaboration, users may develop fatigue or resistance due to the single behavior pattern of the robot.
[0004] In addition, there are obvious defects in the quantitative evaluation of human - machine collaboration degree in the prior art. Traditional robot systems usually evaluate the collaboration effect only through single indicators such as task completion rate or execution time, lacking a comprehensive quantitative description of the degree of human - machine collaboration. For example, in rehabilitation training, the system may only record the number of steps or training duration of patients, and is unable to perceive the physiological state (such as heart rate, electromyogram signal), behavioral intention (such as movement amplitude, step frequency), and interaction performance (such as synchronization with the robot) of patients in real time. This single evaluation method is difficult to reflect the real quality of human - machine collaboration and cannot provide a scientific basis for robot behavior optimization. At the same time, in multi - task scenarios, existing systems lack a systematic modeling and regulation mechanism for dynamically adjusting robot behavior. For example, when a user is performing complex tasks (such as walking and carrying items simultaneously), the robot is unable to dynamically adjust the action speed, intonation, or prompt frequency according to the user's fatigue level or task load, resulting in incoordination or even safety hazards during the collaboration process.
[0005] In recent years, multi-modal perception technology has made certain progress in the field of human-computer interaction. By integrating multi-source data such as vision, speech, and physiological signals, it can comprehensively perceive the user's state. For example, action recognition technology based on depth cameras can capture the user's limb movements, speech analysis technology based on microphone arrays can extract the user's emotional characteristics, and physiological signal monitoring based on wearable devices can reflect the user's fatigue level. However, there are still deficiencies in the fusion and modeling of multi-modal data in the existing technology. On the one hand, the fusion methods of multi-modal features are relatively simple, usually using direct splicing or weighted averaging, which are difficult to effectively capture the correlation and dynamic changes between different modalities; on the other hand, the existing systems lack a method for modeling the synergy degree based on multi-modal data, and cannot convert the user's state perception into a quantifiable synergy score to guide the dynamic adjustment of robot behavior. In addition, there are also limitations in personalized interaction in the existing technology. It is difficult for robots to perform adaptive optimization according to the user's long-term behavior habits or preferences, resulting in a lack of personalization and continuity in the interaction experience.
[0006] To sum up, the existing technology has the following problems in human-machine cooperation measurement, natural interaction, and personalized regulation: First, there is a lack of a real-time quantitative evaluation method for the degree of human-machine cooperation, which is difficult to comprehensively reflect the cooperation quality; second, the robot behavior adjustment mechanism is imperfect, and it cannot dynamically optimize the interaction strategy according to the user's state; third, the application of multi-modal perception and modeling technology is insufficient, which limits the naturalness and efficiency of human-machine cooperation. Therefore, there is an urgent need for a human-machine cooperation degree evaluation system that integrates multi-modal perception and intelligent modeling technology, which can realize real-time quantitative analysis and personalized regulation of the human-machine cooperation process, provide a scientific basis for robot personalized design and behavior optimization, so as to improve the level of human-machine coexistence and meet the cooperation needs in complex scenarios.
[0007] In view of the above problems, there is an urgent need for a human-machine cooperation degree modeling and evaluation system and method based on multi-modal perception. Summary of the Invention
[0008] The purpose of the present invention is to solve the existing technical problems proposed in the above background technology, and provide a human-machine cooperation degree modeling and evaluation system and method based on multi-modal perception.
[0009] The present invention achieves the above purpose through the following solutions:
[0010] On the one hand, the solution of the present invention provides a human-machine cooperation degree modeling and evaluation system based on multi-modal perception, and the system includes:
[0011] A multi-modal perception module, which is used to collect multi-source information of the user in real time during the task execution process, and the multi-source information includes visual image information, speech intonation information, physiological signal information, and interaction action data;
[0012] A feature fusion and extraction module, connected to the multi-modal perception module, for preprocessing, feature extraction, and modal fusion of the multi-source information to generate a unified human-computer interaction feature vector;
[0013] A human-machine collaboration degree modeling module, connected to the feature fusion and extraction module, for constructing a human-machine collaboration degree evaluation model based on the human-computer interaction feature vector through supervised learning or rule inference methods, and outputting a human-machine collaboration degree score;
[0014] A personalized interaction strategy feedback module, connected to the human-machine collaboration degree modeling module, for dynamically adjusting the robot behavior strategy according to the human-machine collaboration degree score to improve the naturalness and efficiency of human-machine collaboration;
[0015] A task recording and evaluation module, for recording the original data, feature extraction results, collaboration degree score, and robot behavior adjustment records of human-computer interaction tasks, and supporting user portrait construction and long-term trend analysis.
[0016] As a preferred technical solution of the present invention, the multi-modal perception module includes:
[0017] A visual acquisition unit, which acquires the facial expressions, body movements, and gait trajectories of the user through a depth camera and an RGB camera;
[0018] A voice acquisition unit, which acquires the language content, intonation intensity, speech rate, and tone characteristics of the user through a microphone array;
[0019] A physiological signal acquisition unit, which acquires the heart rate, electro-dermal activity (EDA), surface electromyogram (sEMG), and oxyhemoglobin (HbO) data of the user through a wearable device;
[0020] An interaction action acquisition unit, which records the interaction trajectory, step frequency, grip force, and posture changes between the user and the robot through an inertial measurement unit (IMU), a force sensor, and a displacement encoder;
[0021] Among them, the multi-modal perception module performs preliminary synchronization and caching on the acquired data through a local edge computing module, and uses timestamp annotation for subsequent multi-modal fusion.
[0022] As a preferred technical solution of the present invention, the feature fusion and extraction module includes:
[0023] A data preprocessing sub-module, for normalizing and background filtering of visual images, endpoint detection and MFCC parameter extraction of voice signals, and filtering, denoising, and peak recognition of physiological signals;
[0024] A feature extraction sub-module, using the feature extraction function Fm = φ m (X m ; θ m ) extracts features from the original data X of the m-th modality m to generate a feature vector F m , where φ m (·; θ m ) is a feature extraction function and θ m is a parameter;
[0025] The modality fusion sub-module fuses multi-modal features through an attention mechanism to generate a unified human-computer interaction feature vector F. The fusion formula is:
[0026]
[0027] where ɑ m is the attention weight of the m-th modality, w m is an attention weight vector, and F includes emotional state, physiological load, and action synchronization degree.
[0028] As a preferred technical solution of the present invention, the human-machine cooperation degree modeling module includes:
[0029] The model training sub-module is used to input the human-computer interaction feature vector and expert labels and construct a human-machine cooperation degree evaluation model using a supervised learning method;
[0030] The scoring output sub-module is used to output a human-machine cooperation degree score S. The scoring formula is:
[0031]
[0032] Or in the form of a kernel function:
[0033]
[0034] where S ∈ [0, 100], β i is a model parameter, F[i] is a component of the feature vector, F j is a support vector, and σ is a kernel width parameter;
[0035] The scoring configuration sub-module supports users or managers to set scoring sensitivity and personalized weight allocation strategies according to scenario requirements.
[0036] As a preferred technical solution of the present invention, the human-machine cooperation degree modeling module supports continuous online learning or periodic fine-tuning to improve the generalization ability of the model in long-term cooperation.
[0037] As a preferred technical solution of the present invention, the personalized interaction strategy feedback module includes: an intonation and speech rate adjustment unit that adjusts the speech rate v according to the score S. The formula is:
[0038] v = v0 + k v ·(S - S ref );
[0039] The action adjustment unit adjusts the action amplitude a according to the score S, and the formula is:
[0040] a = a0 + k a ·(S - S ref );
[0041] The prompt optimization unit adjusts the prompt frequency f according to the score S, and the formula is:
[0042] f = f0 - k f ·(S - S ref );
[0043] Wherein, v0, a0, f0 are reference values, S ref is the reference score, k v , k a , k f is the adjustment coefficient;
[0044] The preference learning unit is used to record user feedback and adjust default behavior parameters to achieve adaptive optimization.
[0045] As a preferred technical solution of the present invention, the task recording and evaluation module includes:
[0046] The task log storage unit is used to save user number, task type, timestamp and cooperation degree score information;
[0047] The user portrait construction unit is used to statistically analyze the physiological response fluctuations, behavior habits and preference response patterns of users based on historical task data;
[0048] The trend analysis unit calculates the long-term cooperation efficiency trend T t , and the formula is:
[0049] T t = λS t +(1 - λ)T t-1 ;
[0050] Wherein, S t is the cooperation degree score at time t, and λ is the smoothing factor; the data export unit supports exporting data in CSV, JSON or XML format.
[0051] As a preferred technical solution of the present invention, the system is applicable to the human-robot cooperation scenarios of rehabilitation robots, service robots or nursing robots, and supports the personalized design and behavior optimization of robots.
[0052] On the other hand, the solution of the present invention also provides a method for modeling and evaluating human-machine collaboration based on multi-modal perception, and the method includes the following steps:
[0053] Real-time collect the user's visual image information, speech intonation information, physiological signal information and interaction action data through a multi-modal perception module;
[0054] Use the feature fusion and extraction module to preprocess, extract features and perform modal fusion on the collected multi-source information to generate a unified human-machine interaction feature vector F, where the feature extraction formula is F m =φ m (X m ; θ m ), and the fusion formula is Attention weight
[0055] Based on the human-machine interaction feature vector, construct an evaluation model through the human-machine collaboration modeling module, and output the human-machine collaboration score S. The scoring formula is
[0056] According to the human-machine collaboration score, dynamically adjust the robot behavior strategy through the personalized interaction strategy feedback module;
[0057] Record the original data, feature extraction results, collaboration score and behavior adjustment records of the human-machine interaction task through the task record and evaluation module, and perform user portrait construction and long-term trend analysis.
[0058] As a preferred technical solution of the present invention, the method further includes, during the feature extraction process, using a convolutional neural network, a long short-term memory network and a fast Fourier transform to process images, dynamic pose sequences and physiological signals respectively;
[0059] During the modal fusion process, perform weighted fusion on multi-modal features through an attention mechanism; during the behavior strategy adjustment process, dynamically optimize the robot's speech rate v, action amplitude a and prompt frequency f according to the score S. The adjustment formula is: v = v0 + k v ·(S - S ref ), a = a0 + k a ·(S - S ref ), f = f0 - k f ·(S - S ref );
[0060] During the trend analysis process, use the exponentially weighted moving average method to calculate the long-term trend T t , and the formula is T t =λS t +(1 - λ)T t-1 .
[0061] Compared with the prior art, the present invention has the following beneficial effects:
[0062] 1. A human - machine collaboration degree modeling and evaluation system and method based on multi - modal perception provided by the present invention significantly improves the naturalness, efficiency, and user experience of human - machine collaboration through the collaborative work of modules such as multi - modal perception, feature fusion, collaboration degree modeling, personalized feedback, and task recording. First, the system comprehensively collects data on the user's vision, voice, physiological signals, and interaction actions through the multi - modal perception module, overcoming the defect of single - dimensional perception in traditional technologies and being able to more accurately perceive the user's state, providing a rich data basis for subsequent modeling. For example, in rehabilitation training, the system can simultaneously perceive the patient's movement trajectory and physiological state, avoiding training risks caused by misjudgment of the state.
[0063] 2. In the solution of the present invention, the effective integration of multi - modal data is achieved through the feature fusion and extraction module. The attention mechanism is used to dynamically adjust the contribution degree of different modalities, generating a more representative unified feature vector, significantly improving the accuracy and robustness of feature expression. The human - machine collaboration degree modeling module constructs an evaluation model based on this feature vector and outputs a quantitative score, filling the gap in the prior art lacking a collaboration measurement method. At the same time, the personalized interaction strategy feedback module dynamically adjusts the robot's behavior according to the score, such as speech rate, action amplitude, and prompt frequency, making the interaction more natural and the user acceptance higher. Especially in long - term collaboration, the user's trust and comfort in the robot are significantly enhanced.
[0064] 3. In the solution of the present invention, the systematic management and analysis of collaboration data are realized through the task recording and evaluation module, supporting long - term trend analysis and user portrait construction, providing a scientific basis for collaboration optimization. The system is applicable to various scenarios such as rehabilitation robots, service robots, and nursing robots, and can dynamically adapt to the user's state, improving collaboration efficiency and safety. For example, in the service robot scenario, the system can adjust its behavior according to the user's fatigue state to improve the user experience; in the nursing scenario, the system optimizes the assisting behavior to reduce the operation difficulty. Generally speaking, the present invention effectively solves the problems of insufficient collaboration measurement, poor natural interaction, and limited personalized regulation ability in the prior art, providing important support for human - machine co - integration. Description of the Drawings
[0065] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0066] Figure 1It is the system block diagram of a human - machine collaboration degree modeling and evaluation system based on multi - modal perception of the present invention;
[0067] Figure 2 It is the multi - modal perception data acquisition flow chart of a human - machine collaboration degree modeling and evaluation system based on multi - modal perception of the present invention;
[0068] Figure 3 It is the human - machine collaboration degree modeling and evaluation flow chart of a human - machine collaboration degree modeling and evaluation system based on multi - modal perception of the present invention;
[0069] Figure 4 It is the schematic diagram of the personalized feedback control mechanism of a human - machine collaboration degree modeling and evaluation system based on multi - modal perception of the present invention. Specific embodiments
[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0071] The solution of this application provides a human - machine collaboration degree modeling and evaluation system and method based on multi - modal perception, aiming to achieve real - time quantitative evaluation and dynamic regulation in the process of human - machine collaboration. The system includes a multi - modal perception module, a feature fusion and extraction module, a human - machine collaboration degree modeling module, a personalized interaction strategy feedback module, and a task record and evaluation module. The implementation methods of each module and the application of relevant algorithm formulas are described in detail below.
[0072] First, the system uses the multi - modal perception module to collect multi - source information of the user in real - time during the task execution process, constructing a comprehensive data basis for the user's current state. The acquisition process involves multiple modal data: Visual data is collected by a depth camera (such as Intel RealSense D435) and an RGB camera to obtain the user's facial expressions, body movements, and gait trajectories, with a collection frequency of 30 frames per second and a resolution of 1280×720; Voice data is collected by a microphone array (4 - channel array microphone) to obtain the user's language content, intonation intensity, speech rate, and tone characteristics, with a sampling rate of 16 kHz;
[0073] Physiological signals are collected by wearable devices (such as BioHarness 3.0) for the user's heart rate (range 0 - 200 bpm), electrodermal activity (EDA, range 0 - 100 μS), surface electromyogram (sEMG, range 0 - 5000 μV), and oxyhemoglobin (Hbo, based on near-infrared spectroscopy); interaction action data is recorded by an inertial measurement unit (IMU, such as MPU-6050), force sensors, and displacement encoders for the interaction trajectory, step frequency, grip force (range 0 - 100 N), and attitude changes (angle range -180° to 180°) between the user and the robot. All the collected data is preliminarily synchronized and cached by the local edge computing module, marked with timestamps to ensure the time alignment of multimodal data. The data is stored in JSON format, including the modality type, timestamp, and raw data values.
[0074] Next, the feature fusion and extraction module receives the raw data stream from the multimodal perception module, performs preprocessing, feature extraction, and modality fusion to generate a unified human-computer interaction feature vector. In the preprocessing stage, the visual images are normalized (pixel values are normalized to [0, 1]) and background is filtered (using the background segmentation algorithm in OpenCV), the speech signals are subjected to endpoint detection (based on energy threshold) and MFCC parameter extraction (extracting 13-dimensional MFCC features), and the physiological signals are filtered (Butterworth filter, cut-off frequency 0.5 - 50 Hz), denoised (wavelet denoising), and peak identification (detecting heart rate peaks).
[0075] In the feature extraction stage, different methods are used for different modalities: visual features are extracted by a convolutional neural network (CNN, such as ResNet-18), outputting a 128-dimensional feature vector F1, and the formula is F1 = φ1(X1; θ1), where X1 is the image data and θ1 is the CNN parameter;
[0076] Speech features are processed by a long short-term memory network (LSTM, hidden layer dimension 64), outputting a 64-dimensional feature vector F2, and the formula is F2 = φ2(X2; θ2);
[0077] Physiological signal features are extracted by fast Fourier transform (FFT) to obtain spectral features, outputting a 32-dimensional feature vector F3, and the formula is F3 = φ3(X3; θ3);
[0078] Interaction action features are obtained through statistical analysis (mean, variance) and time series modeling, outputting a 32-dimensional vector F4.
[0079] In the modality fusion stage, the attention mechanism is used to generate a unified feature vector F, and the fusion formula is where the attention weight M = 4, w m is the attention weight vector, The fused feature vector F includes dimensions such as emotional state, physiological load, and action synchronization degree.
[0080] Based on the fused feature vector F, the human-machine collaboration degree modeling module constructs an evaluation model and outputs the human-machine collaboration degree score S. In the model training stage, the fused feature vector F and expert labels (collaboration smoothness, range 0 - 100) are input, and a prediction model is constructed using the support vector machine (SVM). The training dataset contains 10,000 groups of samples, the feature dimension is 128 dimensions, and the scoring formula is where d = 128, β i Obtained through SVM training, S ∈ [0, 100]. To introduce non-linear relationships, the kernel function form can be adopted: where σ = 1.0, N is the number of support vectors (about 200), and the score output is a continuous value. For example, S = 80 indicates a high collaboration degree, and S = 30 indicates a low collaboration degree.
[0081] Users can configure the scoring sensitivity and weight allocation through the interface. For example, set the weight of the emotional state to 0.4, the weight of the physiological load to 0.3, and the weight of the action synchronization degree to 0.3. The model supports continuous online learning and is fine-tuned weekly using new data (about 100 groups of samples) to improve the generalization ability in long-term collaboration.
[0082] The personalized interaction strategy feedback module dynamically adjusts the robot's behavior strategy according to the collaboration degree score S. The speech rate adjustment formula is v = v0 + k v ·(S - S ref ), where v0 = 1.0 (benchmark speech rate, unit: multiple of speed), k v = 0.02, S ref = 50. For example, if S = 80, then v = 1.6, and the speech rate speeds up; the movement amplitude adjustment formula is a = a0 + k a ·(S - S ref ), where a0 = 0.5 (benchmark amplitude, unit: meter), k a = 0.01. For example, if S = 30, then a = 0.3, and the movement amplitude decreases; the prompt frequency adjustment formula is f = f0 - k f ·(S - S ref ), where f0 = 0.1 (benchmark frequency, unit: times per second), k f = 0.002. For example, if S = 20, then f = 0.16, and the prompt frequency increases. The module also records user feedback (such as preferring a slower speech rate) and adjusts the default parameters v0, a0, f0 to achieve adaptive optimization.
[0083] The task recording and evaluation module records the data of the human - machine interaction task and analyzes it. The task log is saved in the form of an SQLite database, including the user ID, task type (such as "walking assistance"), timestamp, and the cooperation degree score S. The user profile is based on the statistical analysis of historical data on heart rate fluctuations, behavior habits, and preference response patterns. The long - term trend analysis uses the Exponentially Weighted Moving Average (EWMA) method, and the formula is T t = λS t +(1 - λ)T t-1 , where λ = 0.1, T0 = S0. For example, if S1 = 70 and S2 = 75, then T1 = 70 and T2 = 70.5. The data support can be exported in CSV format, including the timestamp, score, and behavior adjustment records for doctors or engineers to analyze.
[0084] This system is applicable to the human - machine collaboration scenarios of rehabilitation robots, service robots, and nursing robots, and supports the personalized design and behavior optimization of robots. This will be further illustrated through specific embodiments below.
[0085] Embodiment 1: Rehabilitation robot scenario;
[0086] A stroke patient uses a rehabilitation robot for lower - limb gait training. The system helps evaluate the human - machine cooperation degree and optimize the robot's assisting behavior. When the patient walks during the training, the depth camera captures a step length of 0.4 meters, a step frequency of 0.5 steps per second. The wearable device measures a heart rate of 90 bpm, an average EMG signal of 200 μV, and the force sensor records the supporting force between the patient and the robot as 50 N. The system processes these data through the feature extraction module, extracts gait features (32 - dimensional), physiological features (32 - dimensional), and interaction features (32 - dimensional), and uses the attention mechanism to fuse them, calculating the attention weights as 0.4, 0.3, and 0.3 respectively, to obtain a 128 - dimensional unified feature vector F.
[0087] Based on the feature vector F, the human - machine cooperation degree modeling module calculates the score through the SVM model and obtains S = 60, indicating a medium cooperation degree. The personalized interaction strategy feedback module adjusts the robot's behavior according to the score: the speech rate is adjusted to v = 1.0+0.02·(60 - 50)=1.2, and the robot provides voice guidance at 1.2 times the speech rate; the action amplitude is adjusted to a = 0.5+0.01·(60 - 50)=0.6, and the assisting amplitude increases to 0.6 meters; the prompting frequency is adjusted to f = 0.1 - 0.002·(60 - 50)=0.08, and the prompting frequency is slightly reduced. The task recording module saves these data, and the trend analysis shows that the patient's cooperation efficiency trend T t rises from 60 to 65, indicating that the training effect is gradually improving.
[0088] Embodiment 2: Service robot scenario;
[0089] The service robot assists the user in carrying items in a home environment, and the system evaluates the human-robot collaboration degree and optimizes the interaction experience. During the user's carrying process, the RGB camera detects that the user's facial expression is "fatigued", the microphone records that the user's intonation intensity is low, the speech rate is 2 words per second, the wearable device measures the heart rate to be 100 bpm, and the EDA is 50 μS, indicating that the user has a high load. The system extracts the facial expression features (64 dimensions), speech features (64 dimensions), and physiological features (32 dimensions), fuses them through an attention mechanism, and the attention weights are 0.3, 0.4, and 0.3 respectively to generate a unified feature vector F. The collaboration degree score is calculated to be S = 40, indicating a low collaboration degree. The robot adjusts its behavior according to the score: the speech rate is adjusted to v = 1.0 + 0.02·(40 - 50) = 0.8, slowing down the speech rate to 0.8 times; the movement amplitude is adjusted to a = 0.5 + 0.01·(40 - 50) = 0.4, reducing the carrying amplitude; the prompting frequency is adjusted to f = 0.1 - 0.002·(40 - 50) = 0.12, increasing the prompting frequency to provide more guidance. The task recording module records the user's preference (likes slow interaction), and the trend analysis shows the collaboration efficiency trend T t Stabilizes at around 45, and it is recommended to reduce the task difficulty to improve the user experience.
[0090] Example 3: Nursing robot scenario;
[0091] The nursing robot assists the elderly user to get up, and the system evaluates the collaboration degree and optimizes the robot's behavior. When the user gets up, the depth camera records the posture change as 30° / second, the microphone captures the user's voice command "slow down", the intonation is stable, the wearable device measures the heart rate to be 80 bpm, and the average myoelectric signal is 150 μV. The system extracts the posture features (32 dimensions), speech features (64 dimensions), and physiological features (32 dimensions), fuses them through an attention mechanism, and the attention weights are 0.5, 0.3, and 0.2 respectively to generate a unified feature vector F. The collaboration degree score is calculated to be S = 75, indicating a high collaboration degree. The robot adjusts its behavior according to the score: the speech rate is adjusted to v = 1.0 + 0.02·(75 - 50) = 1.5, speeding up the speech rate to 1.5 times; the movement amplitude is adjusted to a = 0.5 + 0.01·(75 - 50) = 0.75, increasing the assisting amplitude; the prompting frequency is adjusted to f = 0.1 - 0.002·(75 - 50) = 0.05, reducing the prompting frequency to avoid interference. The task recording module saves the data, and the trend analysis shows the collaboration efficiency trend T t Rises from 70 to 78, indicating that the user's adaptability to the robot is gradually improving.
[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A human-machine collaboration degree modeling and evaluation system based on multimodal perception, characterized in that The system includes: A multi-modal perception module, which is used to collect multi-source information of the user in the process of performing tasks in real time. The multi-source information includes visual image information, speech intonation information, physiological signal information, and interaction action data; A feature fusion and extraction module, connected to the multi-modal perception module, which is used to preprocess, extract features, and perform modal fusion on the multi-source information to generate a unified human-computer interaction feature vector; A human-machine cooperation degree modeling module, connected to the feature fusion and extraction module, which is used to construct a human-machine cooperation degree evaluation model based on the human-computer interaction feature vector through supervised learning or rule inference methods, and output a human-machine cooperation degree score; A personalized interaction strategy feedback module, connected to the human-machine cooperation degree modeling module, which is used to dynamically adjust the robot behavior strategy according to the human-machine cooperation degree score to improve the naturalness and efficiency of human-machine collaboration; A task recording and evaluation module, which is used to record the original data of human-computer interaction tasks, feature extraction results, cooperation degree scores, and robot behavior adjustment records, and support the construction of user portraits and long-term trend analysis.
2. The human-machine collaboration degree modeling and evaluation system based on multi-modal perception according to claim 1, characterized in that The multi-modal perception module includes: A visual acquisition unit, which acquires the user's facial expressions, body movements, and gait trajectories through a depth camera and an RGB camera; A voice acquisition unit, which acquires the user's language content, intonation intensity, speech rate, and tone characteristics through a microphone array; A physiological signal acquisition unit, which acquires the user's heart rate, electro-dermal activity (EDA), surface electromyogram (sEMG), and oxyhemoglobin (HbO) data through wearable devices; An interaction action acquisition unit, which records the interaction trajectories, step frequencies, grip forces, and posture changes between the user and the robot through an inertial measurement unit (IMU), a force sensor, and a displacement encoder; Among them, the multi-modal perception module performs preliminary synchronization and caching on the acquired data through a local edge computing module, and uses timestamp annotation for subsequent multi-modal fusion.
3. The human-machine collaboration degree modeling and evaluation system based on multi-modal perception according to claim 1, wherein The feature fusion and extraction module includes: A data preprocessing sub-module, which is used to normalize and filter the background of visual images, perform endpoint detection and MFCC parameter extraction on voice signals, and filter, denoise, and identify peaks on physiological signals; Feature extraction sub-module, using feature extraction function F m = φ m (X m ; θ m ) to extract features from the original data X m of the m-th modality, generating a feature vector F m , where φ m (·; θ m ) is the feature extraction function and θ m is the parameter; A modal fusion sub-module, which fuses multi-modal features through an attention mechanism to generate a unified human-computer interaction feature vector F. The fusion formula is: Among them, α m is the attention weight of the m-th modality, w m is the attention weight vector, and F includes emotional state, physiological load, and action synchronization degree.
4. The human-machine collaboration degree modeling and evaluation system based on multimodal perception according to claim 1, characterized in that, The human-machine cooperation degree modeling module includes: A model training sub-module, which is used to input the human-computer interaction feature vector and expert labels, and construct a human-machine cooperation degree evaluation model by using supervised learning methods; A score output sub-module, which is used to output a human-machine cooperation degree score S. The score formula is: Or in the form of a kernel function: where S ∈ [0, 100], β i is a model parameter, F[i] is a feature vector component, F j is a support vector, and σ is a kernel width parameter; A score configuration sub-module, which supports users or managers to set score sensitivity and personalized weight distribution strategies according to scenario requirements.
5. The human-machine collaboration degree modeling and evaluation system based on multi-modal perception according to claim 4, characterized in that The human-machine cooperation degree modeling module supports continuous online learning or periodic fine-tuning to improve the generalization ability of the model in long-term collaboration.
6. The human-machine collaboration degree modeling and evaluation system based on multi-modal perception according to claim 1, characterized in that, The personalized interaction strategy feedback module includes: an intonation and speech rate adjustment unit, which adjusts the speech rate v according to the score S. The formula is: v = v0 + k v ·(S - S ref ); An action adjustment unit, which adjusts the action amplitude a according to the score S. The formula is: a = a0 + k a ·(S - S ref ); The hint optimization unit adjusts the hint frequency f according to the score S, and the formula is: f = f0 - kx·(S - S ref ); Among them, v0, a0, and f0 are reference values, and S ref is the reference score, and k v , k a , k f are adjustment coefficients; The preference learning unit is used to record user feedback and adjust default behavior parameters to achieve adaptive optimization.
7. The human-machine collaboration degree modeling and evaluation system based on multimodal perception according to claim 1, wherein The task recording and evaluation module includes: The task log storage unit is used to save user number, task type, timestamp and cooperation degree score information; The user portrait construction unit is used to statistically analyze the user's physiological response fluctuations, behavior habits and preference response patterns based on historical task data; Trend analysis unit, which calculates the long-term collaboration efficiency trend T using the exponentially weighted moving average method t , and the formula is: T t = λS t + (1 - λ)T t-1 ; Among them, S t is the synergy degree score at time t, and λ is the smoothing factor; the data export unit supports exporting data in CSV, JSON or XML formats.
8. The human-machine collaboration degree modeling and evaluation system based on multi-modal perception according to claim 1, characterized in that The system is applicable to the human-robot collaboration scenarios of rehabilitation robots, service robots or nursing robots, and supports robot personalized design and behavior optimization.
9. A method for modeling and evaluating human-machine collaboration degree based on multi-modal perception, characterized in that, The method includes the following steps: Real-time collect the user's visual image information, speech intonation information, physiological signal information and interaction action data through the multimodal perception module; The multi-source information collected is preprocessed, feature extracted, and modality fused by using the feature fusion and extraction module to generate a unified human-computer interaction feature vector F, where the feature extraction formula is F m = φ m (X m ; θ m ), and the fusion formula is Attention weight Based on the human-machine interaction feature vector, an evaluation model is constructed by the human-machine collaboration degree modeling module, and the human-machine collaboration degree score S is output. The scoring formula is According to the human-robot cooperation degree score, dynamically adjust the robot behavior strategy through the personalized interaction strategy feedback module; Record the original data, feature extraction results, cooperation degree score and behavior adjustment records of the human-robot interaction task through the task recording and evaluation module, and conduct user portrait construction and long-term trend analysis.
10. The method for modeling and evaluating human-machine collaboration degree based on multi-modal perception according to claim 9, wherein The method also includes that in the feature extraction process, a convolutional neural network, a long short-term memory network and a fast Fourier transform are respectively used to process images, dynamic pose sequences and physiological signals; During the modal fusion process, multi-modal features are weighted and fused through an attention mechanism; during the behavior strategy adjustment process, the speech rate v, action amplitude a, and prompting frequency f of the robot are dynamically optimized according to the score S, and the adjustment formulas are: v = v0 + k v ·(S - S ref ), a = a0 + k a ·(S - S ref ), f = f0 - k f ·(S - S ref ); During the trend analysis process, the exponentially weighted moving average method is used to calculate the long-term trend T t , and the formula is T t = λS t + (1 - λ)T t-1 .
Citation Information
Cited By
Robot control method and system based on multi-modal fusion
CN121061897A
Construction quality data analysis method and system using neural network
CN121257960A
Old-age care robot interaction interface design method based on multiple modes
CN121349585A