Robot behavior mode dynamic adjustment method based on multi-mode perception
Through multimodal perception technology, the integration of visual, voice and environmental data, dynamically adjusts the robot's behavior pattern, solving the accuracy and stability of a single modal interaction method in complex environments, and achieving more efficient and natural human-computer interaction.
Patent Information
- Application Number
- CN202510331368.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing robot interaction methods mainly rely on single modal information, resulting in reduced accuracy and stability in complex environments, making it difficult to adapt to different scenarios and user attributes.
Multimodal perception technology is adopted to collect data in real time through visual, voice and environmental sensors, extract various modal features, and use a weighted fusion algorithm to generate fusion feature vectors, identify scene types and user attributes, dynamically adjust the interaction mode, and update the weight allocation rules in real time.
It improves the robot's interaction accuracy and intelligence level in different scenarios and users, and enhances the adaptability and naturalness of interactive experience.
Smart Images

Figure CN120257050A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human - computer interaction, and particularly to a method for dynamically adjusting the behavior pattern of a robot based on multi - modal perception. Background Art
[0002] With the rapid development of artificial intelligence and human - computer interaction technologies, intelligent robots are increasingly widely used in various application scenarios such as homes, public places, and emergency rescue. Existing robot interaction methods mainly rely on single - modal information, for example, responding only based on voice commands, visual recognition, or environmental perception. However, single - modal information is limited by specific environmental conditions. For example, the accuracy of speech recognition decreases in high - noise environments, the stability of visual recognition decreases under large lighting changes, and it is difficult to effectively infer user intentions in unstructured scenarios through environmental parameter analysis.
[0003] To improve the adaptability of robots to complex scenarios, multi - modal perception technology has gradually become an important research direction for intelligent interaction systems. By fusing visual, voice, and environmental data, robots can more comprehensively perceive user states and interaction needs. However, how to efficiently fuse different modal information and achieve dynamic adjustment of the robot's behavior pattern to adapt to different scenarios and user attributes remains a key issue in the current technological development. Summary of the Invention
[0004] Based on the above - mentioned purpose, the present invention provides a method for dynamically adjusting the behavior pattern of a robot based on multi - modal perception.
[0005] The method for dynamically adjusting the behavior pattern of a robot based on multi - modal perception includes the following steps:
[0006] S1: Deploy visual sensors, voice sensors, and environmental sensors to collect visual images, voice signals, and environmental parameters of the target scene in real - time;
[0007] S2: Extract features of each modality according to the data collected in S1;
[0008] S21: Perform face key - point detection on the visual image and extract user facial expression features;
[0009] S22: Perform voiceprint analysis on the voice signal and extract user age features and emotional features;
[0010] S23: Filter noise from the environmental parameters and extract illumination intensity and temperature features;
[0011] S3: Based on the features of each modality extracted in S2, use a weighted fusion algorithm to dynamically assign weights and generate a fused feature vector;
[0012] S4: Identify the current scene type and user attributes based on the fused feature vector generated in S3, where the user attributes include age, emotional state, and interaction intention;
[0013] S5: Based on the scene type and user attributes identified in S4, match the corresponding interaction mode from a preset policy library, and adjust the tone, speech rate, and robot movement amplitude of the voice output;
[0014] S6: Execute the interaction mode matched in S5, and collect the user voice response duration and body movement following degree in real time as feedback data, and then update the weight allocation rule of the fused feature vector in S3 according to the feedback data.
[0015] Optionally, the specific steps of S1 include:
[0016] S11: Deploy visual sensors on the robot body, where the visual sensors include an RGB camera and an infrared depth camera. The RGB camera is used to collect the color visual image of the target scene, and the infrared depth camera is used to obtain the depth information of the scene and the spatial position data of the target object. The acquisition frame rate is set to 30 frames per second, and the acquisition resolution is set to 1920×1080 pixels;
[0017] S12: Deploy a voice sensor on the robot's head or the sound source direction to collect the voice signal of the target user. The sampling rate of the voice signal is set to 16kHz, and the sampling precision is set to 16bit;
[0018] S13: Deploy environmental sensors on the robot body or inside a fixed environment. The environmental sensors include a light sensor and a temperature sensor. Among them, the light sensor is used to generate light data with a numerical range of 0 - 65535lux; the temperature sensor generates temperature data with an accuracy of ±0.1°C;
[0019] S14: Store the visual image, voice signal, and environmental parameter data collected in S11 - S13 in a time - synchronized manner, and mark the data frames with time stamps.
[0020] Optionally, the specific steps of S21 include:
[0021] S211: Divide the visual image into multiple candidate regions, perform feature convolution calculation on each candidate region, and perform binary classification screening of face and non - face at the network output layer to obtain the face detection result;
[0022] S212: Perform face alignment operation on the detected face region, and correct the tilt and orientation of the face using the normalized affine transformation method, where the affine transformation matrix is mapped according to the face image;
[0023] S213: Locate N key points on the standardized face image and obtain the pixel coordinates of each key point;
[0024] S214: Calculate the geometric distance D between adjacent or designated area key points based on the key point coordinates m,n ;
[0025] S215: D obtained in S24 m,n The facial expression feature vector is formed by arranging them in order and normalized so that all feature data are mapped to the [0,1] interval.
[0026] Optionally, the S22 specifically includes:
[0027] S221: performing noise reduction processing on the speech signal, and performing frame-by-frame windowing to extract short-term speech segments;
[0028] S222: Calculate the fundamental frequency and the first three formant frequencies of the speech signal, wherein the fundamental frequency range is used to distinguish age characteristics;
[0029] S223: extracting the short-time energy, zero-crossing rate and Mel-frequency cepstrum coefficients of the speech signal, and combining the fundamental frequency fluctuation amplitude and speech rate change rate to quantify the emotional characteristics;
[0030] S224: combining the fundamental frequency range, the resonance peak frequency and the emotion quantization result into an age feature vector and an emotion feature vector.
[0031] Optionally, the S23 specifically includes:
[0032] S231: performing time series smoothing processing on the original light intensity signal and temperature signal collected by the environmental sensor;
[0033] S232: Remove background noise from the smoothed light intensity signal and calculate the mean value L of the light intensity change in the continuous time window m and standard deviation L d , and set the threshold θ L For L m +2L d , for values greater than θ L The outliers are removed, and finally the denoised light intensity L is obtained f ;
[0034] S233: Smoothed temperature signal T s Perform anomaly detection and calculate the temperature change rate ΔT within the sliding window. If ΔT exceeds the set threshold θ T , then the temperature data at the time point T s (t) is judged as abnormal, and the window mean is finally used to replace the abnormal value, and finally the denoised temperature data T is obtained. f .
[0035] Optionally, S3 specifically includes:
[0036] S31: Calculate the signal quality scores for the visual features, speech features, and environmental features extracted in S2 respectively. Let the signal quality score of the visual features be Q v , the signal quality score of the speech features be Q a , and the signal quality score of the environmental features be Q e ;
[0037] S32: According to the signal quality scores of each modality, allocate weights according to the following rules:
[0038] Visual feature weight
[0039] Speech feature weight
[0040] Environmental feature weight Among them, W v , W a , W e respectively represent the weighting ratios of the visual, speech, and environmental modalities, such that W v +W a +W e = 100%;
[0041] S33: Standardize each modality feature vector so that its numerical range is unified to 0, 1, and perform linear superposition according to the allocated weights to generate a fused feature vector F;
[0042] S34: Calculate the information loss rate of the fused feature vector. If the information loss rate ≤ 5%, it is determined as effective fusion; otherwise, re-allocate the weights.
[0043] Optionally, S4 specifically includes:
[0044] S41: Based on the average illumination intensity, average temperature, and environmental noise level in the fused feature vector, determine the scene type according to the following rules;
[0045] If the average illumination intensity > 500 lux, the average temperature ≤ 25 °C, and the noise level < 50 dB, it is determined as a home scene;
[0046] If the average illumination intensity ≤ 500 lux, the average temperature > 25 °C, and the noise level ≥ 50 dB, it is determined as a public scene;
[0047] If the average temperature > 35 °C or the illumination intensity variance > 100 lux 2 , and the keywords "help" or "alarm" are detected in the speech features, it is determined as an emergency scene;
[0048] S42: Determine the user's age based on the voice fundamental frequency range and formant frequency in the fused feature vector;
[0049] If the fundamental frequency > 200 Hz and the formant frequency F1 > 500 Hz, determine that the user is a child;
[0050] If 150 Hz ≤ fundamental frequency ≤ 200 Hz and F1 ≤ 500 Hz, determine that the user is an adult;
[0051] If the fundamental frequency < 150 Hz and the formant frequency F1 < 400 Hz, determine that the user is an elderly person;
[0052] S43: Determine the emotional state based on the facial expression features and voice emotion features in the fused feature vector;
[0053] If the mouth corner arc of the facial expression > 0.3, the eyelid opening and closing degree > 0.8, and the short-time voice energy > 60 dB, determine that the emotion is an excited state;
[0054] If the eyebrow tilt angle of the facial expression < -0.2, the eyelid opening and closing degree < 0.5, and the voice zero-crossing rate < 20 Hz, determine that the emotion is a frustrated state;
[0055] If the facial expression features and voice emotion features do not meet the above conditions, determine that the emotion is a calm state;
[0056] S44: Based on the semantic keyword matching and context association in the fused feature vector, further determine the interaction intention, extract the keywords in the voice signal. If the keyword contains temperature or weather, the interaction intention is environmental query; if the keyword contains play or stop, the interaction intention is media control; if the proportion of negative words detected in the continuous dialogue > 30%, the interaction intention is termination request;
[0057] S45: Combine the scene type, user age, emotional state, and interaction intention into a structured label and input it into the adaptive behavior strategy generation module.
[0058] Optionally, the specific steps of S5 are as follows:
[0059] S51: Load the basic interaction template from the policy library according to the recognized scene type;
[0060] S52: Adjust the basic template parameters according to the user's age. If the user is a child, increase the voice fundamental frequency by 10% - 15% and reduce the speech rate to 80% of the basic value; if the user is an elderly person, reduce the voice fundamental frequency by 5% - 10% and reduce the speech rate to 70% of the basic value; if the user is an adult, keep the basic template parameters unchanged;
[0061] S53: Modify the interaction mode according to the user's emotional state. If the emotion is in an excited state, increase the speech rate by 10%, increase the amplitude of actions by 20%, and insert random nodding actions; if the emotion is in a depressed state, decrease the speech rate by 15%, decrease the amplitude of actions by 30%, and adopt a uniform and slow movement; if the emotion is in a calm state, maintain the current parameters.
[0062] S54: Invoke the preset instruction set according to the interaction intention. If the intention is environment query, load the weather data interface from the policy library and output the result in the form of a question and answer; if the intention is media control, call the media player API and synchronously adjust the robot's gestures to match the play / pause actions; if the intention is a termination request, immediately stop the current task.
[0063] S55: Combine the adjusted speech rate, tone, amplitude of actions, and function instructions into a complete interaction strategy and transmit it to the execution module.
[0064] Optionally, the basic interaction template includes a home mode template, a public mode template, and an emergency mode template; where:
[0065] The home mode template corresponds to the home scenario, sets the speech rate to 120 - 150 words per minute, the amplitude of actions to medium, and the default tone to the cordial mode.
[0066] The public mode template corresponds to the public scenario, sets the speech rate to 150 - 180 words per minute, the amplitude of actions to a small amplitude, and the default tone to the formal mode.
[0067] The emergency mode template corresponds to the emergency scenario, sets the speech rate to 180 - 200 words per minute, the amplitude of actions to a large amplitude, and the default tone to the urgent mode.
[0068] Optionally, the specific content of S6 includes:
[0069] S61: Calculate the speech response duration score S t and the limb movement following degree score S m ;
[0070] S62: Calculate the comprehensive feedback score S according to the predetermined weight ratio 综合 , and the formula is: S 综合 = 0.6×S t + 0.4×S m ;
[0071] S63: Adjust the weight allocation rules of each modality in S3 according to S 综合 ;
[0072] When S 综合 ≥ 80%, then keep the current weight allocation rules unchanged;
[0073] When 60% ≤ S 综合 <80%, the voice feature weight ΔW is increased a = 80% - S 综合 × 0.5%; and the environmental feature weight ΔW is decreased e = 80% - S 综合 × 0.3%;
[0074] When S 综合 <60%, the weights are reset to the initially assigned ratios;
[0075] S64: Update the adjusted weight assignment rule to the weighted fusion algorithm in S3 for feature fusion in the next interaction cycle.
[0076] Advantages of the present invention:
[0077] In the present invention, the facial expressions, voice features, and environmental information of the user are obtained through multi-modal feature extraction technology, and the weights of each modal feature are dynamically adjusted based on the signal quality scoring mechanism, making the data fusion more accurate; secondly, through scene type recognition, user attribute analysis, and interaction intention inference, a personalized interaction strategy is established, enabling the robot to dynamically match an adaptive behavior pattern according to different user groups and different scenarios, improving the naturalness and flexibility of the interaction experience.
[0078] In the present invention, by introducing a user feedback mechanism during the interaction process, real-time evaluation is carried out using the voice response duration and the degree of following of body movements, and the weight assignment of modal features is dynamically adjusted based on the comprehensive score, thereby realizing the iterative optimization of the robot's behavior pattern; compared with the traditional methods of fixed weights or rule matching, the present invention can optimize the weight assignment rule according to real-time interaction data, enabling the robot to have stronger adaptive capabilities and maintain efficient and natural human-computer interaction under different environmental conditions. Description of the Drawings
[0079] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0080] Figure 1 Schematic diagram of the method for dynamically adjusting the robot's behavior pattern according to an embodiment of the present invention;
[0081] Figure 2 Schematic diagram of the process for extracting the facial expression features of the user according to an embodiment of the present invention. Detailed Embodiments
[0082] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the accompanying drawings are only for more specifically describing the embodiments and are not intended to specifically limit the present invention.
[0083] It should be noted that in the specification, when referring to "an embodiment", "embodiment", "exemplary embodiment", "some embodiments", etc., it indicates that the described embodiment may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. Additionally, when combining an embodiment to describe a specific feature, structure or characteristic, implementing such feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge scope of those skilled in the relevant art.
[0084] Generally, terms can be understood at least in part from their use in context. For example, at least in part depending on the context, the term "one or more" used herein can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but instead, at least in part depending on the context, allowing for the existence of other factors that may not be explicitly described.
[0085] As Figure 1 - Figure 2 shown, the method for dynamically adjusting the robot behavior pattern based on multi-modal perception includes the following steps:
[0086] S1: Deploy a visual sensor, a voice sensor and an environmental sensor for real-time acquisition of visual images, voice signals and environmental parameters of the target scene;
[0087] S2: Extract features of each modality according to the data collected in S1;
[0088] S21: Perform face key point detection on the visual image and extract the user's facial expression features;
[0089] S22: Perform voiceprint analysis on the voice signal and extract the user's age features and emotional features;
[0090] S23: Perform noise filtering on the environmental parameters and extract the light intensity and temperature features;
[0091] S3: Based on the features of each modality extracted in S2, use a weighted fusion algorithm to dynamically allocate weights and generate a fused feature vector;
[0092] S4: Based on the fused feature vectors generated in S3, identify the current scene type and user attributes, where user attributes include age, emotional state, and interaction intention;
[0093] S5: Based on the scene type and user attributes identified in S4, match the corresponding interaction mode from the preset policy library, and adjust the tone, speech rate, and robot movement amplitude of the voice output;
[0094] S6: Execute the interaction mode matched in S5, and collect the user voice response duration and body movement following degree in real time as feedback data, and then update the weight allocation rule of the fused feature vectors in S3 according to the feedback data.
[0095] S1 specifically includes:
[0096] S11: Deploy visual sensors on the robot body. The visual sensors include an RGB camera and an infrared depth camera. The RGB camera is used to collect the color visual image of the target scene, and the infrared depth camera is used to obtain the depth information of the scene and the spatial position data of the target object. The acquisition frame rate is set to 30 frames per second, and the acquisition resolution is set to 1920×1080 pixels;
[0097] S12: Deploy voice sensors on the robot head or the sound source direction. The voice sensors include a MEMS microphone array. The MEMS microphone array determines the source direction of the voice signal through beamforming technology, and suppresses background noise using an adaptive filter during signal reception, and is used to collect the voice signal of the target user. The sampling rate of the voice signal is set to 16 kHz, and the sampling accuracy is set to 16 bit;
[0098] S13: Deploy environmental sensors on the robot body or inside the fixed environment. The environmental sensors include a light sensor and a temperature sensor. The light sensor adopts a photodiode array structure and is used to generate light data with a numerical range of 0 - 65535 lux by measuring the change in the light intensity of the target scene; the temperature sensor adopts a thermal resistance temperature measuring element and generates temperature data with an accuracy of ±0.1°C by measuring the change in the environmental temperature;
[0099] S14: Store the visual images, voice signals, and environmental parameter data collected in S11 to S13 in a time-synchronized manner, and mark the data frames with timestamps to ensure the consistency and timing matching of the data during the subsequent feature extraction process; Through the above steps, it is possible to ensure that the robot's multi-modal data acquisition of the target scene is accurate, real-time, and has timing consistency, providing high-quality data input for subsequent feature extraction, fusion, and behavior pattern adjustment, and improving the reliability and response speed of the overall system.
[0100] The specific steps for extracting the user's facial expression features in S21 include:
[0101] S211: Divide the visual image into multiple candidate regions, perform feature convolution calculations on each candidate region, and perform binary classification screening of human faces and non-human faces at the network output layer to obtain the face detection result;
[0102] S212: Perform face alignment operations on the detected face regions, and use the normalized affine transformation method to correct the tilt and orientation of the face. The affine transformation matrix is mapped according to the face image, and the expression of the affine transformation matrix is:
[0103]
[0104] Among them, X and Y respectively represent the pixel coordinates in the original image, X' and Y' respectively represent the normalized pixel coordinates after transformation, and A, B, C, D, E, F are the parameters of the affine transformation matrix, which are obtained by feature point matching calculation;
[0105] S213: Locate N key points on the normalized face image and obtain the pixel coordinates X n , Y n , where X n and Y n are the pixel coordinates of the nth key point;
[0106] S214: Calculate the geometric distance D m,n between adjacent or specified region key points according to the key point coordinates. Specifically, use the Euclidean distance formula to calculate the geometric distance between the mth key point and the nth key point. The formula is: Among them, D m,n represents the Euclidean distance between key points m and n, X m , Y m and X n , Y n are the pixel coordinates of key points m and n respectively;
[0107] S215: Arrange the D m,n obtained in S24 in order to form a facial expression feature vector, and perform normalization processing to map all feature data to the [0,1] interval; Through the above steps, the detection of face key points and the extraction of expression features in the visual image can ensure accuracy and maintain data consistency, providing high-quality facial feature data for subsequent multi-modal feature fusion and scene recognition.
[0108] The specific extraction of the user's age feature and emotion feature in S22 includes:
[0109] S221: Perform noise reduction processing on the voice signal, and frame and window it to extract short-time voice segments;
[0110] S222: Calculate the fundamental frequency (F0) and the first three formant frequencies (F1 - F3) of the speech signal, where the fundamental frequency range is used to distinguish age characteristics;
[0111] S223: Extract the short - time energy, zero - crossing rate, and Mel - Frequency Cepstral Coefficients (MFCC) of the speech signal,
[0112] and combine the fundamental frequency fluctuation amplitude and the speech rate change rate to quantify the emotional characteristics;
[0113] S224: Combine the fundamental frequency range, formant frequencies, and the emotional quantification results into an age feature vector and an emotional feature vector for subsequent fusion feature generation; Through the above steps, the robot can efficiently distinguish and quantify age and emotion during the analysis of the speech signal, form a more accurate user portrait information, and provide highly reliable speech data support for subsequent multi - modal feature fusion and robot behavior strategy selection.
[0114] The specific extraction of light intensity and temperature features in S23 includes:
[0115] S231: Perform time - series smoothing processing on the original light intensity signal and temperature signal collected by the environmental sensor. Specifically, use the first - order exponential smoothing algorithm to calculate the smoothed light intensity and temperature. The calculation formulas are: L s (t) = αL r (t)+(1 - α)L s (t - 1); T s (t) = βT r (t)+(1 - β)T s (t - 1); where, L s (t) and T s (t) are the smoothed light intensity and temperature at time point t respectively, L r (t) and T r (t) are the original light intensity and temperature at time point t respectively, and α and β are the smoothing coefficients, with the value range of 0.1 - 0.3 to balance the real - time performance and stability of the data;
[0116] S232: Remove the background noise from the smoothed light intensity signal, calculate the mean value L m and the standard deviation L d of the light intensity change within the continuous time window, and set the threshold θ L to be L m +2L d , remove the outliers greater than θ L , and finally obtain the denoised light intensity L f , and the calculation formula is:
[0117] Among them, L f (t) is the denoised light intensity at time point t, and θ L is the light noise rejection threshold, L m is the mean value of the light intensity, and L d is the standard deviation of the light intensity;
[0118] S233: Perform anomaly detection on the smoothed temperature signal T s , calculate the temperature change rate ΔT within the sliding window. If ΔT exceeds the set threshold θ T (the threshold is taken as 0.5 °C / s), then the temperature data T s (t) at the time point is determined to be abnormal. Finally, the window mean value is used to replace the abnormal value, and finally the denoised temperature data T f is obtained; Through the above method, the environmental parameters are smoothed and noise-filtered, effectively removing mutation values and abnormal data, ensuring the stability of the light intensity and temperature characteristics, providing high-quality environmental data support for subsequent feature fusion and robot interaction mode adjustment, and improving the adaptability and data reliability of the robot in complex environments.
[0119] S3 specifically includes:
[0120] S31: Calculate the signal quality scores for the visual features, speech features, and environmental features extracted in S2 respectively. Let the signal quality score of the visual features be Q v , the signal quality score of the speech features be Q a , and the signal quality score of the environmental features be Q e ;
[0121] The signal quality score Q v of the visual features is jointly calculated through the image sharpness and the key point detection confidence. The calculation formula is: Q v = 0.6×S c + 0.4×S k , where S c represents the image sharpness score, and S k represents the mean value of the key point confidence;
[0122] The signal quality score Q a of the speech features is jointly calculated through the signal-to-noise ratio and the fundamental frequency stability. The calculation formula is: Q a = 0.7×S s + 0.3×1V f , where S s represents the signal-to-noise ratio, and V f represents the fundamental frequency fluctuation variance;
[0123] The signal quality score Q eThrough the combined calculation of sensor data stability and the proportion of outliers, the calculation formula is: Q e = 0.5×1V w + 0.5×(1 - R e ), where V w represents the variance of the data sliding window, and R e represents the proportion of outliers;
[0124] S32: According to the quality scores of each modal signal, assign weights according to the following rules:
[0125] Visual feature weight
[0126] Voice feature weight
[0127] Environmental feature weight Among them, W v , W a , W e represent the weighted ratios of the visual, voice, and environmental modalities respectively, such that W v + W a + W e = 100%;
[0128] S33: Standardize each modal feature vector so that its numerical range is unified to 0, 1, and perform linear superposition according to the assigned weights to generate a fused feature vector F. The calculation formula is: F = W v × F v + W a × F a + W e × F e , where F v represents the normalized visual feature vector, F a represents the normalized voice feature vector, and F e represents the normalized environmental feature vector;
[0129] S34: Calculate the information loss rate of the fused feature vector. If the information loss rate ≤ 5%, it is determined as effective fusion; otherwise, reassign weights. The information loss rate is specifically calculated through the cosine similarity between the original feature and the fused feature. The formula is: Through the above steps, adopt a signal quality scoring mechanism to dynamically adjust the modal feature weights, realize highly reliable multi-modal information fusion, ensure that the final fused feature vector can maintain information integrity in different scenarios, and improve the robot's accurate recognition ability of the user's state and the stability of interaction decisions.
[0130] S4 specifically includes:
[0131] S41: Determine the scene type based on the average illumination intensity, average temperature, and ambient noise level in the fused feature vector according to the following rules;
[0132] If the average illumination intensity > 500 lux, the average temperature ≤ 25 °C, and the noise level < 50 dB, it is determined to be a home scene;
[0133] If the average illumination intensity ≤ 500 lux, the average temperature > 25 °C, and the noise level ≥ 50 dB, it is determined to be a public scene;
[0134] If the average temperature > 35 °C or the illumination intensity variance > 100 lux 2 , and the keywords "help" or "alarm" are detected in the speech features, it is determined to be an emergency scene;
[0135] S42: Determine the user's age based on the voice fundamental frequency range and formant frequency in the fused feature vector;
[0136] If the fundamental frequency > 200 Hz and the formant frequency F1 > 500 Hz, it is determined that the user is a child;
[0137] If 150 Hz ≤ fundamental frequency ≤ 200 Hz and F1 ≤ 500 Hz, it is determined that the user is an adult;
[0138] If the fundamental frequency < 150 Hz and the formant frequency F1 < 400 Hz, it is determined that the user is an elderly person;
[0139] S43: Determine the emotional state based on the facial expression features and speech emotion features in the fused feature vector;
[0140] If the mouth corner curvature of the facial expression > 0.3, the eyelid opening and closing degree > 0.8, and the short-time energy of the speech > 60 dB, the emotion is determined to be an excited state;
[0141] If the eyebrow tilt angle of the facial expression < -0.2, the eyelid opening and closing degree < 0.5, and the zero-crossing rate of the speech < 20 Hz, the emotion is determined to be a depressed state;
[0142] If the facial expression features and speech emotion features do not meet the above conditions, the emotion is determined to be a calm state;
[0143] S44: Based on the semantic keyword matching and context association in the fused feature vector, further determine the interaction intention, extract the keywords in the speech signal. If the keyword contains "temperature" or "weather", the interaction intention is environmental query; if the keyword contains "play" or "stop", the interaction intention is media control; if the proportion of negative words ("no", "cancel") detected in the continuous dialogue > 30%, the interaction intention is a termination request;
[0144] S45: Combine the scene type, user age, emotional state, and interaction intention into structured tags and input them into the adaptive behavior strategy generation module; Through the above solution, the system can perform refined classification and determination of the scene type and user attributes based on the fused feature vectors, accurately identify the user's age and emotional state under diverse environmental conditions, analyze the user's interaction intention based on keywords, and provide high-accuracy situational information support for the dynamic adjustment of subsequent behavior strategies.
[0145] S5 specifically includes:
[0146] S51: Load the basic interaction template from the policy library according to the identified scene type;
[0147] S52: Adjust the basic template parameters according to the user's age. If the user is a child, increase the voice fundamental frequency by 10%-15%, reduce the speech rate to 80% of the base value, and increase the frequency of body movements; If the user is an elderly person, reduce the voice fundamental frequency by 5%-10%, reduce the speech rate to 70% of the base value, and reduce the amplitude of body movements; If the user is an adult, keep the basic template parameters unchanged;
[0148] S53: Modify the interaction mode according to the user's emotional state. If the emotion is in an excited state, increase the speech rate by 10%, increase the movement amplitude by 20%, and insert random nodding actions; If the emotion is in a depressed state, reduce the speech rate by 15%, reduce the movement amplitude by 30%, and adopt a uniform and slow movement; If the emotion is in a calm state, maintain the current parameters;
[0149] S54: Invoke the preset instruction set according to the interaction intention. If the intention is environmental query, load the weather data interface from the policy library and output the result in a question-and-answer form; If the intention is media control, call the media player API and synchronously adjust the robot's gestures to match the play / pause actions; If the intention is a termination request, immediately stop the current task;
[0150] S55: Combine the adjusted speech rate, tone, movement amplitude, and function instructions into a complete interaction strategy and transmit it to the execution module; Through the above steps, after identifying the scene type and user attributes, it is possible to implement differentiated interaction modes based on the preset policy library for different scene conditions and user needs, and flexibly adjust the speech output and robot actions in combination with age, emotion, and intention factors to achieve a more natural and user-friendly robot behavior response and enhance the user's interaction experience in different situations.
[0151] The basic interaction templates include the home mode template, the public mode template, and the emergency mode template; Among them:
[0152] The home mode template corresponds to the home scene, sets the speech rate to 120-150 words per minute, the movement amplitude to medium, and the default tone to the kind mode;
[0153] The common mode template corresponds to the common scenario, with the speech speed set at 150 - 180 words per minute, the amplitude of movements being small, and the default tone being the formal mode;
[0154] The emergency mode template corresponds to the emergency scenario, with the speech speed set at 180 - 200 words per minute, the amplitude of movements being large, and the default tone being the hasty mode.
[0155] S6 specifically includes:
[0156] S61: Calculate the speech response duration score S t and the body movement follow - up score S m , and the formula is as follows:
[0157]
[0158] Among them, T 预设 is the standard response duration of the current interaction mode; T 实际 is the actual usage duration of the user; N 同步 is the number of frames in which the user's body movements are synchronized with the robot's movements; N 总 is the total number of frames used for detection in the entire interaction;
[0159] S62: Calculate the comprehensive feedback score S 综合 according to the predetermined weight ratio, and the formula is: S 综合 = 0.6×S t + 0.4×S m ;
[0160] S63: Adjust the weight distribution rules of each modality in S3 according to S 综合 ;
[0161] When S 综合 ≥ 80%, then keep the current weight distribution rules unchanged;
[0162] When 60% ≤ S 综合 < 80%, then increase the voice feature weight ΔW a = 80% - S 综合 × 0.5%; and reduce the environmental feature weight ΔW e = 80% - S 综合 × 0.3%;
[0163] When S 综合 < 60%, then reset the weights to the initial allocation ratio;
[0164] S64: Update the adjusted weight assignment rule to the weighted fusion algorithm in S3 for feature fusion in the next interaction cycle. Through the above steps, it is possible to adaptively adjust the weights of multimodal features based on the user's speech response duration and limb movement follow-up degree, quickly capture the user feedback signal while ensuring the basic interaction fluency, and achieve dynamic optimization of the weight assignment rule.
[0165] This invention covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To enable the public to have a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments of this invention. However, those skilled in the art can fully understand this invention without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of this invention.
[0166] The above are only the preferred embodiments of this invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of this invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this invention.
Claims
1. A method for dynamically adjusting the behavior pattern of a robot based on multi-modal perception, characterized in that, It includes the following steps: S1: Deploy visual sensors, voice sensors and environmental sensors to collect visual images, voice signals and environmental parameters of the target scene in real time; S2: Extract multi-modal features according to the data collected in S1; S21: Perform face key point detection on the visual image and extract the user's facial expression features; S22: Perform voiceprint analysis on the voice signal and extract the user's age features and emotional features; S23: Filter the environmental parameters for noise and extract the light intensity and temperature features; S3: Based on the multi-modal features extracted in S2, use a weighted fusion algorithm to dynamically allocate weights and generate a fused feature vector; S4: Identify the current scene type and user attributes according to the fused feature vector generated in S3, where the user attributes include age, emotional state and interaction intention; S5: Based on the scene type and user attributes identified in S4, match the corresponding interaction mode from the preset policy library and adjust the tone, speech rate and robot movement amplitude of the voice output; S6: Execute the interaction mode matched in S5, and collect the user's voice response duration and body movement following degree in real time as feedback data, and then update the weight allocation rule of the fused feature vector in S3 according to the feedback data.
2. The method for dynamically adjusting the robot behavior pattern based on multi-modal perception according to claim 1, wherein The specific content of S1 includes: S11: Deploy a visual sensor on the robot body. The visual sensor includes an RGB camera and an infrared depth camera. The RGB camera is used to collect the color visual image of the target scene, and the infrared depth camera is used to obtain the depth information of the scene and the spatial position data of the target object. The acquisition frame rate is set to 30 frames per second, and the acquisition resolution is set to 1920×1080 pixels; S12: Deploy a voice sensor on the robot's head or the sound source direction to collect the voice signal of the target user. The sampling rate of the voice signal is set to 16kHz, and the sampling accuracy is set to 16bit; S13: Deploy an environmental sensor on the robot body or inside the fixed environment. The environmental sensor includes a light sensor and a temperature sensor. The light sensor is used to generate light data with a numerical range of 0-65535lux; the temperature sensor generates temperature data with an accuracy of ±0.1°C; S14: Store the visual image, voice signal and environmental parameter data collected in S11 to S13 in a time-synchronized manner and mark the data frames with time stamps.
3. The method for dynamically adjusting the robot behavior pattern based on multi-modal perception according to claim 1, wherein The specific content of S21 includes: S211: Divide the visual image into multiple candidate regions, perform feature convolution calculation on each candidate region, and perform binary classification screening of face and non-face at the network output layer to obtain the face detection result; S212: Perform face alignment operation on the detected face region, and use the normalized affine transformation method to correct the tilt and orientation of the face, where the affine transformation matrix is mapped according to the face image; S213: Locate N key points on the standardized face image and obtain the pixel coordinates of each key point; S214: Calculate the geometric distance D between adjacent or specified area key points based on the key point coordinates m,n ; S215: Arrange the D obtained in S24 m,n in sequence to form a facial expression feature vector, and perform normalization processing so that all feature data are mapped to the interval [0, 1].
4. The method for dynamically adjusting the robot behavior pattern based on multi-modal perception according to claim 1, characterized in that The specific content of S22 includes: S221: Perform noise reduction processing on the voice signal and add windows to each frame to extract short-time voice segments; S222: Calculate the fundamental frequency and the first three formant frequencies of the speech signal, where the fundamental frequency range is used to distinguish age characteristics; S223: Extract the short-time energy, zero-crossing rate, and Mel-frequency cepstral coefficients of the speech signal, and combine the fundamental frequency fluctuation amplitude and the speech rate change rate to quantify the emotional characteristics; S224: Combine the fundamental frequency range, formant frequencies, and the emotional quantification results into an age feature vector and an emotional feature vector.
5. The method for dynamically adjusting the robot behavior pattern based on multi-modal perception according to claim 1, wherein The specific steps of S23 include: S231: Perform time series smoothing processing on the original light intensity signal and temperature signal collected by the environmental sensor; S232: Remove the background noise from the smoothed light intensity signal, and calculate the mean value L of the light intensity change within a continuous time window m and the standard deviation L d , and set the threshold θ L to be L m + 2L d , remove the outliers greater than θ L , and finally obtain the denoised light intensity L f ; S233: Perform anomaly detection on the smoothed temperature signal T s Calculate the temperature change rate ΔT within the sliding window. If ΔT exceeds the set threshold θ T , then the temperature data T s (t) at the time point is determined to be abnormal. Finally, replace the abnormal value with the window mean to finally obtain the denoised temperature data T f .
6. The method for dynamically adjusting the robot behavior pattern based on multi-modal perception according to claim 1, characterized in that The specific steps of S3 include: S31: Calculate the signal quality scores for the visual features, speech features, and environmental features extracted in S2 respectively. Let the signal quality score of the visual features be Q v and the signal quality score of the speech features be Q a and the signal quality score of the environmental features be Q e ; S32: According to the quality scores of each modality signal, assign weights according to the following rules: Visual feature weight Voice feature weight Environmental feature weight Among them, W v , W a , W e respectively represent the weighted ratios of the visual, speech, and environmental modalities, such that W v + W a + W e = 100%; S33: Standardize each modality feature vector so that its numerical range is unified to [0, 1], and perform linear superposition according to the assigned weights to generate a fused feature vector F; S34: Calculate the information loss rate of the fused feature vector. If the information loss rate ≤ 5%, it is determined as an effective fusion; otherwise, reassign the weights.
7. The method for dynamically adjusting the robot behavior pattern based on multi-modal perception according to claim 1, wherein The specific steps of S4 include: S41: Based on the average light intensity, average temperature, and environmental noise level in the fused feature vector, determine the scene type according to the following rules; If the average light intensity > 500 lux, the average temperature ≤ 25 °C, and the noise level < 50 dB, it is determined as a home scene; If the average light intensity ≤ 500 lux, the average temperature > 25 °C, and the noise level ≥ 50 dB, it is determined as a public scene; If the average temperature > 35°C or the variance of the light intensity > 100 lux 2 , and keywords "help" or "alarm" are detected in the voice features, it is determined as an emergency scenario; S42: Based on the fundamental frequency range and formant frequencies in the fused feature vector, determine the user's age; If the fundamental frequency > 200 Hz and the formant frequency F1 > 500 Hz, it is determined that the user is a child; If 150 Hz ≤ the fundamental frequency ≤ 200 Hz and F1 ≤ 500 Hz, it is determined that the user is an adult; If the fundamental frequency < 150 Hz and the formant frequency F1 < 400 Hz, it is determined that the user is an elderly person; S43: Based on the facial expression features and speech emotional features in the fused feature vector, determine the emotional state; If the mouth corner curvature of the facial expression > 0.3, the eyelid opening and closing degree > 0.8, and the short-time energy of the speech > 60 dB, it is determined that the emotion is an excited state; If the eyebrow tilt angle of the facial expression < -0.2, the eyelid opening and closing degree < 0.5, and the zero-crossing rate of the speech < 20 Hz, it is determined that the emotion is a depressed state; If the facial expression features and speech emotional features do not meet the above conditions, it is determined that the emotion is a calm state; S44: Based on the semantic keyword matching and context association in the fused feature vector, and then determine the interaction intention. Extract the keywords in the speech signal. If the keyword contains temperature or weather, the interaction intention is an environment query; if the keyword contains play or stop, the interaction intention is media control; if the proportion of negative words detected in the continuous dialogue > 30%, the interaction intention is a termination request; S45: Combine the scene type, user age, emotional state, and interaction intention into a structured label and input it into the adaptive behavior strategy generation module.
8. The method for dynamically adjusting the robot behavior pattern based on multi-modal perception according to claim 1, characterized in that The specific steps of S5 include: S51: Load the basic interaction template from the policy library according to the identified scene type; S52: Adjust the basic template parameters according to the user's age. If the user is a child, increase the fundamental frequency of the voice by 10%-15% and reduce the speech rate to 80% of the base value; if the user is an elderly person, reduce the fundamental frequency of the voice by 5%-10% and reduce the speech rate to 70% of the base value; if the user is an adult, keep the basic template parameters unchanged; S53: Modify the interaction mode according to the user's emotional state. If the emotion is in an excited state, increase the speech rate of the voice by 10%, increase the amplitude of the action by 20%, and insert a random nodding action; if the emotion is in a depressed state, reduce the speech rate of the voice by 15%, reduce the amplitude of the action by 30%, and adopt a uniform and slow movement; if the emotion is in a calm state, maintain the current parameters; S54: Invoke the preset instruction set according to the interaction intention. If the intention is environmental query, load the weather data interface from the policy library and output the result in the form of a question and answer; if the intention is media control, call the media player API and synchronously adjust the robot's gestures to match the play / pause action; if the intention is a termination request, immediately stop the current task; S55: Combine the adjusted speech rate, tone, action amplitude, and function instructions into a complete interaction strategy and transmit it to the execution module.
9. The method for dynamically adjusting the robot behavior pattern based on multi-modal perception according to claim 8, wherein, The basic interaction template includes a home mode template, a public mode template, and an emergency mode template; among them: The home mode template corresponds to the home scene, sets the speech rate to 120-150 words per minute, the action amplitude to medium, and the default tone to the kind mode; The public mode template corresponds to the public scene, sets the speech rate to 150-180 words per minute, the action amplitude to a small amplitude, and the default tone to the formal mode; The emergency mode template corresponds to the emergency scene, sets the speech rate to 180-200 words per minute, the action amplitude to a large amplitude, and the default tone to the hasty mode.
10. The method for dynamically adjusting the robot behavior pattern based on multi-modal perception according to claim 1, wherein The specific content of S6 includes: S61: Calculate the voice response duration score S based on the user feedback data collected in real time t and the limb movement follow-up score S m ; S62: Calculate the comprehensive feedback score S according to a predetermined weight ratio 综合 , the formula is: S 综合 = 0.6 × S t + 0.4 × S m ; S63: According to S 综合 Adjust the modal weight allocation rules in S3; When S 综合 ≥ 80%, the current weight allocation rule remains unchanged; When 60% ≤ S 综合 <80%, the voice feature weight ΔW is increased a =(80% - S 综合 ) × 0.5%; and the environmental feature weight ΔW is decreased e =(80% - S 综合 ) × 0.3%; When S 综合 < 60%, the reset weight is the initial allocation ratio; S64: Update the adjusted weight distribution rule to the weighted fusion algorithm in S3 for feature fusion in the next interaction cycle.
Citation Information
Cited By
Multi-modal intelligent robot system and interaction method
CN120588256A
Morality behavior demonstration device and system based on humanoid robot
CN120735065A
Smart phone and method
CN120856822A
Action generation method and device based on multi-modal emotion perception, equipment and medium
CN120872157A
Motion generation method and device based on multi-modal emotion perception, equipment and medium
CN120872157B