Humanoid robot situational social behavior expression method
By constructing user behavior and emotional perception steps, combining sensors and neural networks to recognize voice emotions, humanoid robots can naturally express emotions in multiple situations, solving the problem of insufficient emotional interaction in the existing technology, and improving the flexibility and nature of human-computer interaction.
Patent Information
- Application Number
- CN202510382214.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, humanoid robots lack emotional interaction ability in multiple situations, and it is difficult to express emotions naturally in social tasks, which affects their application in education, medical care and other fields.
By constructing user behavior and emotion perception steps, combining sensors to perceive user location and voice commands, using neural networks to recognize voice emotions, generate multimodal social behavior decisions, and realize emotional expression of humanoid robots in different situations.
It improves the flexible interaction capabilities of human-like robots in multiple situations, enhances the naturalness and personalized experience of human-computer interaction, and is suitable for human-computer interaction in actual service environments.
Smart Images

Figure CN120236578A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robots, and specifically to a method for expressing context-aware social behaviors of a humanoid robot based on user behavior and emotion perception. This method takes into account the social tasks and social emotion expressions in the interaction process of the humanoid robot. Background Art
[0002] Good human-computer interaction capabilities are particularly important for the further widespread application of humanoid robots in multiple fields such as education, medical care, and public services. Due to the similar appearance of humanoid robots to humans, users often have expectations such as intelligence and natural interaction. This requires humanoid robots to not only perform tasks in specific contexts but also possess the ability to express emotions related to the context and users.
[0003] Current research on the behavior expression of humanoid robots mainly focuses on aspects such as the speech expression of humanoid robots (such as speech synthesis) and action expression (such as robot behavior control), and less consideration is given to the context-aware social expression of humanoid robots. In the actual environment, the interaction between humans and humanoid robots occurs in multiple contexts, and these contexts may have temporal continuity or spatial consistency. Therefore, for a humanoid robot, it should be able to handle multiple context tasks in a specific environment. At the same time, in the process of human-computer interaction, emotional interaction helps to improve the efficiency of human-computer interaction and the understanding of interaction intentions. There are also certain deficiencies in the existing technology in the research on the social behavior expression of humanoid robots that integrates emotions, lacking a social behavior expression method that can consider both multi-context applications and emotional interaction.
[0004] To solve the above problems, a method for expressing context-aware social behaviors of a humanoid robot based on user behavior perception is proposed, which improves the flexible interaction ability of the humanoid robot in multiple contexts and also establishes the interaction between humans and humanoid robots from an emotional perspective, providing help for the research on flexible and natural interaction of humanoid robots. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method for expressing anthropomorphic robot's contextual social behavior in view of the deficiencies of the above-mentioned prior art. This method constructs the relationships among the interaction context, interaction emotion, and the social behavior of the anthropomorphic robot in the human-computer interaction environment. By combining user behavior and emotion perception, the anthropomorphic robot's contextual social behavior decision-making, and the anthropomorphic robot's multi-modal social behavior expression, this method enables the anthropomorphic robot to perceive the interaction context and thus express more natural social behaviors that integrate emotion and task implementation, enhancing the emotional interaction between humans and anthropomorphic robots in the actual interaction task environment, improving the flexible interaction ability of the anthropomorphic robot in multiple contexts, and also establishing the interaction between humans and anthropomorphic robots from an emotional perspective. This method can provide certain technical support for the situation where robots replace humans in the actual service environment and enhance the user experience in interacting with robots.
[0006] To achieve the above technical objectives, the technical solution adopted by the present invention is as follows:
[0007] A method for expressing anthropomorphic robot's contextual social behavior, including steps of user behavior and emotion perception, the anthropomorphic robot's contextual social behavior decision-making, and the anthropomorphic robot's multi-modal social behavior expression;
[0008] The steps of user behavior and emotion perception include steps of user location perception S101, user voice command recognition S102 - S105, and user voice emotion prediction S106 - S108, including:
[0009] S101: Sense the user's location through sensors and define the categories of interaction contexts;
[0010] S1011: Sense the user's location in the interactive area in the physical environment through sensors;
[0011] S1012: Sense the distance between the user and the anthropomorphic robot through sensors and set three thresholds, which are respectively used as the recognition conditions for different contexts to occur. The three thresholds are, from large to small, threshold one, threshold two, and threshold three;
[0012] S1013: Define five categories of interaction contexts, namely attention, welcome, indication, apology, and obstruction contexts. When the distance obtained in step S1012 is between threshold one and threshold two, the context is the attention context; when the distance is between threshold two and threshold three, the context is the welcome context; when the distance is less than threshold three, the context is the indication, apology, or obstruction context. Specifically, whether it is the indication, apology, or obstruction context is identified by step S105;
[0013] S102: Collect the user's voice signal and perform filtering processing;
[0014] S103: Extract voice feature parameters and construct a voice feature vector;
[0015] S104. Identify the vocabulary of the user's voice command;
[0016] S1041. Pre-define a voice recognition sample set, which is divided into three types of sample sets: instruction, apology, and obstruction, according to the situation. Each sample set contains multiple words that will appear in that situation;
[0017] S1042. Train the voice feature parameter template of the sample set according to the steps of S103;
[0018] S105. Perform user voice command recognition: According to the reference template of S1042 and the user voice command, perform time warping on the user voice feature vector to obtain the optimal warping path, and obtain the situation category to which the user voice command belongs;
[0019] S106. Construct voice emotion features and construct a voice emotion feature template;
[0020] S107. Train a neural network using the data in the voice emotion feature template to obtain a trained neural network model;
[0021] After the user audio passes through the steps 1061 - 1065 in S102, S103, and S106, the voice emotion features corresponding to each frame of voice are obtained (including the cepstrum and its first-order difference parameters extracted in S1033, the short-time energy, fundamental frequency, first-order fundamental frequency jitter, and first formant parameters extracted in steps S1061 - S1063). Input the voice emotion features corresponding to each frame of voice into the trained neural network model in S107 to obtain the prediction of the emotion category of the user audio;
[0022] Among them, the decision-making steps for the anthropomorphic robot's contextual social behavior include:
[0023] S201. Generate an interactive situation category code according to the output of step S1013 and the output of step S105;
[0024] S202. Pre-define a dataset of the anthropomorphic robot's contextual social actions. This dataset includes the contextual social task actions of five situation categories, namely attention, welcome, instruction, apology, and obstruction situations;
[0025] S203. Define social emotion action sub-categories in each contextual social task category respectively;
[0026] S204. Taking the emotion category predicted in step S108 as a condition, according to the output of S201, search for the social emotion action sub-category of the anthropomorphic robot that matches the emotion category predicted in step S108 under the interactive situation category; and output the anthropomorphic robot social emotion action code through the search result;
[0027] S205. Pre-define voice templates, defining voice tags and light mode tags in scenarios of welcome, instruction, apology, and obstruction; according to the output of step S105, output the voice tags in the scenario category to which the user voice command belongs and encode them, output the light mode tags in the scenario category to which the user voice command belongs, and perform light mode encoding;
[0028] S206. Integrate the encodings of steps S204 and S205, and define them as the encoded context-aware social behaviors of the humanoid robot;
[0029] Among them, the multi-modal social behavior expression steps of the humanoid robot include:
[0030] S301. According to the output of step S206, pre-define the multi-modal social behavior template of the humanoid robot;
[0031] S302. Adopt the TCP / IP protocol, and preset the data packet integration order as: action data, light mode data, voice data, and send the integrated context-aware social behavior data packet of the humanoid robot to the humanoid robot;
[0032] S303. The humanoid robot receives the data packet and decodes and extracts the interaction content;
[0033] S304. Through the API, call the pre-defined action library and voice library in the humanoid robot, and perform light mode switching to complete the multi-modal context-aware social behavior expression.
[0034] As a further improved technical solution of the present invention, the sensor adopted in S1011 is an infrared pyroelectric sensor, and the sensor adopted in S1012 is an ultrasonic sensor;
[0035] The specific content of S102 is as follows:
[0036] S1021. Collect the user voice signal through a microphone;
[0037] S1022. Amplify and filter the collected signal;
[0038] S1023. Through a USB data acquisition card, transmit the signal after step S1022 to the PC side;
[0039] S1024. Adopt a filtering algorithm to filter the audio signal to further reduce the signal noise.
[0040] As a further improved technical solution of the present invention, the specific content of S103 is as follows:
[0041] S1031. Perform pre-emphasis, framing, and windowing on the voice, and perform voice endpoint detection;
[0042] S1032. Perform fast Fourier transform on each frame of signal and calculate the spectral line energy;
[0043] S1033. Use an auditory filter bank to calculate the speech energy spectrum and discrete cosine transform cepstrum of each frame, extract the cepstral features and their first-order difference parameters, and construct a speech feature vector.
[0044] As a further improved technical solution of the present invention, the S105 is specifically as follows:
[0045] S1051. Denote the speech feature vector in the reference template of S1042 as the reference speech feature vector, and denote the speech feature vector corresponding to the user speech command obtained after step S1033 as the user speech feature vector; assume that the reference speech feature vector has J frame vectors and the user speech feature vector has I frame vectors;
[0046] Perform time warping on the user speech feature vector, and the formula is:
[0047]
[0048] where d[T(i),R(ω(i))] is the distance between the i-th frame user speech feature vector T(i) and the j-th frame reference speech feature vector R(j), j = ω(i), and D is the vector distance under the time warping optimization condition;
[0049] S1052. When performing time warping in S1051, for the i-th frame user speech feature vector T(i) and the j-th frame reference speech feature vector R(j), use the reverse decision to obtain the optimal warping path, and then obtain the context category to which the user speech command belongs. Assume that their vector lengths are t and r respectively. For each matching R(j), all possible points that can reach its current value need to be considered. The minimum distance from (t,r) to (t,r - 1), (t - 1,r - 1), (t - 1,r) is the current cost. Calculate the total cost function from (t,r) to (1,1) to obtain the best warping path.
[0050] As a further improved technical solution of the present invention, to reduce the computational amount, in the calculation process of S1052, an adaptive hierarchical matching method is performed, and the one-time full-length matching is changed to two-time matching with an adaptive threshold. The S1052 is specifically as follows:
[0051] S10521. Use the user speech feature vector corresponding to the user speech command extracted in step S1033 as the test sample, and the reference speech feature vector in the reference template of S1042 as the training sample. When calculating the distance between the test sample and the training sample, randomly extract several frames and calculate the cumulative distance through formula (1), and set this cumulative distance as D c ;
[0052] S10522. Determine the cumulative distance D according to S10521 c Determine the adaptive threshold condition and perform adaptive layering; the threshold condition adopted in the present invention is
[0053] S10523. Match and calculate the test sample with the training sample. When calculating the cumulative distance, when the cumulative distance D c is less than the threshold, store the training sample in the new template library, and store the cumulative distance and the calculated frame length; if the cumulative distance is greater than or equal to the threshold, ignore the training sample;
[0054] S10524. Define the new template library obtained in S10523 as the first-layer matching; call the cumulative distance and frame length obtained in S10523, calculate the cumulative distance of the remaining frame length, and select the optimal regularization path with the minimum cumulative distance to obtain the matching result of the test sample, that is, the context category to which the user voice command belongs.
[0055] As a further improved technical solution of the present invention, the S106 is specifically as follows:
[0056] S1061. Extract the short-time energy of the audio;
[0057] S1062. Extract the fundamental frequency and the first-order fundamental frequency jitter of the audio by using the autocorrelation function method;
[0058] S1063. Extract the first formant parameter of the audio by using the linear prediction method;
[0059] S1064. Combine the cepstrum and its first-order difference parameters extracted in S1033, the short-time energy, fundamental frequency, first-order fundamental frequency jitter and first formant parameter extracted in steps S1061 to S1063 to construct the speech emotion feature;
[0060] S1065. Perform dimensionality reduction on the speech emotion feature by using linear discriminant analysis;
[0061] S1066. Use the CASIA emotional speech database to generate the speech emotion feature template according to steps S1061 to S1065.
[0062] As a further improved technical solution of the present invention, the S107 is specifically as follows:
[0063] S1071. Construct the topology structure of the backpropagation neural network;
[0064] S1072. Train the normalized speech emotion feature sample data in step S1066;
[0065] S1073, initializing the network structure, using the normalized speech emotion feature sample data obtained in step S1072 as the input of the neural network;
[0066] S1073, encoding the speech of the CASIA emotional speech database into emotion categories, and using the encoded samples as the output of the neural network, wherein the corresponding emotion categories of the output codes include happiness, neutrality, surprise, anger, fear and sadness;
[0067] S1074, using logistic regression logsig as the neural network activation function;
[0068] S1075, initializing the learning rate, smoothing parameter, cumulative sum of squared gradients, and mean square error vector, and setting the maximum number of training iterations;
[0069] S1076, calculate forward propagation;
[0070] S1077, calculating propagation error and back propagation gradient;
[0071] S1078. Optimize the adaptive learning rate to update the optimal solution of the neural network weights and biases. When the iteration reaches the maximum number of iterations set in step S1075 or reaches the target error, an optimized neural network model is obtained, that is, a trained neural network model.
[0072] As a further improved technical solution of the present invention, the S201 is specifically as follows:
[0073] S2011, generating the first level code of the interaction situation category outputted in step S1013;
[0074] S2012, generating the second level code of the interaction situation category outputted in step S105;
[0075] S2013, integrating the first layer and the second layer coding to obtain the interaction context category coding;
[0076] The S203 is specifically:
[0077] Combined with the attribute characteristics of the humanoid robot situational social tasks, social emotion action subcategories are defined in each situational social task category, and the social emotion action subcategories include neutral, happy, sad, angry and / or fear. However, depending on the characteristics of the interaction situation, the specific division of subcategories under different situational social task action categories is different. For example, under the welcome situational social task action category, the subcategories include neutral and happy. The neutral subcategory is a social emotion action subcategory included in each situational social task category.
[0078] As a further improved technical solution of the present invention, the S204 is specifically as follows:
[0079] S2041: Taking the emotional category predicted in step S108 as a condition, according to the output of S201, search for the sub-category of the social emotional actions of the humanoid robot that matches the emotional category predicted in step S108 under the interaction situation category;
[0080] S2042: If found, output the social emotional action code of the humanoid robot in this sub-category of social emotional actions;
[0081] S2043: If not found, output the social emotional action code of the humanoid robot in the neutral sub-category.
[0082] As a further improved technical solution of the present invention, the specific content of S205 is as follows:
[0083] S2051: Pre-define voice templates, and define voice labels and LED light mode labels in the situations of welcome, instruction, apology, and obstruction;
[0084] S2052: According to the output of step S103, output the voice label in the welcome situation and encode it;
[0085] S2053: According to the output of step S105, output the voice label in the situation category to which the user voice instruction belongs and encode it;
[0086] S2054: According to the output of step S105, output the LED light mode label in the situation category to which the user voice instruction belongs, and perform LED light mode encoding.
[0087] The beneficial effects of the present invention are as follows:
[0088] (1) The present invention can combine the user's location, voice instructions, and voice emotion perception, fully consider the implementation of social tasks and social emotional communication in the human-computer interaction situation, enable the robot to give social task behaviors with emotional meanings, improve the naturalness of the robot when performing social tasks, and enhance human-computer interaction and understanding.
[0089] (2) The present invention constructs a multi-layer perception structure of the interaction situation based on the user's behavior, enabling the robot's perception ability of the situation to change with the occurrence and implementation of the user's behavior.
[0090] (3) For different interaction situations, the present invention enables the humanoid robot to adaptively give feedback that not only executes the interaction task but also expresses a certain emotion, making the robot more applicable to the actual environment.
[0091] (4) By combining the adaptive hierarchical speech isolated word recognition method with the speech emotion recognition method based on an optimized neural network, the humanoid robot of the present invention can give different multi-modal social behavior expressions for different users in the same situation, enhancing the personalized and flexible experience in human-computer interaction.
[0092] (5) The present invention uses a variety of sensors to perceive user behavior, thereby constructing a multi-layer interactive situation awareness structure and establishing the association between the human-computer interaction situation and user behavior; adopting adaptive hierarchical isolated word recognition and speech emotion recognition based on an optimized neural network, making full use of user speech features, constructing the matching relationship between user emotion and the emotional expression of the robot's contextual social actions, and proposing a decision-making for the robot's contextual social behavior expression that integrates emotion, so as to realize the multi-modal social behavior expression of the humanoid robot that can express emotion while performing interactive tasks. The present invention not only constructs the relationship between user behavior and interaction situation, and between interaction situation and the social behavior of the humanoid robot, but also solves the problem of less emotional interaction when the humanoid robot performs social tasks in the prior art. It fully considers the environmental situation and user emotion, and proposes an implementation path for the multi-situation social expression of the humanoid robot, providing feasible help for the application of the humanoid robot in various actual situations and for enhancing the user's personalized and flexible natural experience. Description of the Drawings
[0093] Figure 1 is a flow chart of the contextual social behavior expression of a humanoid robot based on user behavior provided by the present invention.
[0094] Figure 2 is the details of the sensor perceiving the user provided by the embodiment of the present invention.
[0095] Figure 3 is the multi-layer interactive situation awareness structure provided by the present invention.
[0096] Figure 4 is the system communication scheme provided by the present invention.
[0097] Figure 5 is a flow chart of the social action decision-making of the humanoid robot provided by the present invention.
[0098] Figure 6 is a flow chart of the multi-modal expression of the contextual social behavior of the humanoid robot provided by the present invention. Detailed Embodiments
[0099] The following further describes the specific implementation manners of the present invention with reference to the accompanying drawings to provide a full understanding of the method of the present application. The described features, structures or characteristics can be applied to any humanoid robot's emotional behavior expression. The flowcharts shown in the accompanying drawings do not necessarily include all contents and operation steps. For example, some steps / operations can be decomposed, some steps / operations can be combined or partially combined, so the actual execution order may be changed according to the actual situation.
[0100] This embodiment provides a method for a humanoid robot's contextual social behavior expression. Figure 1 It is the overall flowchart of a method for a humanoid robot's contextual social behavior expression provided by an embodiment of the present invention. It mainly includes steps of user behavior and emotion perception, steps of a humanoid robot's contextual social behavior decision-making, and steps of a humanoid robot's multi-modal social behavior expression.
[0101] Among them, the steps of user behavior and emotion perception include step S101 of user position perception, steps S102 - S105 of user voice command recognition, and steps S106 - S108 of user voice emotion prediction, which are specifically as follows:
[0102] S101. Sense the user's position through sensors.
[0103] S1011. Sense that the user is in an interactive area in the physical environment through an infrared pyroelectric sensor.
[0104] S1012. Sense the distance between the user and the humanoid robot through an ultrasonic sensor, and set three thresholds, which are respectively used as recognition conditions for different situations. The three thresholds are, from large to small, threshold one, threshold two, and threshold three.
[0105] S1013. Define five categories of interaction situations, namely attention, welcome, indication, apology, and blocking situations. When the distance obtained in step S1012 is between threshold 1 and threshold 2, the situation is an attention situation; when the distance is between threshold 2 and threshold 3, the situation is a welcome situation; when the distance is less than threshold 3, the situation is an indication, apology, or blocking situation, and specifically which of the indication, apology, or blocking situations it is is recognized by step S105.
[0106] S102. Collect the user's voice signal through a microphone and perform filtering processing.
[0107] S1021. Collect the user's voice signal through a microphone.
[0108] S1022. Amplify and filter the collected signal.
[0109] S1023. Transmit the signal after step S1022 to the PC side through a USB data acquisition card.
[0110] S1024. Filter the audio signal using a filtering algorithm to further reduce signal noise.
[0111] S103. Extract speech feature parameters and construct a speech feature vector.
[0112] S1031. Perform pre-emphasis, framing, and windowing, and perform voice activity detection.
[0113] S1032. Perform a fast Fourier transform on each frame of the signal and calculate the spectral line energy.
[0114] S1033. Use an auditory filter bank to calculate the speech energy spectrum and discrete cosine transform cepstrum for each frame, extract cepstral features and their first-order difference parameters, and construct a speech feature vector.
[0115] S104. Recognize the user's voice command vocabulary.
[0116] S1041. Pre-define a speech recognition sample set, divide it into three types of sample sets: instructions, apologies, and blocks according to the context, and each sample set contains multiple vocabulary that will appear in that context.
[0117] S1042. Train the speech feature parameter template of the sample set according to step S103.
[0118] S105. Perform user voice command recognition.
[0119] S1051. Assume that the reference template in S1042 has J frame vectors, and the user voice command in S1033 has I frame vectors. Perform time warping on the user audio vector obtained after step S103. The main formula is:
[0120]
[0121] where d[T(i),R(ω(i))] is the distance between the i-th frame user audio vector T(i) and the j-th frame template vector R(j), j = ω(i), and D is the vector distance under the time warping optimization condition.
[0122] S1052. When performing time warping in S1051, for the i-th frame user audio vector T(i) and the j-th frame template vector R(j), use reverse decision to obtain the optimal warping path, and then obtain the context category to which the user voice command belongs.
[0123] To reduce the computational complexity, during the calculation process of S1052, perform an adaptive hierarchical matching method, changing the one-time full-length matching to two-time matching with an adaptive threshold.
[0124] After extracting the user's speech feature parameters, use them as test samples. When calculating the distance by matching with the training samples, randomly select several frames to calculate the cumulative distance, and set this cumulative distance as D c 。
[0125] S10522. Perform adaptive stratification according to the cumulative distance Dc in S10521. In this embodiment, the threshold condition is
[0126] S10523. Match the test samples with the training samples. When calculating the cumulative distance, if the cumulative distance is less than the threshold, store this training sample in the new template library, and store the cumulative distance and the calculated frame length. If the cumulative distance is greater than or equal to the threshold, ignore this training sample.
[0127] S10524. Define the new template library obtained in S10523 as the first-level matching. Call the cumulative distance and frame length obtained in S10523, calculate the cumulative distance of the remaining frame lengths, and select the regularization path with the minimum cumulative distance as the condition to obtain the matching result of the test sample.
[0128] S106. Construct a speech emotion feature template.
[0129] S1061. Extract the short-time energy of the audio.
[0130] S1062. Use the autocorrelation function method to extract the fundamental frequency and the first-order fundamental frequency jitter of the audio.
[0131] S1063. Use the linear prediction method to extract the first formant parameter of the audio.
[0132] S1064. Combine the cepstrum and its first-order difference parameters extracted in S1033, and the steps of S1061 - S1063 to construct the speech emotion feature.
[0133] S1065. Perform dimensionality reduction of the speech emotion feature using linear discriminant analysis.
[0134] S1066. Use the CASIA emotional speech database to generate a speech emotion feature template from the steps of S1061 - S1065.
[0135] S107. Use the neural network model method to identify the user's speech emotion.
[0136] S1071. Construct the topology of the backpropagation neural network;
[0137] S1072. Train the normalized speech emotion feature sample data in step S1066;
[0138] S1073. Initialize the network structure, and use the normalized speech emotion feature sample data obtained in step S1072 as the input of the neural network;
[0139] S1073. Encode the emotions of the voices in the CASIA emotion speech database, and use the encoded samples as the output of the neural network. The corresponding emotion categories of the output encoding include happy, neutral, surprised, angry, fearful, and sad.
[0140] S1074. Use the logistic regression logsig as the activation function of the neural network;
[0141] S1075. Initialize the learning rate, smoothing parameter, cumulative sum of gradient squares, mean square error vector, and set the maximum number of training iterations;
[0142] S1076. Calculate the forward propagation;
[0143] S1077. Calculate the propagation error and the backpropagation gradient;
[0144] S1078. Optimize the adaptive learning rate to update the optimal solutions of the neural network weights and biases. When the iteration reaches the maximum number of iterations set in step S1075 or reaches the target error, an optimized neural network model is obtained.
[0145] S108. After the user audio (user speech) passes through S1061 - S1065 in steps S102, S103, and S106, the speech emotion features of each frame are obtained. The speech emotion features of each frame are input into the neural network model optimized in S107 to obtain the prediction result of the user speech emotion category.
[0146] Among them, the steps of the anthropomorphic robot's contextual social behavior decision include:
[0147] S201. Generate the interactive context category encoding.
[0148] S2011. Generate the first - layer encoding of the interactive context category output in step S1013.
[0149] S2012. Generate the second - layer encoding of the interactive context category output in step S105.
[0150] S2013. Combine the first - layer and second - layer encodings to obtain the interactive context category encoding.
[0151] S202. Pre - define the anthropomorphic robot's contextual social action dataset, which includes five categories of contextual social task actions, namely attention, welcome, indication, apology, and obstruction contexts.
[0152] S203. Combine the attribute characteristics of the humanoid robot's situational social tasks, and define social emotional action subcategories in each situational social task category. The subcategories generally include neutral, happy, sad, angry, and fearful. However, according to different interaction situation characteristics, there are differences in the specific classification of subcategories under different situational social task action categories. For example, under the welcome situational social task action category, the subcategories include neutral and happy. The neutral subcategory is the social emotional action subcategory included in each situational social task category.
[0153] S204. Taking the emotional category predicted in step S108 as a condition, and according to the output of S201, search for the humanoid robot's social emotional action subcategory that matches the emotional category predicted in step S108 under the interaction situation category.
[0154] S2041. If found, output the humanoid robot's social emotional action code in this subcategory.
[0155] S2042. If not found, output the humanoid robot's social emotional action code in the neutral subcategory.
[0156] S205. Pre-define voice templates, and define voice labels and LED light mode labels in the situations of welcome, instruction, apology, and obstruction.
[0157] S2051. According to the output of step S103, output the voice label in the welcome situation and encode it.
[0158] S2052. According to the output of step S105, output the voice label corresponding to the voice instruction and encode it.
[0159] S2053. According to the output of step S105, output the LED light mode label. For example, in the obstruction mode, output the red LED light label and perform LED light mode encoding.
[0160] S206. Integrate the encodings of steps S204 and S205, and define them as the humanoid robot's contextualized social behavior encoding;
[0161] Among them, the multi-modal social behavior expression steps of the humanoid robot include:
[0162] S301. According to step S206, pre-define the humanoid robot's multi-modal social behavior template (used to store the preset data packet in S302);
[0163] S302. Adopt the TCP / IP protocol, and the preset data packet fusion order is: action data, light mode data, voice data. Send the fused humanoid robot's contextualized social behavior data packet (i.e., the encoded data packet) to the humanoid robot.
[0164] S303. The humanoid robot receives the data packet and decodes it to extract the interaction content.
[0165] S304. Call the predefined behavior action library file and voice library file through the API, and execute the lamp mode switching to complete the multi-modal contextual social behavior expression, so as to realize the perception of the situation according to the user's behavior and output the social behavior of the humanoid robot.
[0166] Figure 2 When the sensor perceives the user's behavior, the specific steps of the perceived content include;
[0167] S101. Perceive the user's location through the sensor.
[0168] S1011. Perceive that the user is in the interactive area in the physical environment through the pyroelectric infrared sensor.
[0169] S1012. Perceive the distance between the user and the humanoid robot through the ultrasonic sensor, and set three thresholds as the recognition conditions for different situations to occur.
[0170] S102. Collect the user's voice signal through the microphone and perform filtering processing.
[0171] The audio collected in step S102 is used for the adaptive hierarchical isolated word recognition in step S105 of the user voice command and the user voice emotion prediction in step S107.
[0172] Figure 3 For the multi-layer interaction context perception structure, the main steps in the embodiment of the present invention include,
[0173] Step S1011 is the initial state perception layer of the interaction context, steps S1012 - S1013 are the first-layer structure layer of the interaction context perception, and steps S102 - S105 are the second-layer structure layer of the interaction context perception. Among them, when no user audio is collected in step S1021, the interaction context returns to the welcome context in the first-layer structure.
[0174] Figure 4 The main details of the system communication scheme are described as follows:
[0175] In the embodiment of the present invention, the ultrasonic sensor, pyroelectric infrared sensor, and microphone are used to collect signals, and real-time one-way wired communication of data is carried out using the I / O port and the USB data acquisition card.
[0176] In the embodiment of the present invention, the USB data acquisition card communicates with the PC, and the PC communicates with the humanoid robot using the TCP / IP protocol. It is necessary to place the PC and the humanoid robot in the same local area network for two-way wireless communication. The TCP / IP protocol requires a handshake confirmation to establish communication. The PC sends the contextualized social behavior encoding described in S302 to the humanoid robot.
[0177] In the embodiment of the present invention, the humanoid robot mainly consists of a humanoid robot controller, a robot speaker, a robot LED light, and a robot joint servo.
[0178] Figure 5 It is the social action decision-making process of the humanoid robot, and its implementation details are mainly the steps of S203 to S204.
[0179] Figure 6 It is the multi-modal expression flowchart of the contextualized social behavior of the humanoid robot; its implementation details are mainly the steps of S301 to S304.
[0180] The protection scope of the present invention includes but is not limited to the above embodiments. The protection scope of the present invention is subject to the claims. Any replacement, deformation, and improvement that are easily conceivable by those skilled in the art to this technology fall within the protection scope of the present invention.
Claims
1. A method for expressing situational social behavior of a humanoid robot, characterized in that: It includes the steps of user behavior and emotion perception, the humanoid robot's contextualized social behavior decision-making, and the humanoid robot's multimodal social behavior expression; The steps of user behavior and emotion perception include: S101, sensing the user's position through sensors and defining interaction context categories; S1011, sensing the interactive area of the user in the physical environment through sensors; S1012, using sensors to sense the distance between the user and the humanoid robot, and setting three thresholds as identification conditions for different situations, where the three thresholds are threshold 1, threshold 2, and threshold 3 from largest to smallest; S1013, define five interaction situation categories, namely, attention, welcome, instruction, apology and blocking situations; when the distance obtained in step S1012 is between threshold 1 and threshold 2, the situation is an attention situation; when the distance is between threshold 2 and threshold 3, the situation is a welcome situation; when the distance is less than threshold 3, the situation is an instruction, apology or blocking situation; S102, collecting user voice signals and performing filtering processing; S103, extracting speech feature parameters and constructing a speech feature vector; S104, recognizing user voice command vocabulary; S1041, pre-define speech recognition sample sets, and divide them into three types of sample sets: instruction, apology, and obstruction according to the situation, and each sample set contains multiple words that will appear in the situation; S1042, training the speech feature parameter template of the sample set according to step S103; S105, performing time regularization on the user voice feature vector according to the reference template of S1042 and the user voice command, obtaining an optimal regularization path, and obtaining a context category to which the user voice command belongs; S106, constructing speech emotion features and constructing speech emotion feature templates; S107, using the data in the speech emotion feature template to train the neural network to obtain a trained neural network model; S108: After the user audio passes through steps S102, S103 and S106, the speech emotion feature corresponding to each frame of speech is obtained, and the speech emotion feature corresponding to each frame of speech is input into the neural network model trained in S107 to obtain a prediction of the emotion category of the user audio; Among them, the contextual social behavior decision-making steps of the humanoid robot include: S201, generating an interaction context category code according to the output of step S1013 and the output of step S105; S202, pre-define a contextualized social action dataset of a humanoid robot, the dataset including contextual social task actions of five context categories, namely, attention, welcome, instruction, apology, and blocking contexts; S203, defining social emotion action subcategories in each situational social task category; S204, taking the emotion category predicted by step S108 as a condition, according to the output of S201, searching for the social emotion action subcategory of the humanoid robot that matches the emotion category predicted by step S108 under the interaction context category; and outputting the social emotion action code of the humanoid robot through the search result; S205, pre-define voice templates, define voice tags and light mode tags in the welcome, instruction, apology and blocking situations; according to the output of step S105, output and encode the voice tags in the situation category to which the user's voice command belongs, output the light mode tags in the situation category to which the user's voice command belongs, and perform light mode encoding; S206, integrating the codes of steps S204 and S205 and defining them as contextualized social behavior codes of humanoid robots; Among them, the steps of expressing the multimodal social behavior of the humanoid robot include: S301, pre-defining a multimodal social behavior template of a humanoid robot according to the output of step S206; S302, using the TCP / IP protocol, the preset data packet fusion order is: action data, light mode data, voice data, and the fused humanoid robot contextual social behavior data packet is sent to the humanoid robot; S303, the humanoid robot receives the data packet, and decodes and extracts the interactive content; S304, calling the predefined action library and voice library in the humanoid robot through the API, and executing the light mode switching to complete the multi-modal situational social behavior expression.
2. The method for expressing situational social behavior of a humanoid robot according to claim 1, characterized in that: The sensor used in S1011 is an infrared pyroelectric sensor, and the sensor used in S1012 is an ultrasonic sensor; The S102 is specifically: S1021. Collecting user voice signals through a microphone; S1022, amplifying and filtering the collected signal; S1023, transmitting the signal after step S1022 to the PC via a USB data acquisition card; S1024: Filter the audio signal using a filtering algorithm to further reduce signal noise.
3. The method for expressing situational social behavior of a humanoid robot according to claim 1, characterized in that: The S103 is specifically: S1031, pre-emphasize, frame and window the voice, and perform voice endpoint detection; S1032, performing fast Fourier transform on each frame signal and calculating the spectral line energy; S1033. Using an auditory filter bank, calculate the speech energy spectrum and discrete cosine transform cepstrum of each frame, extract cepstrum features and their first-order difference parameters, and construct a speech feature vector.
4. The method for expressing situational social behavior of a humanoid robot according to claim 1, characterized in that: The S105 is specifically: S1051, record the speech feature vector in the reference template of S1042 as the reference speech feature vector, and record the speech feature vector corresponding to the user speech instruction obtained after step S1033 as the user speech feature vector; suppose the reference speech feature vector has J frame vectors, and the user speech feature vector has I frame vectors; The user's voice feature vector is time-warped, and the formula is: Where d[T(i), R(ω(i))] is the distance between the user speech feature vector T(i) of the i-th frame and the reference speech feature vector R(j) of the j-th frame, j = ω(i), and D is the vector distance under the time warping optimization condition; When S1052 and S1051 perform time regularization, for the user voice feature vector T(i) of the i-th frame and the reference voice feature vector R(j) of the j-th frame, a reverse decision is used to obtain the optimal regularization path, and then the context category to which the user voice command belongs is obtained.
5. The method for expressing situational social behavior of a humanoid robot according to claim 4, characterized in that: The S1052 is specifically: S10521, the user voice feature vector corresponding to the user voice command extracted in step S1033 is used as a test sample, and the reference voice feature vector in the reference template of S1042 is used as a training sample. When the test sample and the training sample are matched to calculate the distance, several frames are randomly selected and the cumulative distance is calculated by formula (1), and the cumulative distance is set to D c ; S10522, according to S10521, the cumulative distance is D c Determine adaptive threshold conditions and perform adaptive stratification; S10523, when the accumulated distance D c If the cumulative distance is less than the threshold, the training sample is stored in the new template library, and the cumulative distance and the calculated frame length are stored; if the cumulative distance is greater than or equal to the threshold, the training sample is ignored; S10524. Define the new template library obtained in S10523 as the first-level matching; call the cumulative distance and frame length obtained in S10523, calculate the cumulative distance of the remaining frame length, and select the optimal regularized path based on the minimum cumulative distance to obtain the matching result of the test sample, that is, the context category to which the user's voice command belongs.
6. The method for expressing situational social behavior of a humanoid robot according to claim 1, characterized in that: The S106 is specifically: S1061, extracting short-time energy of audio; S1062, extracting the fundamental frequency and first-order fundamental frequency jitter of the audio using an autocorrelation function method; S1063, extracting the first formant parameter of the audio using a linear prediction method; S1064, constructing speech emotion features by combining the cepstrum extracted in S1033 and its first-order difference parameters and steps S1061 to S1063; S1065, using linear discriminant analysis to reduce the dimension of speech emotion features; S1066. Using the CASIA emotional speech database, generate a speech emotion feature template through steps S1061 to S1065.
7. The method for expressing situational social behavior of a humanoid robot according to claim 1, characterized in that: The S107 is specifically: S1071. Construct a back propagation neural network topology; S1072, training and normalizing the speech emotion feature sample data of step S1066; S1073, initializing the network structure, using the normalized speech emotion feature sample data obtained in step S1072 as the input of the neural network; S1073, encoding the speech of the CASIA emotional speech database into emotion categories, and using the encoded samples as the output of the neural network, wherein the corresponding emotion categories of the output codes include happiness, neutrality, surprise, anger, fear and sadness; S1074, using logistic regression logsig as the neural network activation function; S1075, initializing the learning rate, smoothing parameter, cumulative sum of squared gradients, and mean square error vector, and setting the maximum number of training iterations; S1076, calculate forward propagation; S1077, calculating propagation error and back propagation gradient; S1078. Optimize the adaptive learning rate to update the optimal solution of the neural network weights and biases. When the iteration reaches the maximum number of iterations set in step S1075 or reaches the target error, an optimized neural network model is obtained, that is, a trained neural network model.
8. The method for expressing situational social behavior of a humanoid robot according to claim 1, characterized in that: The S201 is specifically: S2011, generating the first level code of the interaction situation category outputted in step S1013; S2012, generating the second level code of the interaction situation category outputted in step S105; S2013, integrating the first layer and the second layer coding to obtain the interaction context category coding; The S203 is specifically: Combined with the attribute characteristics of the humanoid robot situational social tasks, social emotion action subcategories are defined in each situational social task category, and the social emotion action subcategories include neutral, happy, sad, angry and / or fearful.
9. The method for expressing situational social behavior of a humanoid robot according to claim 1, characterized in that: The S204 is specifically: S2041, taking the emotion category predicted by step S108 as a condition, according to the output of S201, searching for a social emotion action subcategory of a humanoid robot that matches the emotion category predicted by step S108 under the interaction context category; S2042: if found, output the humanoid robot social emotion action code in the social emotion action subcategory; S2043. If not found, output the social emotion action code of the humanoid robot in the neutral subcategory.
10. The method for expressing situational social behavior of a humanoid robot according to claim 1, characterized in that: The S205 is specifically: S2051, pre-defined voice templates, defining voice labels and LED light mode labels in the welcome, instruction, apology and blocking situations; S2052, outputting the voice tag in the welcome context according to step S103 and encoding it; S2053, according to the output of step S105, output the voice tag under the scenario category to which the user's voice command belongs and encode it; S2054. According to the output of step S105, output the LED light mode label under the scenario category to which the user voice command belongs, and perform LED light mode encoding.
Citation Information
Cited By
Bionic robot complex emotion expression method and system based on multi-Agent game
CN121936501A
Bionic robot complex emotion expression method and system based on multi-agent game
CN121936501B