An intelligent interactive robot control system
By designing an intelligent interactive robot control system, using voice and audio analysis technology to identify user emotional states and automatically adjust the output style, the problem of failure to determine user emotional states based on voice input in the existing technology is solved, and more intelligent and personalized human-computer interaction is achieved.
Patent Information
- Application Number
- CN202510301437.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The prior art fails to determine the user's current emotional state based on voice input and automatically adjust the output style of the robot.
An intelligent interactive robot control system is designed, including a voice receiving module, a text conversion processing module, an audio analysis processing module and an emotion judgment adjustment module. The system determines the user's current emotional state through speech noise reduction, text conversion, audio analysis and emotional judgment, and switches the voice output style according to the emotional state.
It significantly improves the intelligence level of human-computer interaction, can accurately identify the user's emotional state and automatically adjust the output style, provide a more considerate and personalized interactive experience, effectively alleviate user negative emotions and promote the generation of positive emotions.
Smart Images

Figure CN119811362B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly to an intelligent interactive robot control system. Background Art
[0002] With the rapid development of technologies such as artificial intelligence, natural language processing, machine learning, speech recognition and synthesis, interactive robots have been able to have more natural and fluent conversations with humans. The progress of these technologies makes it possible for robots to understand and respond to human emotions, and further enables robots to mobilize the emotions of users during the interaction.
[0003] Chinese Patent Publication No. CN116665273B discloses a robot human-computer interaction method based on facial expression recognition and emotion quantification analysis calculation, including: Step P100: Recognize the user's facial expression; Step P200: Perform quantification and calculation analysis of emotions based on the facial expression recognition result of Step P100 to obtain the current emotion quantification calculation analysis result; Step P300: Transmit the current emotion quantification calculation analysis result to the interactive robot, and the interactive robot determines the corresponding interaction content according to the current emotion quantification calculation analysis result, and outputs the interaction content in a human-computer interaction manner to adjust the user's emotions accordingly. The robot human-computer interaction method of the present invention can recognize various emotions that the user may have simultaneously, so as to more accurately recognize the user's current emotion state. At the same time, the interactive robot makes targeted responses according to the user's current emotion state, and better plays the role of adjusting the user's emotions.
[0004] Therefore, the following problems exist in this invention:
[0005] The prior art fails to determine the user's current emotional state based on voice input and automatically adjust the output style of the robot according to the current emotional state. Summary of the Invention
[0006] For this reason, the present invention provides an intelligent interactive robot control system to overcome the problem in the prior art that the user's emotional state cannot be determined based on voice input and the output style of the robot cannot be switched.
[0007] To achieve the above object, the present invention provides an intelligent interactive robot control system, including:
[0008] A voice receiving module for receiving the original audio signal emitted by the user and performing noise reduction processing on the original audio signal to form a purified audio signal;
[0009] A text conversion processing module, which is connected to the voice receiving module, is used to convert the purified audio signal into text content to determine the text word length, determine the number of adjectives in a single text content, the emotional vocabulary categories of each adjective, the number of nouns, and the emotional vocabulary categories of each noun according to the text content, determine the corresponding adjective ratio according to the number of adjectives and the number of nouns, and determine the negative word ratio according to the judgment results of the emotional vocabulary categories of each adjective and the emotional vocabulary categories of each noun;
[0010] An audio analysis processing module, which is respectively connected to the voice receiving module and the text conversion processing module, is used to determine the speech rate of a single purified audio signal according to the time length of the single purified audio signal and its corresponding text word length, and determine the time-domain audio image according to the purified audio signal to determine the pitch feature trend and frequency fluctuation tendency;
[0011] An emotion judgment and adjustment module, which is respectively connected to the text conversion processing module and the audio analysis processing module, is used to determine the text emotion tendency according to the text word length, the adjective ratio, and the negative word ratio of the text content corresponding to a single purified audio signal, determine the audio emotion tendency according to the pitch feature trend and frequency fluctuation tendency of the single purified audio signal, and combine the text emotion tendency and the audio emotion tendency to determine the current emotional state of the user to determine whether to switch the speech output style of the interactive robot;
[0012] Among them, the emotional vocabulary categories include positive emotional vocabulary, negative emotional vocabulary, and neutral emotional vocabulary.
[0013] Further, the text conversion processing module includes a positive emotion word library and a negative emotion word library to determine the emotional vocabulary categories of adjectives and nouns, and determine the negative word ratio according to the judgment results of the emotional vocabulary categories of each adjective and the emotional vocabulary categories of each noun;
[0014] Among them, the negative word ratio is the ratio of the sum of the numbers of negative emotional vocabulary of adjectives and nouns to the sum of the numbers of adjectives and nouns.
[0015] Further, the audio analysis processing module determines the speech rate of a single purified audio signal according to the time length of the single purified audio signal and its corresponding text word length;
[0016] Among them, the speech rate of a single purified audio signal is the ratio of the time length of the single purified audio signal to the corresponding text word length.
[0017] Further, the audio analysis processing module determines the pitch frequency and fundamental frequency slope of a single time-domain audio image through a deep learning model, and determines the pitch feature trend according to whether the pitch frequency and the fundamental frequency slope meet the pitch feature conditions, including,
[0018] If the fundamental frequency and the fundamental frequency slope satisfy the fundamental tone feature condition, it is determined that the fundamental tone feature trend is a negative fundamental tone trend;
[0019] If the fundamental frequency and the fundamental frequency slope do not satisfy the fundamental tone feature condition, it is determined that the fundamental tone feature trend is a positive fundamental tone trend;
[0020] Wherein, the fundamental tone feature condition is that the fundamental frequency is less than the frequency threshold and the fundamental frequency slope is less than the slope threshold.
[0021] Furthermore, the audio analysis and processing module determines the frequency difference between each pair of extreme values, and determines the characteristic extreme value pairs according to the comparison result between the frequency difference and the preset frequency difference value, including,
[0022] If the frequency difference is greater than or equal to the preset frequency difference value, it is determined that the corresponding pair of extreme values is a characteristic extreme value pair;
[0023] If the frequency difference is less than the preset frequency difference value, it is determined that the corresponding pair of extreme values is not a characteristic extreme value pair
[0024] Wherein, the pair of extreme values consists of an adjacent maximum value and a minimum value;
[0025] The preset frequency difference value is determined according to the maximum audio difference value, and the maximum audio difference value is the difference between the maximum frequency value and the minimum frequency value in the time-domain audio image.
[0026] Furthermore, the audio analysis and processing module determines the frequency fluctuation tendency according to the ratio of the number of characteristic extreme value pairs to the number of all extreme value pairs, including,
[0027] If the ratio of the number of characteristic extreme value pairs to the number of all extreme value pairs is less than the preset ratio, it is determined that the frequency fluctuation tendency is a negative frequency fluctuation;
[0028] If the ratio of the number of characteristic extreme value pairs to the number of all extreme value pairs is greater than or equal to the preset ratio, it is determined that the frequency fluctuation tendency is a positive frequency fluctuation;
[0029] Wherein, the preset ratio is determined according to the speech rate of the corresponding purified audio signal.
[0030] Furthermore, the emotion judgment and adjustment module determines the text emotion tendency according to the text word length, the adjective ratio, and the negative word ratio of the text content corresponding to a single purified audio signal, including,
[0031] When the text word length is less than the preset length, the text emotion tendency is determined according to the adjective ratio;
[0032] Wherein, if the adjective ratio is less than the preset adjective ratio, it is determined that the text emotion tendency is negative emotion text.
[0033] Furthermore, the emotion judgment and adjustment module determines the text emotional tendency based on the text word count length, the adjective ratio, and the negative word ratio of the text content corresponding to a single purified audio signal. It also includes that
[0034] when the text word count length is greater than or equal to the preset length, the text emotional tendency is determined according to the adjective ratio and the negative word ratio;
[0035] wherein, if the adjective ratio is greater than or equal to the preset adjective ratio and the negative word ratio is greater than or equal to the preset negative word ratio, it is determined that the text emotional tendency is a negative emotional tendency.
[0036] Furthermore, the emotion judgment and adjustment module determines the audio emotional tendency according to the fundamental tone feature trend and the frequency fluctuation tendency of a single purified audio signal. Among them,
[0037] if the fundamental tone feature trend is a negative fundamental tone trend and the negative frequency fluctuation is a negative frequency fluctuation, it is determined that the audio emotional tendency is a negative emotional audio.
[0038] Furthermore, the emotion judgment and adjustment module combines the text emotional tendency and the audio emotional tendency to determine the user's current emotional state to determine whether to switch the voice output style of the interactive robot, including,
[0039] if the text emotional tendency is negative emotional text and / or the audio emotional tendency is negative emotional audio, it is determined that the current emotional state is a negative emotion and the voice output style of the interactive robot is switched;
[0040] if the text emotional tendency is positive emotional text and the audio emotional tendency is positive emotional audio, it is determined that the current emotional state is a positive emotion and the voice output style of the interactive robot is not switched.
[0041] Compared with the prior art, the beneficial effects of the present invention are that the intelligent interactive robot control system provided by the present invention significantly improves the intelligent level of human-computer interaction by integrating efficient voice processing and emotion analysis functions. It can not only effectively reduce noise and accurately convert speech into text, deeply analyze the emotional elements in the text, but also capture the subtle signs of the user's emotional changes through audio analysis.
[0042] Furthermore, this system can intelligently identify the user's emotional state, and automatically switch to the user-preferred cartoon-style voice output when detecting negative emotions. Through the affinity and fun of the cartoon character, it effectively relieves the user's sad emotions and stimulates their positive emotions; by changing the output voice of the robot, this system not only provides a more considerate and personalized interaction experience, but also plays an invisible role in emotion regulation, promoting the positive development of the user's mental health.
[0043] Furthermore, the emotion judgment and adjustment module realizes the intelligent recognition of the audio emotion tendency by accurately analyzing the pitch feature trend and frequency fluctuation tendency of the audio signal, effectively improving the accuracy of user emotion perception in human-computer interaction. At the same time, by combining the comprehensive judgment of the text emotion tendency and the audio emotion tendency, it can capture the current emotion state of the user more comprehensively and meticulously, and then intelligently adjust the speech output style of the interactive robot to adapt to and respond to the emotional needs of the user. This mechanism not only enhances the natural fluency of human-computer interaction, but also promotes emotional resonance and personalized communication, significantly improving the satisfaction and comfort of the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a connection diagram of the intelligent interactive robot control system according to an embodiment of the present invention;
[0045] Figure 2 It is a working flowchart of the text conversion processing module according to an embodiment of the present invention;
[0046] Figure 3 It is a working flowchart of the audio analysis and processing module according to an embodiment of the present invention;
[0047] Figure 4 It is a working flowchart of the emotion judgment and adjustment module according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to make the objectives and advantages of the present invention more clear, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0049] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0050] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0051] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and defined, the terms "installation", "connection", and "linkage" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0052] Please refer to Figures 1 to 4 as shown, which are respectively the connection diagram of the intelligent interaction robot control system according to the embodiment of the present invention, the working flowchart of the text conversion processing module according to the embodiment of the present invention, the working flowchart of the audio analysis processing module according to the embodiment of the present invention, and the working flowchart of the emotion judgment and adjustment module according to the embodiment of the present invention. The embodiment of the present invention provides an intelligent interaction robot control system, including:
[0053] A voice receiving module, configured to receive the original audio signal emitted by the user and perform noise reduction processing on the original audio signal to form a purified audio signal;
[0054] A text conversion processing module, which is connected to the voice receiving module, configured to convert the purified audio signal into text content to determine the text word length, determine the number of adjectives in a single text content, the emotional vocabulary categories of each adjective, the number of nouns, and the emotional vocabulary categories of each noun according to the text content, determine the corresponding adjective ratio according to the number of adjectives and the number of nouns, and determine the negative word ratio according to the determination results of the emotional vocabulary categories of each adjective and the emotional vocabulary categories of each noun;
[0055] An audio analysis processing module, which is respectively connected to the voice receiving module and the text conversion processing module, configured to determine the speech rate of a single purified audio signal according to the time length of the single purified audio signal and its corresponding text word length, and determine the time-domain audio image according to the purified audio signal to determine the pitch feature trend and the frequency fluctuation tendency;
[0056] An emotion judgment and adjustment module, which is respectively connected to the text conversion processing module and the audio analysis processing module, configured to determine the text emotion tendency according to the text word length, the adjective ratio, and the negative word ratio of the text content corresponding to a single purified audio signal, determine the audio emotion tendency according to the pitch feature trend and the frequency fluctuation tendency of a single purified audio signal, and determine the current emotional state of the user by combining the text emotion tendency and the audio emotion tendency to determine whether to switch the speech output style of the interaction robot;
[0057] Wherein, the emotional vocabulary categories include positive emotional vocabulary, negative emotional vocabulary, and neutral emotional vocabulary.
[0058] In implementation, the voice output styles include a normal style and a cartoon style; the cartoon style includes several cartoon voice packs, and each cartoon voice pack is the cartoon voice of the protagonist of the most-played animated cartoon within each time period. The corresponding cartoon voice pack is determined according to the age of the user. In one implementation, the user is a person born after 1995 and is 28 years old. The steps to determine his cartoon voice pack include: (1) determining the animated cartoon year according to the current year, the current age of the user and a preset age. The preset age is generally set to be about 10 years old, and preferably 10 years old; (2) determining the cartoon voice of the protagonist of the most-played animated cartoon in that animated cartoon year.
[0059] It can be understood that the default voice output style of the intelligent interaction robot at each startup is the normal style, and its voice output style is automatically switched when it is determined that the current emotional state of the user is a negative emotion.
[0060] It can be understood that the intelligent interaction robot control system provided by the present invention significantly improves the intelligent level of human-computer interaction by integrating efficient voice processing and emotion analysis functions. It can not only effectively reduce noise and accurately convert voice into text, deeply analyze the emotional elements in the text, but also capture subtle signs of changes in the user's emotional state through audio analysis; it can also intelligently identify the emotional state of the user, and automatically switch to the cartoon-style voice output preferred by the user when detecting a negative emotion, effectively relieving the user's sad emotion through the affinity and interestingness of the cartoon character, and stimulating his positive emotion; by changing the output voice of the robot, this system not only provides a more considerate and personalized interaction experience, but also plays a role in emotion regulation invisibly, promoting the positive development of the user's mental health.
[0061] Specifically, the text conversion processing module includes a positive emotion word library and a negative emotion word library to determine the emotional word categories of adjectives and nouns, and determines the negative word ratio according to the judgment results of the emotional word categories of each adjective and the emotional word categories of each noun;
[0062] Among them, the negative word ratio is the ratio of the sum of the numbers of negative emotional words of adjectives and nouns to the sum of the numbers of adjectives and nouns.
[0063] It can be understood that the positive emotion word library and the negative emotion word library will be updated online regularly. The update frequency is usually once a week to three times a week, and usually once a week.
[0064] It can be understood that by introducing a positive sentiment word library and a negative sentiment word library and setting up a mechanism for regular online updates, the present invention ensures the timeliness and accuracy of the content in the word libraries, enables the text conversion and processing module to more accurately identify and analyze the sentiment tendency in the text, and enables the system to keep up with the trend of the times and capture the latest sentiment expression methods and vocabulary changes.
[0065] Specifically, the audio analysis and processing module determines the speech rate of a single purified audio signal according to the time length of the single purified audio signal and the length of the corresponding text in terms of the number of words.
[0066] Wherein, the speech rate of a single purified audio signal is the ratio of the time length of the single purified audio signal to the length of the corresponding text in terms of the number of words.
[0067] It can be understood that when a person is in a happy or unhappy mood, the speech rate often shows obvious differences. When a person is in a positive mood, the speech rate usually speeds up because a happy mood is often accompanied by excitement and agitation, making people tend to express their inner joy and pleasure in a faster speech rate. At this time, the intonation also becomes relatively high-pitched, and the voice is full of vitality and enthusiasm. In contrast, when a person is in a negative mood, the speech rate will relatively slow down because negative emotions are often accompanied by emotions such as frustration and disappointment, which will make people feel heavy and depressed, thus slowing down the speech rate.
[0068] Specifically, the audio analysis and processing module determines the fundamental frequency and the fundamental frequency slope of a single time-domain audio image through a deep learning model, and determines the fundamental frequency feature trend according to whether the fundamental frequency and the fundamental frequency slope meet the fundamental frequency feature conditions, including
[0069] If the fundamental frequency and the fundamental frequency slope meet the fundamental frequency feature conditions, it is determined that the fundamental frequency feature trend is a negative fundamental frequency trend;
[0070] If the fundamental frequency and the fundamental frequency slope do not meet the fundamental frequency feature conditions, it is determined that the fundamental frequency feature trend is a positive fundamental frequency trend;
[0071] Wherein, the fundamental frequency feature condition is that the fundamental frequency is less than the frequency threshold and the fundamental frequency slope is less than the slope threshold.
[0072] It can be understood that when a person is in a positive mood such as happy or glad, the vocal cords may be more relaxed, and the relatively high vibration frequency causes the fundamental frequency to rise. This rising fundamental frequency is often accompanied by an increase in speech rate and an upward intonation, conveying positive and pleasant emotions. Moreover, when in a positive mood, people's emotions are relatively stable, so the change in the fundamental frequency may also be relatively stable without much fluctuation. When a person is in a negative mood of sadness, the vocal cords tend to be more tense, resulting in a decrease in vibration frequency and a drop in the fundamental frequency. In addition, negative emotions may also cause an increase in the fluctuation of the fundamental frequency, reflecting emotional instability.
[0073] It can be understood that the fundamental frequency slope is a parameter used to describe the rate of change of the fundamental frequency. The fundamental frequency slope (i.e., the slope of the pitch regression line) is an important parameter for measuring the change trend of the pitch curve. In different emotional states, the change of the fundamental frequency slope is also different: when a person is in a happy or glad mood, the speech is usually more vivid and expressive. This emotion prompts the speaker to adjust the intonation more frequently, resulting in an increase in the fluctuation range of the fundamental frequency, and thus an increase in the fundamental frequency slope. Compared with positive emotions, the fundamental frequency slope shows different characteristics in negative emotions. When sad, the fundamental frequency often drops, and may be accompanied by unstable phenomena such as tremors, resulting in a complex change trend of the fundamental frequency slope.
[0074] It can be understood that the application of the deep learning model in the audio analysis and processing module can accurately capture and analyze the fundamental frequency and the fundamental frequency slope in a single time-domain audio image, and then determine the trend of the fundamental frequency characteristics based on these acoustic features, realizing in-depth insight into the user's emotional state. This mechanism cleverly links the changes in the fundamental frequency and the fundamental frequency slope with the user's emotional state. By setting the fundamental frequency characteristic conditions, it effectively distinguishes between positive and negative fundamental frequency trends. This emotion recognition method based on acoustic features not only enriches the basis for emotion judgment but also improves the accuracy and sensitivity of emotion recognition. This meticulous emotion analysis ability helps the intelligent interaction robot to provide more considerate and personalized feedback and adjustment during the human-computer interaction process, effectively alleviating the user's negative emotions and promoting the generation of positive emotions, thereby improving the overall quality of the user experience.
[0075] Specifically, the audio analysis and processing module determines the frequency difference between each pair of extreme values, and determines the characteristic extreme value pairs according to the comparison result between the frequency difference and the preset frequency difference value, including,
[0076] If the frequency difference of a single pair of extreme values is greater than or equal to the preset frequency difference value, then determine the corresponding pair of extreme values as the characteristic extreme value pair;
[0077] If the frequency difference of a single pair of extreme values is less than the preset frequency difference value, then determine that the corresponding pair of extreme values is not a characteristic extreme value pair;
[0078] Among them, the extreme value pair consists of an adjacent maximum value and a minimum value;
[0079] The preset frequency difference is determined according to the maximum audio difference, and the maximum audio difference is the difference between the maximum frequency and the minimum frequency in the time-domain audio image; in implementation, the value range of the preset frequency difference is 0.5 times the maximum audio difference to 0.9 times the maximum audio difference. The larger the preset frequency difference, the greater the frequency difference of the characteristic extreme value pair.
[0080] It can be understood that a single extreme value should be able to form two different extreme value pairs. In implementation, a maximum value forms two different extreme value pairs with two adjacent minimum values respectively. Similarly, a single minimum value forms two different extreme value pairs with two adjacent maximum values respectively.
[0081] Specifically, the audio analysis and processing module determines the frequency fluctuation tendency according to the ratio of the characteristic extreme value pair to the total number of extreme value pairs, including,
[0082] If the ratio of the characteristic extreme value pair to the total number of extreme value pairs is less than the preset ratio, it is determined that the frequency fluctuation tendency is a negative frequency fluctuation;
[0083] If the ratio of the characteristic extreme value pair to the total number of extreme value pairs is greater than or equal to the preset ratio, it is determined that the frequency fluctuation tendency is a positive frequency fluctuation;
[0084] Among them, the preset ratio is determined according to the speech rate of the corresponding purified audio signal.
[0085] In implementation, the preset ratio has a negative correlation with the speech rate. It can be understood that the larger the preset ratio, the looser the condition for determining the frequency fluctuation tendency as a negative frequency fluctuation; the change of the speech rate is an important way of emotional expression, and the emotional state conveyed by the speaker can be roughly judged accordingly. The slower the speech rate, the more likely the emotional state of the speaker is sad. Therefore, the value of the preset ratio can be increased to relax the condition for determining the negative frequency fluctuation through the ratio; when the speech rate is normal (greater than or equal to the historical average speech rate of positive emotions), the frequency fluctuation tendency of the speaker is determined only through the ratio. Therefore, a relatively strict negative frequency fluctuation condition (i.e., a relatively low preset ratio) is required to determine the frequency fluctuation tendency.
[0086] In implementation, the preset ratio ∈[0.2, 0.5], and preferably set to 0.25, that is, when the speech rate is normal, the preset ratio takes 0.25; when the speech rate is less than the historical average speech rate of positive emotions, the preset ratio = current speech rate ÷ historical average speech rate of positive emotions × 0.25.
[0087] It can be understood that the larger the ratio of the number of characteristic extreme value pairs to the number of all extreme value pairs, the more frequent the frequency fluctuation of a single purified audio data, which may indicate that the speaker is currently in a positive mood; the voice in a negative mood may be lower and slower, with relatively fewer frequency fluctuations.
[0088] Specifically, the emotion judgment and adjustment module determines the text emotion tendency according to the text word length, the adjective ratio, and the negative word ratio of the text content corresponding to a single purified audio signal, including,
[0089] When the text word length is less than the preset length, determine the text emotion tendency according to the adjective ratio;
[0090] Among them, if the adjective ratio is less than the preset adjective ratio, it is determined that the text emotion tendency is negative emotion text.
[0091] It can be understood that the shorter the text word length, the fewer the number of adjectives and nouns will be. In normal expressions, usually at least one adjective is used to modify a noun; if there are fewer words and fewer adjectives used in a sentence, the emotion may be relatively low.
[0092] In implementation, the preset adjective ratio ∈ [1, 2]. Preferably, the preset adjective ratio is 1. The smaller the value of the preset adjective ratio, the looser the condition for judging the text emotion tendency.
[0093] In implementation, the preset length is the historical average word count of the positive emotion of a single text content.
[0094] It can be understood that a single original audio data, its corresponding purified audio data, and the corresponding text content are the single voice input command of the user to the interactive robot.
[0095] Specifically, the emotion judgment and adjustment module determines the text emotion tendency according to the text word length, the adjective ratio, and the negative word ratio of the text content corresponding to a single purified audio signal, and further includes,
[0096] When the text word length is greater than or equal to the preset length, determine the text emotion tendency according to the adjective ratio and the negative word ratio;
[0097] Among them, if the adjective ratio is greater than or equal to the preset adjective ratio and the negative word ratio is greater than or equal to the preset negative word ratio, it is determined that the text emotion tendency is negative emotion tendency.
[0098] It can be understood that the larger the adjective ratio, the longer the corresponding text length should be. At this time, the larger the ratio of the number of negative adjectives to the number of all adjectives, the more obvious it can be determined that the text emotion tendency is negative emotion tendency.
[0099] In implementation, the preset negative word ratio ∈ [0.4, 0.8], and preferably the preset negative word ratio is 0.5, that is, when the number of negative adjectives is half of the total number of adjectives, it is determined that the text emotional tendency is a negative emotional tendency.
[0100] It can be understood that the emotion judgment adjustment module comprehensively considers three major elements: the text word length, the adjective ratio, and the negative word ratio to determine the text emotional tendency, showing higher accuracy and comprehensiveness. When the text reaches a certain length, the module not only relies on the richness of adjectives but also deeply analyzes the proportion of negative words in the overall adjectives. This dual standard effectively avoids misjudgments that may be caused by a single factor; when the adjective ratio reaches the preset threshold and the negative word ratio is also significant, the system determines it as a negative emotional tendency. This mechanism greatly enhances the robot's ability to capture the user's emotions; in addition, the reasonable setting of the preset negative word ratio not only ensures the rigor of the judgment but also takes into account the diversity of emotional expressions in different contexts, enabling the robot to more accurately understand the user's emotional state and thus provide a more considerate and personalized interaction experience; in summary, it not only improves the accuracy of emotion judgment but also builds a more solid bridge for emotional communication between the user and the robot.
[0101] Specifically, the emotion judgment adjustment module determines the audio emotional tendency according to the pitch feature trend and frequency fluctuation tendency of a single purified audio signal, where,
[0102] If the pitch feature trend is a negative pitch trend and the negative frequency fluctuation is a negative frequency fluctuation, it is determined that the audio emotional tendency is a negative emotional audio.
[0103] Specifically, the emotion judgment adjustment module combines the text emotional tendency and the audio emotional tendency to determine the user's current emotional state to determine whether to switch the voice output style of the interactive robot, including,
[0104] If the text emotional tendency is negative emotional text and / or the audio emotional tendency is negative emotional audio, it is determined that the current emotional state is a negative emotion and the voice output style of the interactive robot is switched;
[0105] If the text emotional tendency is positive emotional text and the audio emotional tendency is positive emotional audio, it is determined that the current emotional state is a positive emotion and the voice output style of the interactive robot is not switched.
[0106] It can be understood that the emotional judgment and regulation module realizes the intelligent recognition of the emotional tendency of audio by accurately analyzing the pitch feature trend and frequency fluctuation tendency of the audio signal, effectively improving the accuracy of user emotion perception in human-computer interaction. At the same time, by combining the comprehensive judgment of the text emotional tendency and the audio emotional tendency, it can capture the current emotional state of the user more comprehensively and meticulously, and then intelligently adjust the speech output style of the interactive robot to adapt to and respond to the emotional needs of the user. This mechanism not only enhances the natural fluency of human-computer interaction, but also promotes emotional resonance and personalized communication, significantly improving the satisfaction and comfort of the user experience.
[0107] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
[0108] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An intelligent interactive robot control system, characterized in that: include: A voice receiving module, used for receiving an original audio signal sent by a user and performing noise reduction processing on the original audio signal to form a purified audio signal; A text conversion processing module, which is connected to the voice receiving module, and is used to convert the purified audio signal into text content to determine the length of the text word count, determine the number of adjectives, the emotional vocabulary category of each adjective, the number of nouns and the emotional vocabulary category of each noun in a single text content according to the text content, determine the corresponding adjective ratio according to the number of adjectives and the number of nouns, and determine the negative word ratio according to the determination results of the emotional vocabulary category of each adjective and the emotional vocabulary category of each noun; an audio analysis processing module, which is connected to the speech receiving module and the text conversion processing module respectively, and is used to determine the speech rate of a single purified audio signal according to the time length of the single purified audio signal and the length of the corresponding text word count, and to determine the time domain audio image according to the purified audio signal to determine the fundamental tone feature trend and the frequency fluctuation tendency; an emotion judgment adjustment module, which is connected to the text conversion processing module and the audio analysis processing module respectively, and is used to determine the text emotion tendency according to the text word length, the adjective ratio and the negative word ratio of the text content corresponding to a single purified audio signal, determine the audio emotion tendency according to the pitch feature trend and the frequency fluctuation tendency of the single purified audio signal, and determine the current emotional state of the user in combination with the text emotion tendency and the audio emotion tendency to determine whether to switch the voice output style of the interactive robot; Wherein, the emotional vocabulary categories include positive emotional vocabulary, negative emotional vocabulary and neutral emotional vocabulary; The emotion judgment adjustment module determines the text emotion tendency according to the text word length, the adjective ratio and the negative word ratio of the text content corresponding to the single purified audio signal, including: When the length of the text is less than a preset length, determining the sentiment tendency of the text according to the adjective ratio; Among them, if the adjective ratio is less than the preset adjective ratio, the emotional tendency of the text is determined to be negative emotional text.
2. The intelligent interactive robot control system according to claim 1, characterized in that: The text conversion processing module includes a positive sentiment word library and a negative sentiment word library for determining the sentiment word categories of adjectives and nouns, and determines the negative word ratio according to the determination results of the sentiment word category of each adjective and the sentiment word category of each noun; The negative word ratio is the ratio of the sum of the number of negative sentiment words of adjectives and nouns to the sum of the number of adjectives and nouns.
3. The intelligent interactive robot control system according to claim 1, characterized in that: The audio analysis and processing module determines the speech rate of a single purified audio signal according to the time length of the single purified audio signal and the length of the corresponding text word count; The speaking speed of a single purified audio signal is the ratio of the time length of the single purified audio signal to the length of the corresponding text word count.
4. The intelligent interactive robot control system according to claim 1, characterized in that: The audio analysis and processing module determines the fundamental frequency and fundamental frequency slope of a single time-domain audio image through a deep learning model, and determines the fundamental frequency characteristic trend according to whether the fundamental frequency and the fundamental frequency slope meet the fundamental frequency characteristic condition, including: If the fundamental frequency and the fundamental frequency slope meet the fundamental frequency characteristic condition, determining that the fundamental frequency characteristic trend is a negative fundamental frequency trend; If the fundamental frequency and the fundamental frequency slope do not satisfy the fundamental frequency characteristic condition, determining that the fundamental frequency characteristic trend is a positive fundamental frequency trend; The fundamental frequency characteristic condition is that the fundamental frequency is less than a frequency threshold and the fundamental frequency slope is less than a slope threshold.
5. The intelligent interactive robot control system according to claim 1, characterized in that: The audio analysis and processing module determines the frequency difference of each extreme value pair, and determines the characteristic extreme value pair according to the comparison result between the frequency difference and the preset frequency difference, including: If the frequency difference is greater than or equal to the preset frequency difference value, determining the corresponding extreme value pair as a characteristic extreme value pair; If the frequency difference is less than the preset frequency difference, it is determined that the corresponding extreme value pair is not a characteristic extreme value pair. Wherein, the extreme value pair consists of an adjacent maximum value and a minimum value; The preset frequency difference is determined according to the maximum audio difference, and the maximum audio difference is the difference between the maximum frequency and the minimum frequency in the time domain audio image.
6. The intelligent interactive robot control system according to claim 5, characterized in that: The audio analysis and processing module determines the frequency fluctuation tendency according to the ratio of the number of characteristic extreme value pairs to the number of all extreme value pairs, including: If the ratio of the number of characteristic extreme value pairs to the number of all extreme value pairs is less than a preset ratio, the frequency fluctuation tendency is determined to be a negative frequency fluctuation; If the ratio of the number of characteristic extreme value pairs to the number of all extreme value pairs is greater than or equal to a preset ratio, the frequency fluctuation tendency is determined to be a positive frequency fluctuation; The preset number ratio is determined according to the speech speed of the corresponding purified audio signal.
7. The intelligent interactive robot control system according to claim 1, characterized in that: The emotion judgment adjustment module determines the text emotion tendency according to the text word length, the adjective ratio and the negative word ratio of the text content corresponding to the single purified audio signal, and also includes: When the length of the text word count is greater than or equal to a preset length, determining the text sentiment tendency according to the adjective ratio and the negative word ratio; Among them, if the adjective ratio is greater than or equal to the preset adjective ratio and the negative word ratio is greater than or equal to the preset negative word ratio, the text sentiment tendency is determined to be a negative sentiment tendency.
8. The intelligent interactive robot control system according to claim 1, characterized in that: The emotion judgment adjustment module determines the audio emotion tendency according to the fundamental tone characteristic trend and frequency fluctuation tendency of a single purified audio signal, wherein: If the fundamental pitch feature trend is a negative fundamental pitch trend and the negative frequency fluctuation is a negative frequency fluctuation, then the audio emotional tendency is determined to be negative emotional audio.
9. The intelligent interactive robot control system according to claim 1, characterized in that: The emotion judgment adjustment module determines the current emotional state of the user by combining the text emotion tendency and the audio emotion tendency to determine whether to switch the voice output style of the interactive robot. include, If the emotional tendency of the text is negative emotional text and / or the emotional tendency of the audio is negative emotional audio, then the current emotional state is determined to be negative and the voice output style of the interactive robot is switched; If the emotional tendency of the text is positive emotional text and the emotional tendency of the audio is positive emotional audio, the current emotional state is determined to be positive emotion and the voice output style of the interactive robot is not switched.
Citation Information
Patent Citations
Robot-Human Interaction Method Based on Facial Expression Recognition and Emotion Quantization Analysis
CN116665273B
Man-machine interaction method and device and terminal equipment
CN117083581A
Voice text emotion analysis method and device based on emotion word bank
CN117253509A