Multi-modal man-machine interaction optimization system integrating multiple large language models

By integrating multiple large language models, the human-computer interaction of intelligent humanoid robots is optimized, solving the problems of single mode and low flexibility, realizing a personalized and natural multimodal interactive experience, and improving the flexibility and personalization of user interaction.

CN120909547APending Publication Date: 2025-11-07LINGTONG ROBOT (SHANGHAI) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511015630.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing human-computer interaction technologies for intelligent humanoid robots are limited in their modes of operation, flexibility, and adaptability, lacking personalization. Furthermore, the interaction systems are insufficient in utilizing user contextual information, resulting in an unnatural and unpersonalized interactive experience.

Method used

By integrating multiple large language models, and through natural language processing, prompting engineering, and multimodal interaction design, it generates customized feedback that matches the persona, emotion, and context. Combined with speech recognition, speech synthesis, function calls, and keyword matching units, it optimizes the human-computer interaction experience.

Benefits of technology

It improves the naturalness, flexibility, and personalization of human-computer interaction, achieving accurate speech recognition and rapid response in diverse language environments, and dialogue responses that are consistent with emotional and behavioral representations and character personalities, thus enhancing the user's interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909547A_ABST
    Figure CN120909547A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal man-machine interaction optimization system integrating multiple large language models, and focuses on the man-machine interaction field of intelligent humanoid robots. And by fusing technologies such as prompt engineering, natural language processing and personalized role configuration, the flexibility, adaptability and personalized level of human-computer interaction are remarkably improved. According to the technology, the advantages of various large language models are utilized, complex dialogue logic is processed, personalized voice, emotion and action feedback matched with human settings, emotions and contexts is generated, and therefore real-time human-computer interaction is optimized. According to the specific implementation, the system comprises a voice awakening unit, a voice recognition unit, a function calling unit, an anthropomorphic dialogue unit, a keyword matching unit and a voice synthesis unit, and finally highly natural and immersive man-machine interaction experience is achieved. The method shows excellent performance in the aspects of speech recognition accuracy, response speed, emotion action representation, personalized interaction and the like, and provides more vivid, immersive and personalized user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of human-computer interaction of intelligent humanoid robots, and particularly relates to a multi-modal human-computer interaction optimization system integrating multiple large language models. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, the market has higher requirements for the personalization, interactivity and immersion of intelligent humanoid robots, so it is urgent to improve and optimize the existing human-computer interaction technology. Traditional human-computer interaction systems usually rely on single-mode interaction, such as controlling robots to perform tasks through specific voice commands. However, this method has obvious limitations in realizing multi-round, multi-theme and multi-style dialogue interaction, and lacks sufficient flexibility and adaptability. In addition, the existing interaction system lacks sufficient consideration of user context information, resulting in unnatural and lack of personalized interaction experience. Large language models are deep neural network models trained using massive amounts of unlabeled data, with powerful natural language understanding, generation and generalization capabilities, and have been widely applied in virtual assistants, game development and other fields. Based on this, a multi-modal human-computer interaction optimization technology integrating multiple large language models is proposed. This technology integrates the advantages of multiple large language models, expands the creativity and possibilities of human-computer interaction, including personalized character personality, language style, etc. At the same time, by comprehensively considering the context information, processing complex dialogue logic, and generating replies that adapt to the character, emotion, and context, real-time voice interaction and personalized emotional action representation functions are realized. SUMMARY

[0003] The present application proposes a multi-modal human-computer interaction optimization technology integrating multiple large language models to address the problems of single mode, low flexibility, poor adaptability and lack of personalization in current human-computer interaction technology of intelligent humanoid robots. By deeply analyzing the dialogue content and generating customized feedback that conforms to the character personality characteristics according to the character setting or personalized needs of the user, the user interaction experience is improved.

[0004] The present application is implemented by the following technical solutions:

[0005] The present application adopts natural language processing, prompt engineering, personalized feedback generation and multi-modal interaction design key technologies, and significantly optimizes the real-time human-computer interaction experience by integrating multiple large language models. With the help of large language model API, natural language processing technology and prompt engineering are used to deeply analyze and accurately extract key information from user dialogue content. On this basis, customized feedback that adapts to the character, emotion and context is generated according to the character setting or personalized needs of the user, further enhancing the interactivity and immersion between the user and the intelligent humanoid robot.

[0006] The method specifically includes:

[0007] The first step is to collect voice input based on a sound card, convert the voice into text through voice recognition technology, and preprocess the text, such as removing spaces, line breaks, and other special symbols, and determining whether the text is empty. This step is the key to realizing natural dialogue with an intelligent humanoid robot, and ensures accurate input of dialogue content as the basis for subsequent processing.

[0008] The second step is to set up a function call (FunctionCall) large language model, use the logical capabilities of the large language model to do tool routing, and customize a tool set containing functions such as perceiving time, querying weather, and playing music to expand the capabilities of the large language model.

[0009] The third step is to pre-set the role library of the anthropomorphic dialogue large language model (including role characteristics and intimacy) through the prompting project, and transmit the text data to the large language model to generate a reply. With the deep analysis capabilities of the large language model, key information is extracted from the dialogue content, and a matching reply is generated according to the role personality characteristics or the user's personalized needs, ensuring efficient understanding and response to complex dialogues.

[0010] The fourth step is to pre-set the general large language model through the emotion and action library and the prompting project to realize the keyword extraction function. Emotion and action representations are extracted from the generated reply content and input into the large language model, and the corresponding emotion and action instructions are matched in the pre-set library to express the emotional state and reaction of the robot, enhancing the naturalness and immersion of the interaction.

[0011] The fifth step is to set the tone, speed, etc. according to the role, and use text-to-speech technology to output the dialogue interaction.

[0012] The technology involved in the present application includes a voice wake-up unit, a voice recognition unit, a function call unit, an anthropomorphic dialogue unit, a keyword matching unit, and a speech synthesis unit. The voice wake-up unit monitors and responds to the wake-up word in real time to switch the system to the working state; the voice recognition unit converts the sound signal into a text signal and performs preprocessing to ensure that the text meets the input format of the large language model; the function call unit uses the logical capabilities of the large language model to do tool routing, and expands the functions of the large language model through tool customization, thereby enhancing its learning and generalization capabilities; the anthropomorphic dialogue unit realizes personalized dialogue interaction and emotion and action representation through the native interface of the anthropomorphic dialogue large language model; the keyword matching unit matches the emotion and action instructions in the pre-set library based on the general large language model to control the robot to perform tasks and express emotions; the speech synthesis unit generates and plays sound signals with personalized tones based on the output text information and the speech large model. The above steps form the core technical route of the present application, which optimizes and enhances the human-computer interaction experience, and provides users with more accurate and rich interactive ways.

[0013] The beneficial effects of this invention are:

[0014] This invention demonstrates high accuracy in speech recognition, accurately recognizing user voice commands in diverse language environments, including supporting both Chinese and English, and under varying speech rates and background noise conditions. Regarding response speed, the response time from the user issuing a voice command to the robot's reply is controlled within a short period (less than 15 seconds) to provide a smooth interactive experience. In terms of emotional and behavioral representation consistent with the character's personality, advanced large-scale language model technology is used to ensure that the robot's dialogue responses and emotional actions are adapted to its role and context. In terms of emotion and action keyword matching, it can accurately extract emotion and action keywords from the dialogue, thereby triggering specific actions that match the character's personality and emotional state. The proposed method enhances the naturalness, flexibility, and personalization of human-computer interaction, providing users with a more vivid and personalized interactive experience. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the system flow of the present invention;

[0017] Figure 2 This is a schematic diagram of the system flow of the present invention. Detailed Implementation

[0018] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should also be noted that, to make the embodiments more comprehensive, the following embodiments are the best and preferred embodiments, and those skilled in the art can use other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.

[0019] It should be noted that the use of terms such as "an embodiment," "an embodiment," "an exemplary embodiment," and "some embodiments" in the specification indicates that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment necessarily includes that specific feature, structure, or characteristic. Furthermore, when a specific feature, structure, or characteristic is described in connection with an embodiment, implementing such a feature, structure, or characteristic in conjunction with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the art.

[0020] Generally, terms can be understood at least in part from the context in which they are used. For example, the term "one or more" as used herein, depending at least in part upon context, can be used to describe any feature, structure, or characteristic in a singular sense or can be used to describe combinations of features, structures or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily being confined to factors that are explicitly enumerated, but rather can instead be broadly understood as including factors that are not explicitly enumerated, at least in part depending on the context.

[0021] As shown in Figure 1 , Figure 2 The present embodiment takes the Elsa character in Frozen as an example, based on the general large language model, anthropomorphic dialogue large language model and voice large language model of PatSnap, carries out multi-modal human-computer interaction optimization technology integrating multiple large language models, specifically including:

[0022] First, develop a wake-up word detection algorithm and set the wake-up word to "HiElsa". When the wake-up word is detected, the system switches to the working state.

[0023] Second, convert the user's voice signal into text through voice recognition technology. This step ensures that the user's input voice content can be accurately captured and understood by the system.

[0024] Third, preprocess the converted text to determine if the text is empty and remove special symbols such as spaces and line breaks, standardizing the text into the input format preset by the large language model to ensure the accuracy and consistency of subsequent processing.

[0025] Fourth, set the function call function based on the general large language model. The specific parameters are shown in Table 1. Based on the logical capabilities of the large language model, develop tool routing to customize a set of tools including time perception, weather query, music playback, and other functions to expand the capabilities of the large language model. Specifically, develop a time query function based on Python, input the current time, and calculate past, present or future time through the model's logical capabilities; based on Python, access the API of the openweathermap weather query website to realize functions such as climate query and dressing suggestion; develop music query and playback functions based on Python.

[0026] Table 1 Function call function based on general large language model

[0027]

[0028]

[0029] The fifth step is to set the role library of the personified dialogue large language model based on the prompt engineering, set the intimacy and role characteristics of the role, access the large language model through the native interface, and generate dialogue content consistent with the characteristics of the Elsa role. The personalized role configuration based on the personified dialogue large language model is shown in Table 2.

[0030] Table 2 Personalized role configuration based on personified dialogue large language model

[0031]

[0032]

[0033] The sixth step is to separate the output content of the personified dialogue large language model. The output content of the personified dialogue large language model includes: the content in the brackets contains emotions and actions, and the content outside the brackets is the dialogue content. The program is designed to extract the content in and outside the brackets respectively, and save it as "emotion and action reply" and "dialogue reply".

[0034] The seventh step is to set the keyword matching function based on the general large language model, input the preset emotion library, action library and "emotion and action reply" based on natural language, realize the matching of emotion keywords and action keywords, and generate emotion and action sequences matched with the Elsa role. The keyword matching function setting based on the general large language model is shown in Table 3.

[0035] Table 3 Keyword matching function setting based on general large language model

[0036]

[0037]

[0038]

[0039] The eighth step is to convert the "dialogue reply" content through text-to-speech technology, select the tone, speed and other characteristics consistent with the character, and realize natural and smooth human-computer interaction.

[0040] The present application covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present application. In order to make the public have a thorough understanding of the present application, specific details are described in the following preferred embodiments of the present application, and the present application can also be fully understood without the description of these details to those skilled in the art. In addition, in order to avoid unnecessary confusion to the essence of the present application, well-known methods, processes, procedures, elements and circuits are not described in detail.

[0041] The above merely describes the preferred embodiments of the present application, and it should be pointed out that, for those skilled in the art, some improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as falling within the protection scope of the present application.

Claims

1. A multi-modal human-computer interaction optimization system integrating multiple large language models, characterized in that, Comprise the following steps: The first step is to collect voice input based on sound card, convert voice to text through voice wake-up and voice recognition technology, and preprocess the text to ensure accurate input of conversation content; The second step is to set up a function to call a large language model, use the logical ability of the large language model to do tool routing, and customize a tool set containing time perception, weather query, and music playing functions to expand the capabilities of the large language model The third step is to generate a reply by transmitting text data to the large language model through a prompt engineering preset role library of the personified dialogue large language model (containing role characteristics and intimacy), extracting key information from the dialogue content, and generating a matching reply; The fourth step is to achieve keyword extraction function by presetting a general large language model through emotion, action library and prompt engineering, extracting emotion and action representation in the generated reply content and inputting it into the large language model to match the corresponding emotion and action instructions; The fifth step is to select the tone according to the characteristics of the role and set the speed, and output the dialogue interaction by using the text-to-speech technology.

2. The multi-modal human-computer interaction optimization system integrating multiple large language models according to claim 1, wherein, The system integrates three large language models, which realize the functions of function call, personified dialogue and keyword matching through differentiated prompt engineering.

3. The multi-modal human-machine interaction optimization system integrating multiple large language models according to claim 1, wherein, In the second step, based on the function call of the large language model, the tool routing is developed, and a tool set containing time perception, weather query, and music playing functions is customized to expand the learning and generalization ability of the large language model.

4. The multi-modal human-machine interaction optimization system integrating multiple large language models according to claim 1, wherein, In the third step, the personified dialogue large language model API is accessed, and the role library and dialogue role are preset. Through the prompt engineering, the role characteristics and the intimacy between roles are customized. The personified dialogue large language model considers the context information comprehensively, processes complex dialogue logic, generates a plot that matches the role personality, emotion, and context, and realizes personalized dialogue interaction and emotion and action representation.

5. The multi-modal human-machine interaction optimization system integrating multiple large language models according to claim 1, wherein, In the fourth step, based on the general large language model, a keyword matching unit is constructed to match the plot of the personified dialogue large model to the emotion and action instructions in the preset library to control the robot to perform tasks and express emotions. This method can accurately extract emotion and action keywords from the dialogue, thereby triggering specific actions that match the role IP personality and emotional state.

6. The multi-modal human-machine interaction optimization system integrating multiple large language models according to claim 1, wherein, In the fifth step, the speech synthesis algorithm is designed according to the output text information to generate and play sound signals that meet the role characteristics.

Citation Information

Cited By

  • Biomimetic robot cluster dance action consistency correction method based on state feedback

    CN122378754A