Terminal interaction method and device based on multi-modal perception, terminal and medium
Through multimodal perception technology to obtain user information and combine it with large model analysis, the problem of single interaction mode in the existing technology is solved, rich and natural user-terminal interactions are achieved, and the interaction effect is improved.
Patent Information
- Application Number
- CN202510082905.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the interaction between the user and the terminal is single, especially the interaction between the user and the electronic pet in the terminal, and the lack of multimodal perception, resulting in the interaction being not smooth and natural enough, and the inability to accurately understand complex emotional expressions and personalized needs.
Using a terminal interaction method based on multimodal perception, by obtaining the user's voice information, expression information, gesture information and body information, combined with the big model to perform semantic recognition and emotional analysis, determine the interaction intention, and control the target object to perform corresponding interactive actions.
It realizes rich interaction forms, accurately determines the user's emotional expression and personalized needs, improves the interaction effect, and makes the interaction between the user and the electronic pets in the terminal more natural and intelligent.
Smart Images

Figure CN119987552A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of terminal interaction technology, and in particular to a terminal interaction method, device, terminal and medium based on multimodal perception. Background Art
[0002] Currently, the interaction between users and terminals is relatively simple, especially the interaction between users and specific objects in the terminal has certain limitations. For example, the interaction between users and electronic pets is generally only achieved through voice dialogue interaction, without other diversified interaction functions, and unable to achieve a rich interactive experience.
[0003] In addition, the existing technology also has certain limitations in speech recognition and understanding, resulting in unsmooth and unnatural interaction with the terminal. Moreover, the existing interaction methods are not accurate enough in understanding complex emotional expressions and personalized needs.
[0004] Therefore, the prior art still needs to be improved and enhanced. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a terminal interaction method, device, terminal and medium based on multimodal perception in view of the above-mentioned defects of the prior art. The technical solution adopted by the present invention is as follows:
[0006] In a first aspect, the present invention provides a terminal interaction method based on multimodal perception, wherein the method comprises:
[0007] Acquiring multimodal perception data, wherein the multimodal perception data includes any one or more of the user's voice information, expression information, gesture information, and body information;
[0008] Determining a target object displayed in the terminal, and determining an interaction intention for the target object based on the multimodal perception data;
[0009] Based on the interaction intention, the target object is controlled to perform an interaction action corresponding to the interaction intention.
[0010] In one implementation, the acquiring of multimodal sensing data includes:
[0011] Detecting audio data input by a user to obtain the voice information;
[0012] Acquire user image data, and perform expression recognition on the user image data to obtain the expression information;
[0013] Performing motion capture on the user image data to obtain the gesture information and the body information;
[0014] The multimodal perception data is obtained based on any one or more of the voice information, the expression information, the gesture information and the body information.
[0015] In one implementation, the determining the interaction intention for the target object based on the multimodal perception data includes:
[0016] If the multimodal perception data includes the voice information, semantic recognition and understanding of the voice information is performed based on a preset large model to obtain semantic information;
[0017] Based on the semantic information, an interaction intention for the target object is obtained.
[0018] In one implementation, the determining the interaction intention for the target object based on the multimodal perception data includes:
[0019] If the multimodal perception data includes the expression information, the expression information is recognized based on a preset large model to obtain emotion information corresponding to the expression information;
[0020] Based on the emotion information, an interaction intention for the target object is obtained.
[0021] In one implementation, the determining the interaction intention for the target object based on the multimodal perception data includes:
[0022] If the multimodal perception data includes the gesture information and / or body information, the gesture information and / or body information are recognized based on a preset large model to obtain an action intention corresponding to the gesture information and / or body information;
[0023] Based on the action intention, an interaction intention for the target object is obtained.
[0024] In one implementation, the method further includes:
[0025] If it is determined that the interaction intention is an emergency interaction intention, the target object is controlled to perform an interaction action corresponding to the emergency interaction intention, and suggestion information is generated based on the emergency interaction intention.
[0026] In one implementation, the method further includes:
[0027] Obtain user habit data or user preference data;
[0028] Based on the user habit data and the user preference data, an interaction solution is recommended.
[0029] In a second aspect, an embodiment of the present invention further provides a terminal interaction system based on multimodal perception, wherein the system includes:
[0030] A perception data acquisition module, used to acquire multimodal perception data, wherein the multimodal perception data includes any one or more of the user's voice information, expression information, gesture information, and body information;
[0031] An interaction intention determination module, configured to determine a target object displayed in the terminal, and determine an interaction intention for the target object based on the multimodal perception data;
[0032] An interaction action execution module is used to control the target object to execute the interaction action corresponding to the interaction intention based on the interaction intention.
[0033] In a third aspect, an embodiment of the present invention further provides a terminal, wherein the terminal includes a memory, a processor, and a terminal interaction program based on multimodal perception stored in the memory and executable on the processor, and when the processor executes the terminal interaction program based on multimodal perception, the steps of the terminal interaction method based on multimodal perception of any one of the above-mentioned schemes are implemented.
[0034] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein a terminal interaction program based on multimodal perception is stored on the computer-readable storage medium, and when the terminal interaction program based on multimodal perception is executed by a processor, the steps of the terminal interaction method based on multimodal perception described in any one of the above-mentioned schemes are implemented.
[0035] Beneficial effects: Compared with the prior art, the present invention provides a terminal interaction method based on multimodal perception. The present invention first obtains multimodal perception data, wherein the multimodal perception data includes any one or more of the user's voice information, expression information, gesture information, and body information. Then, the target object displayed in the terminal is determined, and the interaction intention for the target object is determined based on the multimodal perception data. Finally, based on the interaction intention, the target object is controlled to perform the interaction action corresponding to the interaction intention. The present invention can realize the interaction with the target object displayed in the terminal through multimodal perception data of voice, gesture, expression and other information, enriching the interaction form. And the present invention can accurately determine the user's emotional expression and personalized needs through multimodal perception data, thereby improving the interaction effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A flowchart of a preferred embodiment of a terminal interaction method based on multimodal perception provided in an embodiment of the present invention.
[0037] Figure 2 A logic diagram of a terminal interaction method based on multimodal perception provided by an embodiment of the present invention.
[0038] Figure 3 A schematic diagram of the architecture of a terminal interaction system based on multimodal perception provided by an embodiment of the present invention.
[0039] Figure 4 A functional block diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0041] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations or steps, nor must they be executed in the order described. For example, some operations or steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0042] It should be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0043] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, words such as "first" and "second" are used to distinguish between identical or similar items with substantially identical functions and effects. For example, the first control information and the second control information are only used to distinguish different control information, and their order is not limited.
[0044] Those skilled in the art can understand that the words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not necessarily limit the differences.
[0045] It should be further understood that the term “and / or” used in the present specification and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0046] Since the interaction mode between users and terminals in the prior art is single, especially the interaction between users and electronic pets in the terminal is basically based on voice interaction to realize the control of certain functions, and there is no interaction mode similar to feeding and nurturing. In addition, the prior art is insufficient in voice recognition and understanding, resulting in the interaction with electronic pets not being smooth and natural enough. For example, some dialects or vague voice commands may not be accurately recognized, and the understanding of complex emotional expressions and personalized needs is not accurate enough. Although the current terminal screen can provide a certain visual display, most electronic pets only interact based on voice and simple animation effects, lacking multi-modal interaction methods such as touch and somatosensory, and cannot provide a rich interactive experience like some mobile electronic pet applications, such as realizing more interesting interactions with pets through the gravity sensor and camera of the mobile phone.
[0047] In order to solve the problems of the prior art, the present embodiment provides a terminal interaction method based on multimodal perception. The method based on the present embodiment can realize the interaction with the target object displayed in the terminal through multimodal perception data of voice, gesture, expression and other information, thereby enriching the interaction form. And the present invention can accurately determine the user's emotional expression and personalized needs through multimodal perception data, thereby improving the interaction effect. Specifically, the present embodiment first obtains multimodal perception data, wherein the multimodal perception data includes any one or more of the user's voice information, expression information, gesture information and body information. Then, the target object displayed in the terminal is determined, and the interaction intention for the target object is determined based on the multimodal perception data. Finally, based on the interaction intention, the target object is controlled to perform the interaction action corresponding to the interaction intention.
[0048] The terminal interaction method based on multimodal perception of this embodiment can be applied to terminals, including intelligent product terminals such as televisions, computers, and mobile phones. Figure 1 As shown in , the terminal interaction method based on multimodal perception includes the following steps:
[0049] Step S100: Acquire multimodal perception data, wherein the multimodal perception data includes any one or more of the user's voice information, expression information, gesture information, and body information.
[0050] This embodiment first collects multimodal perception data, and based on these multimodal perception data, realizes global perception of user status, so as to fully understand the user's personalized needs and emotional expression. The multimodal perception data of this embodiment includes any one or more of the user's voice information, expression information, gesture information, and body information, which greatly improves the dimension of perception data and is conducive to the recognition of the user's interaction intention in subsequent steps.
[0051] The interactive functions of electronic pets in current terminals are relatively simple or even missing. This embodiment can obtain multimodal perception data by collecting the user's voice or image and analyzing the collected data. These multimodal perception data can reflect the user's personalized needs and emotional expressions, and are therefore conducive to enriching the realism and fun of the interaction with the electronic pet. Figure 2 As shown, this embodiment first detects the audio data input by the user to obtain the voice information. When detecting the audio data, this embodiment can use the audio device preset in the terminal to collect and detect. Then, the user image data can also be collected, and the user image data can be subjected to expression recognition to obtain the expression information. When collecting the user image data, this embodiment can use the camera device preset on the terminal to capture the user image data based on the camera device, so as to obtain the expression information, which can reflect the user's emotions. This embodiment can also perform motion capture on the user image data to obtain the gesture information and the body information, such as the user's stroking gesture, and the body information such as jumping, turning in circles and other body movements. These gesture information and body information can be used as the basis for the interaction with the electronic pet in the terminal in the subsequent steps. After collecting the voice information, the expression information, the gesture information and the body information, this embodiment can use any one or more of the voice information, the expression information, the gesture information and the body information as the multimodal perception data.
[0052] Step S200: determine a target object displayed in the terminal, and determine an interaction intention for the target object based on the multimodal perception data.
[0053] This embodiment first determines the target object displayed in the terminal, which is an electronic pet in the terminal (such as a TV) that can interact with the user. In actual applications, the electronic pet in the terminal can also be a voice assistant in the terminal, so that the voice assistant can not only interact with the user, but also control the terminal itself based on the voice assistant.
[0054] Combination Figure 2As shown, after the multimodal perception data is obtained in this embodiment, the interaction intention for the target object can be determined based on these multimodal perception data. The interaction intention reflects the personalized interaction needs of the user. In practical applications, when the multimodal perception data includes the voice information, the embodiment can perform semantic recognition and understanding on the voice information based on the preset large model to obtain semantic information. Then, based on the semantic information, the interaction intention for the target object is obtained. The large model in this embodiment is trained on massive data, which greatly improves the accuracy of voice information recognition and the tolerance of various accents and dialects. Whether it is a dialect from different regions or Mandarin with an accent, it can be accurately recognized, effectively expanding the scope of application of the user group, so that users with different language habits can communicate with electronic pets smoothly by voice. In addition, because traditional technology can only process simple and direct instructions in semantic understanding, it is difficult to understand implicit semantics and complex contexts. The large model in this embodiment has a strong semantic understanding ability, which can deeply analyze the potential intentions in the user's speech and realize multiple rounds of coherent dialogue. For example, when users express content related to their pets' emotions, the system can not only detect the problem, but also further ask for details and give targeted suggestions. This deep semantic understanding and interactive capability makes the communication between electronic pets and users more natural and intelligent, far exceeding the interactive level of existing technologies.
[0055] If the multimodal perception data includes the facial expression information, this embodiment can identify the facial expression information based on a preset large model to obtain the emotional information corresponding to the facial expression information. Then, based on the emotional information, the interaction intention for the target object is obtained. The large model of this embodiment can identify the facial expression information, capture the user's emotional changes, and convert it into a basis for interaction with the electronic pet. Since the prior art rarely involves the application of facial expression recognition in the field of electronic pets. This embodiment uses facial expression recognition technology to enable the terminal to capture the user's facial expression changes and determine the interaction intention of the facial expression information. For example, through the user's joys, sorrows, anger, and happiness, the electronic pet makes corresponding companionship or interactive behaviors. This emotion-based interaction method can establish a closer emotional connection between the user and the electronic pet, so that the electronic pet is no longer a simple virtual image, but a companion who can perceive the user's emotions. This is a major breakthrough in the existing interaction mode.
[0056] If the multimodal perception data includes the gesture information and / or limb information, the present embodiment can identify the gesture information and / or limb information based on a preset large model to obtain the action intention corresponding to the gesture information and / or limb information. Then, the present embodiment can obtain the interaction intention for the target object based on the action intention. Since the gesture interaction function of the electronic pet of the current terminal is relatively simple or even missing. The present embodiment can capture the user's gesture information based on the camera device and determine the interaction intention corresponding to the gesture information, so that the user can interact with the electronic pet through a variety of gestures, such as stroking, controlling actions, etc. So that the electronic pet will make realistic reactions according to the strength and direction of the gesture, greatly enhancing the realism of the interaction. At the same time, the various gesture information can control the actions of the electronic pet, enrich the fun and controllability of the interaction, and bring a new interactive experience to the user, which is difficult to achieve with the existing technology. In addition, the electronic pet of the current terminal has almost no effective limb interaction function. The present embodiment captures the user's actions through the camera device, determines the limb information, and then determines the interaction intention corresponding to the limb information. For example, the user can play interactive imitation games with the electronic pet, and can also issue instructions through limb movements in the pet training scene. This kind of in-depth physical interaction expands the way users interact with electronic pets, enriches the interactive experience, and brings innovative changes to the interactive mode of electronic pets on terminals.
[0057] Step S300: Based on the interaction intention, control the target object to perform an interaction action corresponding to the interaction intention.
[0058] After the interaction intention is determined, the embodiment can control the target object, i.e., the electronic pet, to perform the interaction action corresponding to the interaction intention. Since the interaction intention of the embodiment is determined based on multimodal perception data, and the multimodal perception data includes any one or more of the voice information, the expression information, the gesture information, and the body information. Therefore, combined with Figure 2 As shown, this embodiment can classify the determined interaction intentions based on the interaction classification model to trigger dialogue, expression, gesture or body language engine, thereby facilitating the electronic pet to perform different interaction actions.
[0059] Specifically, if the interaction intention is determined based on the voice information, this embodiment can control the target object in the terminal to perform the interaction action corresponding to the interaction intention of the voice information. For example, if the voice information is "My pet doesn't seem very happy", the corresponding interaction intention is to query the details of the electronic pet and understand the recent behavior of the electronic pet. Therefore, this embodiment can retrieve the recent behavior information of the electronic pet and give targeted suggestions, such as recommending a suitable interactive game, and controlling the electronic pet to play interactive games with the user to improve the pet's mood.
[0060] If the interaction intention is determined based on the expression information, this embodiment can control the target object in the terminal to perform the interaction action corresponding to the interaction intention of the expression information. For example, when the user shows a happy expression, the electronic pet may run around the screen happily, showing an excited state; if the user shows a sad expression, the electronic pet will snuggle in a corner of the screen, showing a comforting gesture. This interaction based on expression information can enable the electronic pet to better perceive the user's emotions and establish a closer emotional connection.
[0061] If the interaction intention is determined based on gesture information and / or body information, this embodiment can control the target object in the terminal to perform the interaction action corresponding to the interaction intention of the gesture information and / or body information. For example, when the user makes a stroking gesture, the electronic pet on the terminal will respond accordingly according to the strength and direction of the gesture, such as squinting comfortably, wagging its tail, etc., so that the user can feel a more realistic interactive experience. In addition, some actions of the pet can also be controlled by gesture information, such as waving to make the pet jump, clenching fists to make the electronic pet sit down, etc., to increase the fun and controllability of the interaction. For example, the user can guide the electronic pet to make the same response by imitating the actions of the electronic pet, such as jumping, turning in circles and other body information, to form an interesting interactive imitation game. In addition, when the user makes some large-scale body movements, such as opening his arms, the electronic pet may approach the edge of the screen like jumping to the owner, creating a stronger emotional communication atmosphere. In the pet training scene, the user can train the electronic pet to complete a specific task through specific body information, such as pointing in a certain direction to let the electronic pet go to a designated location to explore, enriching the way and experience of pet training.
[0062] In other implementations, if the embodiment determines that the interaction intention is an emergency interaction intention based on multimodal perception data, the target object is controlled to perform the interaction action corresponding to the emergency interaction intention, and suggestion information is generated based on the emergency interaction intention. For example, when the user speaks a voice message in an anxious tone: "What should I do if my pet is sick?", while making nervous gesture information and worried expression information, the multimodal perception data includes voice information, expression information and gesture information. The system can quickly determine that the interaction intention is an emergency interaction intention. Therefore, the interaction action corresponding to the emergency interaction intention can be executed first, and professional pet medical advice and related resources and other suggestion information can be provided first. In addition, the embodiment can also obtain user habit data or user preference data. Then, based on the user habit data and the user preference data, an interaction plan is recommended to create a personalized electronic pet growth and interaction model for each user, continuously improve the user's participation and satisfaction, and make the terminal's electronic pet a true companion in the user's life.
[0063] In summary, this embodiment achieves global perception of the user's voice information, gesture information, expression information and other information through the combination of a large model and multimodal perception data. The system can quickly determine the user's needs and emotions, and can quickly provide professional advice and resources when the user encounters a problem, such as the electronic pet is sick. At the same time, it creates a personalized growth and interaction model based on the user's long-term usage habits and preferences, comprehensively improves the user experience, and makes the terminal's electronic pet truly a close partner in the user's life. This is unmatched by existing technologies in optimizing user experience.
[0064] Based on the above embodiments, the present invention also provides a terminal interaction system based on multimodal perception, such as Figure 3 As shown in , the system includes: a perception data acquisition module 10, an interaction intention determination module 20 and an interaction action execution module 30. Specifically, the perception data acquisition module 10 is used to acquire multimodal perception data, wherein the multimodal perception data includes any one or more of the user's voice information, expression information, gesture information and body information. The interaction intention determination module 20 is used to determine the target object displayed in the terminal, and determine the interaction intention for the target object based on the multimodal perception data. The interaction action execution module 30 is used to control the target object to perform the interaction action corresponding to the interaction intention based on the interaction intention.
[0065] In one implementation, the perception data acquisition module 10 includes:
[0066] A voice information determination unit, used to detect the audio data input by the user to obtain the voice information;
[0067] An expression information determination unit, used to obtain user image data, and perform expression recognition on the user image data to obtain the expression information;
[0068] A limb information determination unit, used for performing motion capture on the user image data to obtain the gesture information and the limb information;
[0069] A perception data determination unit is used to obtain the multimodal perception data based on any one or more of the voice information, the expression information, the gesture information and the limb information.
[0070] In one implementation, the interaction intention determination module 20 includes:
[0071] a semantic information determination unit, configured to, if the multimodal perception data includes the voice information, perform semantic recognition and understanding on the voice information based on a preset large model to obtain semantic information;
[0072] The first interaction intention determination unit is used to obtain the interaction intention for the target object based on the semantic information.
[0073] In one implementation, the interaction intention determination module 20 includes:
[0074] an emotion information determination unit, configured to, if the multimodal perception data includes the expression information, identify the expression information based on a preset large model to obtain emotion information corresponding to the expression information;
[0075] The second interaction intention determination unit is used to obtain the interaction intention for the target object based on the emotion information.
[0076] In one implementation, the interaction intention determination module 20 includes:
[0077] an action intention determination unit, configured to, if the multimodal perception data includes the gesture information and / or limb information, identify the gesture information and / or limb information based on a preset macro model to obtain an action intention corresponding to the gesture information and / or limb information;
[0078] The third interaction intention determination unit is used to obtain the interaction intention for the target object based on the action intention.
[0079] In one implementation, the system further includes:
[0080] A suggestion information generating unit is used to control the target object to perform an interaction action corresponding to the emergency interaction intention if it is determined that the interaction intention is an emergency interaction intention, and to generate suggestion information based on the emergency interaction intention.
[0081] In one implementation, the system further includes:
[0082] A data acquisition unit, used to acquire user habit data or user preference data;
[0083] An interaction scheme recommendation unit is used to recommend an interaction scheme based on the user habit data and the user preference data.
[0084] The working principles of each module in the terminal interaction system based on multimodal perception of this embodiment are the same as the principles of each step in the above method embodiment, and will not be repeated here.
[0085] Each module in the terminal interaction system based on multimodal perception can be implemented in whole or in part by software, hardware and their combination. Each module can be embedded in or independent of the processor in the terminal in the form of hardware, or can be stored in the memory in the terminal in the form of software, so that the processor can call and execute the operations corresponding to each module above.
[0086] Based on the above embodiment, the present invention further provides a terminal, the principle block diagram of the terminal can be as follows: Figure 4 The terminal may include one or more processors 100 ( Figure 4 Only one is shown), a memory 101 and a computer program 102 stored in the memory 101 and executable on one or more processors 100.
[0087] In one embodiment, the processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0088] In one embodiment, the memory 101 may be an internal storage unit of an electronic device, such as a hard disk or memory of the electronic device. The memory 101 may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 101 may also include both an internal storage unit of the electronic device and an external storage device. The memory 101 is used to store computer programs and other programs and data required by the terminal. The memory 101 may also be used to temporarily store data that has been output or is to be output.
[0089] Those skilled in the art will understand that Figure 4 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the scheme of the present invention, and does not constitute a limitation on the terminal to which the scheme of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0090] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, operating database or other media used in the embodiments provided by the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A terminal interaction method based on multimodal perception, characterized in that: The method comprises: Acquiring multimodal perception data, wherein the multimodal perception data includes any one or more of the user's voice information, expression information, gesture information, and body information; Determining a target object displayed in the terminal, and determining an interaction intention for the target object based on the multimodal perception data; Based on the interaction intention, the target object is controlled to perform an interaction action corresponding to the interaction intention.
2. The terminal interaction method based on multimodal perception according to claim 1, characterized in that: The acquiring of multimodal perception data includes: Detecting audio data input by a user to obtain the voice information; Acquire user image data, and perform expression recognition on the user image data to obtain the expression information; Performing motion capture on the user image data to obtain the gesture information and the body information; The multimodal perception data is obtained based on any one or more of the voice information, the expression information, the gesture information and the body information.
3. The terminal interaction method based on multimodal perception according to claim 1, characterized in that: The determining the interaction intention for the target object based on the multimodal perception data includes: If the multimodal perception data includes the voice information, semantic recognition and understanding of the voice information is performed based on a preset large model to obtain semantic information; Based on the semantic information, an interaction intention for the target object is obtained.
4. The terminal interaction method based on multimodal perception according to claim 1, characterized in that: The determining the interaction intention for the target object based on the multimodal perception data includes: If the multimodal perception data includes the expression information, the expression information is recognized based on a preset large model to obtain emotion information corresponding to the expression information; Based on the emotion information, an interaction intention for the target object is obtained.
5. The terminal interaction method based on multimodal perception according to claim 1, characterized in that: The determining the interaction intention for the target object based on the multimodal perception data includes: If the multimodal perception data includes the gesture information and / or body information, the gesture information and / or body information are recognized based on a preset macro model to obtain an action intention corresponding to the gesture information and / or body information; Based on the action intention, an interaction intention for the target object is obtained.
6. The terminal interaction method based on multimodal perception according to claim 1, characterized in that: The method further comprises: If it is determined that the interaction intention is an emergency interaction intention, the target object is controlled to perform an interaction action corresponding to the emergency interaction intention, and suggestion information is generated based on the emergency interaction intention.
7. The terminal interaction method based on multimodal perception according to claim 1, characterized in that: The method further comprises: Obtain user habit data or user preference data; Based on the user habit data and the user preference data, an interaction solution is recommended.
8. A terminal interaction system based on multimodal perception, characterized in that: The system comprises: A perception data acquisition module, used to acquire multimodal perception data, wherein the multimodal perception data includes any one or more of the user's voice information, expression information, gesture information, and body information; An interaction intention determination module, configured to determine a target object displayed in the terminal, and determine an interaction intention for the target object based on the multimodal perception data; An interaction action execution module is used to control the target object to execute the interaction action corresponding to the interaction intention based on the interaction intention.
9. A terminal, characterized in that: The terminal includes a memory, a processor, and a terminal interaction program based on multimodal perception stored in the memory and executable on the processor. When the processor executes the terminal interaction program based on multimodal perception, the steps of the terminal interaction method based on multimodal perception as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a terminal interaction program based on multimodal perception. When the terminal interaction program based on multimodal perception is executed by the processor, the steps of the terminal interaction method based on multimodal perception as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Multi-modal interaction method and device
CN114995636A
Multi-modal model for dynamically responsive virtual characters
US20240303891A1