A multi-modal interactive information recognition method, device and equipment and storage medium

By using a multimodal interactive information recognition method, combined with a multimodal knowledge graph and an intelligent question-answering system, the problem of traditional set-top boxes being unable to accurately interpret user intent in voice interaction has been solved, achieving accurate response and personalized answers in multimodal scenarios.

CN116915528BActive Publication Date: 2026-04-07CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Traditional set-top boxes cannot accurately interpret user intent during voice interaction, especially when faced with colloquial statements and multimodal interactions. They struggle to recognize complex user questions, resulting in an inability to provide accurate rich media responses.

Method used

A multimodal interactive information recognition method is adopted. By obtaining the interactive information to be recognized and the multimodal scene recognition information in the interactive scenario, combined with the multimodal knowledge graph and intelligent question answering system, the target question is located and the answer is output in a rich media response mode, including answers in various forms such as animation, sound and video.

Benefits of technology

It achieves accurate parsing and flexible response to user intent in multimodal interaction scenarios, and can provide personalized rich media answers based on user feedback, thereby improving user experience and the accuracy of problem parsing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116915528B_ABST
    Figure CN116915528B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and storage medium for recognizing multimodal interaction information. The method includes: obtaining interaction information to be recognized in an interaction scenario; obtaining multimodal scene recognition information; wherein the multimodal scene recognition information is scene information associated with the interaction information to be recognized; locating the target question hit by the interaction information to be recognized based on the interaction information to be recognized and the multimodal scene recognition information; obtaining the rich media response mode corresponding to the target question, and outputting the answer to the target question in the rich media response mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a method, apparatus, device and storage medium for identifying multimodal interactive information. Background Technology

[0002] With the increasing prevalence of electronic devices such as smart home devices, smart home devices that can provide multimedia content are gradually becoming the first type of smart device that every family chooses. However, during the use of smart home devices, such as traditional set-top boxes, voice interaction relies solely on voice information for intent interpretation, which can lead to inaccurate interpretation of user intentions. Summary of the Invention

[0003] This application aims to provide a method for recognizing multimodal interactive information, solving the problem in related technologies where only voice information is used for related intent parsing, which cannot accurately interpret user intent.

[0004] The technical solution of this application is implemented as follows:

[0005] A method for recognizing multimodal interaction information, the method comprising:

[0006] Obtain the interaction information to be identified in the interaction scenario;

[0007] Obtain multimodal scene recognition information; wherein, multimodal scene recognition information is scene information associated with the interaction information to be recognized;

[0008] Based on the interaction information to be identified and the multimodal scene recognition information, locate the target problem that the interaction information to be identified hits;

[0009] Obtain the rich media response corresponding to the target question, and output the answer to the target question in the rich media response format.

[0010] A device for recognizing multimodal interaction information, the device comprising:

[0011] The acquisition module is used to acquire the interaction information to be identified in the interaction scenario;

[0012] The acquisition module is used to acquire multimodal scene recognition information; wherein, the multimodal scene recognition information is scene information associated with the interaction information to be recognized;

[0013] The processing module is used to locate the target problem hit by the interaction information to be identified based on the interaction information to be identified and the multimodal scene recognition information.

[0014] The acquisition module is used to obtain the rich media response method corresponding to the target question;

[0015] The output module is used to output the answer to the target question in a rich media responsive manner.

[0016] An electronic device, comprising: a processor, a memory, and a communication bus;

[0017] The communication bus is used to realize the communication connection between the processor and the memory;

[0018] The processor is used to execute a recognition program for multimodal interaction information stored in the memory, so as to implement the steps of the multimodal interaction information recognition method described above.

[0019] A storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the multimodal interaction information recognition method described above.

[0020] The multimodal interaction information recognition method provided in this application involves obtaining interaction information to be recognized in an interaction scenario; obtaining multimodal scene recognition information, wherein the multimodal scene recognition information is scene information associated with the interaction information to be recognized; locating the target question hit by the interaction information to be recognized based on the interaction information to be recognized and the multimodal scene recognition information, that is, achieving auxiliary location of the target question by combining the multimodal scene recognition information; furthermore, obtaining the rich media response mode corresponding to the target question, and outputting the answer to the target question in the rich media response mode. Thus, it is also possible to match the corresponding rich media response mode for the target question, achieving the purpose of flexibly matching the output mode, i.e., the answer mode, of the target question during the interaction process. Attached Figure Description

[0021] Figure 1 A flowchart illustrating the method for recognizing multimodal interaction information provided in this application embodiment. Figure 1 ;

[0022] Figure 2 A schematic diagram of an interactive scenario for clarifying a problem by calling a multimodal knowledge graph, provided in an embodiment of this application;

[0023] Figure 3 A flowchart illustrating a multimodal interaction scenario provided in an embodiment of this application;

[0024] Figure 4 A schematic diagram of the structure of the multimodal interaction information recognition device provided in the embodiments of this application;

[0025] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0026] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0027] It should be understood that the phrases "embodiments of this application" or "foreign embodiments" throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, "embodiments of this application" or "in the foreign embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0028] Traditional set-top boxes rely solely on voice information for intent interpretation during voice interaction, which can lead to an inability to accurately interpret user intent.

[0029] Furthermore, traditional set-top boxes often suffer from various challenges. Different licensees offer their own content, and each set-top box possesses unique capabilities. When users encounter constantly updated content or new features, troubleshooting becomes difficult. Accurate content analysis is also challenging when faced with various types of complaints. Moreover, users' spoken language is frequently used, making semantic analysis extremely difficult. Traditional set-top boxes require extensive user data collection for intelligent decision-making during human-computer interaction. This includes voiceprint information, image information, and electronic signature information. Only by fully collecting this type of information can accurate interpretation of user intent be achieved, leading to precise responses. Traditional set-top boxes rely solely on voice information for intent interpretation during voice interaction. However, users also consider on-screen clicks for semantic understanding. Interpreting information solely through voice fails to achieve multimodal intelligent interaction.

[0030] Currently available smart customer service systems installed on set-top boxes cannot troubleshoot problems encountered during function use; they only provide solutions based on configured knowledge base content. For content-related, new feature, and business-related issues, they cannot be updated and troubleshooted in real time. Users often use colloquial language when describing problems, making sentence parsing difficult, and similar words are hard to merge and categorize, making it impossible to accurately interpret user intent.

[0031] During human-computer interaction, set-top boxes struggle to simultaneously collect multi-dimensional information and combine it with voice for intent recognition and response. Current system solutions primarily rely on text descriptions; however, for complex issues and operations, it's more suitable to use videos, images, and text to explain the steps on a large screen. In such cases, it's necessary to categorize and analyze the questions and provide targeted rich media answers based on the specific inquiries.

[0032] The solutions provided by the system are currently mainly text-based. However, for complex issues and operations, it is more suitable to use videos, images, and text to explain the steps on a large screen. In this case, it is necessary to categorize and analyze the problems and provide targeted rich media answers and responses based on the actual inquiries.

[0033] This application provides a method for recognizing multimodal interaction information, applied to a device for recognizing multimodal interaction information, with reference to... Figure 1 As shown, the method includes the following steps:

[0034] Step 101: Obtain the interaction information to be identified in the interaction scenario.

[0035] In this embodiment, the multimodal interaction information recognition device includes, but is not limited to, middleware. Middleware is an independent system software or service program that allows distributed application software to share resources across different technologies. The middleware resides above the client server's operating system and manages computing resources and network communication. The middleware's support for multimodal interaction information recognition services can be viewed as a customer service mechanism supporting multimodal intelligent interaction.

[0036] In some embodiments, the interaction information to be identified in the interaction scenario includes, but is not limited to, questions raised by the user during human-computer interaction. In other embodiments, the interaction information to be identified in the interaction scenario includes, but is not limited to, questions raised by the user after the multimodal interaction information identification device prompts the user to complete the information during human-computer interaction.

[0037] In some interactive scenarios, human-computer interaction can be understood as the interaction between electronic devices and users. These electronic devices include, but are not limited to, smart home devices, mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) and virtual reality (VR) devices, laptops, super mobile personal computers, netbooks, and personal digital assistants (PDAs). They may also include databases, servers, and service response systems based on terminal artificial intelligence. This application does not impose any restrictions on the specific types of electronic devices.

[0038] In this application embodiment, when the electronic device is a smart home device, it includes, but is not limited to, smart home devices such as smart TVs, set-top boxes, central control platforms, speakers, and other smart home devices that provide multimedia information content.

[0039] Step 102: Obtain multimodal scene recognition information.

[0040] Among them, multimodal scene recognition information is scene information associated with the interaction information to be recognized.

[0041] In this embodiment, multimodal scene recognition information is used in the human-computer interaction process to perform intelligent analysis in combination with various types of information.

[0042] Step 103: Based on the interaction information to be identified and the multimodal scene recognition information, locate the target problem that the interaction information to be identified hits.

[0043] In this embodiment, the target problem hit by the interaction information to be identified is located based on the interaction information to be identified and the multimodal scene identification information. In other words, the interaction information to be identified is analyzed and processed in combination with the multimodal scene identification information to achieve auxiliary localization of the target problem.

[0044] Step 104: Obtain the rich media response method corresponding to the target question, and output the answer to the target question in the rich media response method.

[0045] In this application embodiment, rich media refers to information dissemination methods with animation, sound, video or interactivity; when the target problem is located, this application can match the corresponding rich media response method, i.e. output method, for different target problems, and output the answer to the target problem in the matched rich media response method, so as to achieve the purpose of flexibly matching the output method, i.e. answer method, of the target problem during the interaction process.

[0046] The multimodal interaction information recognition method provided in this application embodiment obtains the interaction information to be recognized in the interaction scenario; obtains multimodal scene recognition information, wherein the multimodal scene recognition information is scene information associated with the interaction information to be recognized; locates the target question hit by the interaction information to be recognized based on the interaction information to be recognized and the multimodal scene recognition information, that is, achieves auxiliary location of the target question by combining the multimodal scene recognition information; further, obtains the rich media response mode corresponding to the target question, and outputs the answer to the target question in the rich media response mode. In this way, the corresponding rich media response mode can also be matched for the target question, so as to achieve the purpose of flexibly matching the output mode of the target question, i.e., the answer mode, during the interaction process.

[0047] In some embodiments of this application, step 101, obtaining the interaction information to be identified in the interaction scenario, can be achieved through the following steps:

[0048] A11. Obtain the original interaction information in the interaction scenario.

[0049] A12. Call the multimodal knowledge graph to determine the attribute information of the entity associations contained in the original interaction information.

[0050] In this embodiment, parsing user-generated colloquial language in interactive scenarios is very difficult, as different words often express the same intent, and the amount of information provided by colloquial language is limited. Therefore, the question-answering engine needs to proactively consult the user to clarify the question context. To address this, this application employs a multimodal intelligent interaction approach, utilizing a multimodal knowledge graph to assist in locating user questions. The multimodal knowledge graph is used to associate entities mentioned by the user with corresponding attributes. Then, based on these attribute details, the user is consulted to guide them in completing the information, ultimately providing a clear answer.

[0051] A13. Generate and output prompt information based on attribute information.

[0052] A14. Obtain the interaction information to be identified in response to the prompt information.

[0053] The interactive information to be identified includes information supplemented to the prompt information.

[0054] This application combines a multimodal knowledge graph with a question-answering system to achieve "script tracking," or script tracing. The electronic device constructs a graph database based on the business knowledge structure of the large-screen terminal, implementing attribute-based follow-up questioning logic. When a user encounters a problem, the feedback process often doesn't contain all the problem information in a single sentence. Therefore, it's necessary to prompt the user multiple times to express their desired content, conducting "follow-up questions" from different entity domains to complete the problem information. Furthermore, multimodal information is collected based on the attribute requirements of the follow-up questions.

[0055] In a scenario where information completion is guided, when the original interactive information, such as a user's question, is unclear, it's necessary to continuously prompt the user to complete the information. This type of scenario requires extracting entities from the information provided by the user, querying corresponding attributes and relationships, and then outputting prompts to guide the user in completing the content. Analysis of weekly Q&A report data shows that on average, nearly 45% of complaints are unclear, requiring confirmation and completion of information from the user. Experimental data shows that for Q&A services, when user feedback is limited, 2-3 rounds of information completion logic can be executed. If the problem still cannot be located, similar question prompts can be provided. This completion process is strongly related to the specific business and entities involved in the human-computer interaction scenario. This application utilizes a multimodal knowledge graph to implement the completion process.

[0056] For example, if a user says, "I want to complain about the network," the system resolves the entity to "network" and queries the attributes related to the network, including broadband service issues, network speed issues, router issues, etc. After the user completes the information to "network speed issue," it queries the attributes related to network speed, including network latency, network bandwidth, etc., until a specific problem is identified.

[0057] In an interactive scenario where a multimodal knowledge graph is invoked for question clarification, the interaction flow is as follows: Figure 2 As shown:

[0058] Step 201: The electronic device calls the intelligent question-and-answer portal to obtain the original interaction information in the interaction scenario.

[0059] Here, the original interaction information includes question-and-answer text queries.

[0060] Step 202: The electronic device calls the question-answering algorithm. If a single question can be located, the answer is returned directly.

[0061] In other words, if a user's question is very specific and accurately addresses a problem, then you can directly reply with a solution.

[0062] Step 203: The electronic device parses the slots for the original interactive information.

[0063] Step 204: The electronic device calls the multimodal knowledge graph to confirm the entity based on the slot.

[0064] Step 205: The electronic device queries the graph database for the corresponding attributes based on the entity and receives the attribute results returned by the graph database. The attribute results include the attribute information associated with the entity contained in the original interaction information.

[0065] In this embodiment, different skills correspond to different slot attributes. For example, in a video scene, Zhang San is retrieved as an actor; in a music scene, Zhang San is retrieved as a singer. For text scene recognition, this application combines the scene recognition results to deduce the most likely result for the word slot, avoiding the problem of the same word corresponding to different slots.

[0066] Step 206: The electronic device returns a list of information that the user needs to complete based on the attribute results.

[0067] Step 207: The electronic device calls the intelligent question-and-answer entry point to output the first round of completion prompts. The first round of completion prompts is generated based on the list provided in step 206.

[0068] Step 208: The electronic device calls the intelligent question-and-answer entry to obtain the interactive information to be recognized in response to the feedback from the first round of completion prompts.

[0069] Understandably, this explanation uses a single completion prompt as an example. In practical applications, multiple completion prompts can be provided to ultimately obtain the interaction information to be recognized after the completion information is obtained.

[0070] Step 209: The electronic device calls the question-answering algorithm to return the answer to the interaction information to be identified.

[0071] In other words, if a user's question is unclear, it's necessary to continuously prompt the user to complete the information. In such scenarios, it's necessary to extract entities from the information provided by the user, query the corresponding attributes and relationships, and then prompt the user to complete the content. Analysis of the weekly Q&A report data shows that, on average, nearly 45% of complaints each day are unclear, requiring confirmation and completion of information from the user.

[0072] Therefore, in interactive scenarios, when intelligent question answering cannot confirm the question based on the original interactive information such as the user's statement, it calls the multimodal knowledge graph service to clarify the question, outputs prompts multiple times, guides the completion of information, and obtains the interactive information to be identified in response to the prompts. In this way, the user's intent can be interpreted more accurately.

[0073] Furthermore, in interactive scenarios, besides using multimodal intelligent interaction methods to guide the completion of information to obtain the interaction information to be identified, multimodal scene recognition information can also be combined to achieve auxiliary localization of the target problem. The following describes how to obtain multimodal scene recognition information:

[0074] In some embodiments of this application, obtaining multimodal scene recognition information in step 102 can be achieved through the following steps:

[0075] B21. Call the southbound interface to interact with the operating system of the electronic device to prompt the system services supported in the interaction scenario.

[0076] In this embodiment, the southbound interface is referred to as the southbound interface S, which is used to interact with the operating system of the electronic device to prompt the system services supported in the interaction scenario and expose the various services of the operating system to the outside world.

[0077] B22. Call the first northbound interface to interact with the voice service platform to obtain business data and / or configuration information for the interaction scenarios supported by the system service.

[0078] In this embodiment of the application, the first northbound interface is referred to as the N1 interface, which is used to interact with the voice service platform to obtain service data and / or configuration information in the interactive scenarios supported by the system service.

[0079] In some example scenarios where the N1 interface is called, for instance, the authentication platform obtains authentication information through the N1 interface, the voice service platform obtains device information and sends service data through the N1 interface, and the network management platform sends network management commands, set-top box parameters, configuration information, etc., through the N1 interface. The N1 interface in this application includes, but is not limited to, the player northbound interface, the browser northbound interface, and the terminal network management northbound interface.

[0080] It is evident that multimodal speech recognition can flexibly utilize various capabilities, including those not supported by local electronic devices, and can request the necessary capabilities from the gateway in a home LAN environment. After obtaining the required capabilities from online devices, it can perform relevant operations and obtain corresponding business data and / or configuration information.

[0081] B23. Call the second northbound interface to interact with third-party application software to obtain application data information in the interaction scenario.

[0082] The multimodal scene recognition information includes business data and / or configuration information, as well as application data information.

[0083] In this embodiment, the second northbound interface is called the N2 interface, which is used to interact with third-party application software to obtain application data information in the interaction scenario.

[0084] In some example scenarios where the N2 interface is called, interaction with third-party applications is achieved, including but not limited to the implementation of the following capabilities: playback control, page rendering, and data storage. The N2 interface in this application includes, but is not limited to, the player's northbound interface, the browser's northbound interface, and the data center's northbound interface.

[0085] As can be seen, during human-computer interaction, corresponding interfaces can be called to obtain multimodal scene recognition information in the interaction scenario. This multimodal scene recognition information includes various types of data received from the interface calls, including but not limited to text, voice, images, and visual information. Furthermore, when using multimodal scene recognition information to assist in the localization of the target problem, semantic understanding can be employed, combining various types of information for intelligent analysis. For example, when multimodal scene recognition information includes text information displayed on a large screen, it will be prioritized; when the user inputs information on other terminals, it will be combined with the semantic understanding engine, thereby combining with the interaction information to be recognized to achieve the localization of the target problem. Among these, images include, but are not limited to, facial images and / or facial information. All acquired facial images and facial information comply with legal regulations and have been explicitly disclosed to the parties involved with their consent.

[0086] As described above, this application interacts with the voice-related service platform through the N1 interface, with third-party applications through the N2 interface, and with the terminal operating system through the southbound interface S. Functionally, the middleware provides capability support to upper-layer services through the northbound interface and imposes capability requirements on the terminal operating system through the southbound interface.

[0087] In some embodiments, when a third-party licensee is added, the middleware will be integrated according to the middleware specification. The existing middleware protocol remains unchanged, and the licensee will adapt without requiring an upgrade. The middleware capability requirements include browser, media player, network management, and data-related capabilities, with specific rules mainly covering general capabilities. This middleware is responsible for interacting with the licensee. Upon receiving the content name, the backend will proactively request a testing terminal, which will then request the licensee to develop its own solution to verify resource availability.

[0088] In this embodiment, the middleware includes six main functions: browser, player, network management, settings, error code, and data center, providing unified capability support for set-top box services. Integrating this middleware allows for unified integration of third-party applications while simultaneously exposing various operating system services to the outside world.

[0089] In some embodiments of this application, step 103, which locates the target problem hit by the interaction information to be identified based on the interaction information to be identified and the multimodal scene recognition information, can be achieved through the following steps:

[0090] C31. Obtain the sample representation matrix of the entities contained in the interaction information to be identified.

[0091] Here, the sample representation matrix is ​​represented as

[0092] C32. Obtain the scene weighting parameters and scene adjustment matrix corresponding to the multimodal scene recognition information.

[0093] Here, the scene weighting parameter is represented as s. w The scene adjustment matrix is ​​represented as b w .

[0094] Here, s w It is scene-weighted, where the scene is the current interaction scenario, such as the scene displayed on the screen. If the screen is displaying a video channel, it will be weighted accordingly based on the video.

[0095] C33. By using an activation function, a nonlinear transformation is performed on the sample representation matrix, scene weighting parameters, and scene adjustment matrix to obtain the transformed sample representation features.

[0096] Here, the activation function is represented by squash(); squash() is used to... The purpose of the change is to combine the voice analysis results with the information input on the screen, so as to make a comprehensive judgment based on multi-dimensional input information.

[0097] In some embodiments, a nonlinear transformation is performed on the sample representation matrix, scene weighting parameters, and scene adjustment matrix using an activation function to obtain the transformed sample representation features. This can be achieved using the following formula (1):

[0098]

[0099] C34. Perform scene-weighted summation on the transformed sample representation features to obtain the fused scene features.

[0100] Here, to improve model performance, this application uses the average of k samples in each entity class to represent the class vector. It can be calculated using the following formula (2):

[0101]

[0102] Where, N i Let be the total number of samples in category i. Label the fused scene features as... Then we have: Here, w represents the weight value of the positive and negative scenario scores in a certain scenario, and w1 + w2 = 0. Through the above training, each entity can be labeled with a positive scenario (i.e., high probability) and a negative scenario (i.e., low probability). The k samples are selected from the total sample representations, choosing the k most prominent features for calculation, thus reducing computational complexity.

[0103] It's important to note that this discussion touches on the impact of positive and negative scenarios on user interaction discrimination. Each user exhibits bias when interacting via voice or click. Due to the inherent ambiguity of voice interaction, each interaction can be interpreted in multiple ways. This weighting is used to enhance the accuracy of user preference assessment. For example, after a user makes a voice request, the longer they remain on the current screen, the stronger their current preference, and the higher their w1 score. Conversely, if the user's dwell time is extremely short, it's considered a negative scenario, and subsequent accuracy bias decreases.

[0104] C35. Based on the fused scene features and sample representation matrix, locate the target problem.

[0105] Here, in getting In this case, after squashing with squash(), it is non-linearly mapped to the interval [0, 1], resulting in a new class vector c. i c i It can be calculated using the following formula (3):

[0106]

[0107] Furthermore, the scene vectors corresponding to the entities are obtained through multiple iterations. Finally, the target problem is located based on the scene vectors corresponding to the entities and the sample representation matrix.

[0108] In a feasible scenario, during video playback, a user can call "pause" via voice, causing the video to stop. The user can then select an item on the screen to navigate to an external product link. Furthermore, after the user calls "pause," the electronic device uploads a screenshot containing relevant entity information. The cloud then uses an enhanced Chinese pre-trained model, such as the ALBERT algorithm, to pre-train a new model, which is then synchronized to the electronic device. The next time the user's voice recognizes a similar scenario, the cloud sends the text to the device, where the model performs semantic understanding to enable human-computer interaction. In this way, the electronic device can leverage a Natural Language Processing (NLP) engine to perform targeted semantic understanding based on different user preferences and habits, executing the commands on the device with millisecond-level response times. For example, if a user asks "How much does a blue shirt cost?", and the scenario is recognized as a familiar one, the text is sent to the electronic device, where the model parses it and performs the relevant operations.

[0109] This application employs an enhanced Chinese pre-trained model, such as ALBERT, to decompose the embedding matrix using the E matrix, thereby reducing the overall embedding parameters and transforming V×H into V×E+E×H, where E is an i×j matrix corresponding to the aforementioned... Adjusting the parameters of the E matrix can help reduce the number of computational parameters required for the model.

[0110] In this embodiment, text scene recognition results are integrated into the entity recognition task. The model structure enhances the text's scene understanding capability, ensuring that recognized entities are likely to be mapped to the correct scene. This effectively solves the problem of polysemy in multiple scenarios and significantly improves the accuracy of the entity recognition task. It avoids the problem in related technologies that cannot handle polysemy and cannot accurately map recognized entities to the correct scene.

[0111] In some embodiments of this application, C35 locates the target problem based on the fused scene features and sample representation matrix, which can be achieved through the following steps:

[0112] C351. Based on the fused scene characteristics, select the corresponding scene model from the local scene composite model.

[0113] The cloud generates a matching model and sends it to the electronic device. When the model matches the target scenario, semantic recognition is performed to meet the personalized needs of different users.

[0114] It is evident that the improvement that multimodal technology brings to human-computer interaction customer service systems is also reflected in scene recognition. In this application, for frequently used user scenarios, the electronic device records scene identifiers, including images and voice, and uploads them to the cloud to generate a miniature scene composite model, which is then distributed to the electronic device. Furthermore, the electronic device can filter the local scene composite model; after a scene is matched, the filtered scene model is used for semantic recognition. In this way, semantic recognition during human-computer interaction can meet the personalized needs of different users, ensuring that the recognized entities are likely to correspond to the correct scene.

[0115] C352. Identify the sample representation matrix through the scene model to locate the target problem.

[0116] This application combines scene recognition results to reverse-engineer the most likely result of the slot, thus avoiding the problem of the same word corresponding to different slots.

[0117] In some embodiments of this application, step 104, obtaining the rich media response method corresponding to the target problem, can be achieved through the following steps:

[0118] D41. Obtain the required fill parameters, required follow-up parameters, and entity hit parameters for the entities contained in the interaction information to be identified.

[0119] D42. Obtain the complexity coefficient of the solution steps, the coefficient of the number of follow-up questions, and the coefficient of the number of inquiries corresponding to the target problem.

[0120] D43. Based on the entity's required fill parameters, required follow-up question parameters, entity hit parameters, solution step complexity coefficient, follow-up question number coefficient, and consultation volume coefficient, generate the target question's scoring result.

[0121] D44. Determine the rich media response method corresponding to the target question based on the scoring results.

[0122] The required fill parameters include the number of required fill attributes and the average number of fill attributes; the required follow-up parameters include the number of required follow-up attributes and the average number of follow-up data; and the entity hit parameters include the entity hit frequency and the average question hit frequency.

[0123] In this embodiment, different user feedback issues are comprehensively scored based on the complexity of the solution, the number of follow-up questions, and the volume of user inquiries. The rich media response of the intelligent customer service on the large screen is then set according to the scoring results. An example scoring formula is as follows:

[0124] Overall score for the problem = x × (complexity coefficient of solution steps) + y × (number of follow-up questions coefficient) + z × (number of inquiries coefficient);

[0125] Where x = number of attributes to be filled for the current entity / average number of attributes to be filled; y = number of follow-up attributes to be asked for the current entity / average number of follow-up data; z = current entity hit frequency / average question hit frequency.

[0126] For example, when the overall score is greater than 0.7, the platform suggests that a video introduction is needed, indicating that the current problem is highly complex and has a high hit rate; when the overall score is greater than 0.4, the platform suggests that a text and image introduction is needed, indicating that the current problem is of medium complexity and has a medium hit rate; when the overall score is less than 0.4, the platform suggests that a text introduction is needed, indicating that the current problem is of low complexity and has a low hit rate.

[0127] As can be seen, this application proposes an intelligent setting for the feedback mode of the target problem. In related technologies, the feedback mode is manually configured, and different electronic devices use the same feedback format, failing to present more information. This application can provide intelligent feedback by comprehensively considering factors such as problem complexity, the functional capacity of the electronic device, and user usage data.

[0128] The multimodal interaction information recognition provided in this application has a complete set of rich media rating capabilities and performs real-time message synchronization. Since user feedback issues are diverse, different issues will be comprehensively scored based on the complexity of the solution, the number of follow-up questions, and the number of user inquiries. High-scoring issues will prompt operations personnel to configure video introductions, medium-scoring issues will configure text and image introductions, and the simplest issues will be configured with text answers.

[0129] In this embodiment, for smart devices requiring voice broadcasting, in addition to preset offline voice, online voice streams accessed from the cloud are also cached. The voice cached locally has a retention period; for common phrases, the cached voice is persisted and continuously updated with on-device offline data to achieve rapid and human-like responses. This human-like voice protection feature is included in the capability evaluation indicators of QB-E-067-2018 "Internet TV Set-Top Box Terminal Technical Specification". Smart terminal devices will pre-configure some human-like voices based on their functional scope for broadcasting. Furthermore, this portion can be upgraded as firmware.

[0130] When a smart terminal receives a command requiring text-to-speech (TTS) playback, it first requests the on-device speech library. If the library is not found, it then requests the cloud for speech synthesis. After synthesis, the speech is cached locally, and subsequent calls to the same speech are recorded. If no similar speech is played within a week (configurable), the speech cache will be deleted.

[0131] After a period of iteration, smart terminals have cached the vast majority of anthropomorphic voice files on the device side, and the voice responses of smart terminals in home scenarios have become anthropomorphic.

[0132] During multimodal intelligent interaction, the device can also flexibly call upon various capabilities, including those not supported by the device itself. In a home LAN environment, it will request the necessary capabilities from the gateway and execute them after obtaining the required online devices.

[0133] In a home setting, various smart terminal devices complete online connections via Bluetooth, Wi-Fi, etc. All terminals connect to the smart home control center and report device information, including various device identifiers, capability identifiers (such as broadcasting capability, camera capability, screen display capability, voiceprint capability, etc.), and of course, status.

[0134] After the user's voice is recognized and parsed, the command is sent to the terminal. The command will indicate the capability required to execute the command. If the current smart hardware does not have the capability, it will request the smart home control to check if any connected home devices are available. If a terminal device that supports the capability is available, the command will be forwarded to the corresponding device for execution.

[0135] When a smart terminal receives ambiguous feedback from a user, it needs to clarify the user's intent. At this point, it requests the home control system to check if there are any available devices that support this capability and complete the intent confirmation process. For example, when a user buys a train ticket using a remote control, the camera is invoked for verification during the purchase confirmation process; if a signature confirmation is required, a handwriting pad is invoked for signature confirmation. This fully leverages the capabilities of various terminals to achieve multimodal intelligent interaction.

[0136] It should be noted that electronic devices also possess multimodal learning capabilities for offline voice processing. For electronic devices that need to read the answer to a target question aloud, in addition to the preset offline voice, online voice streams accessed from the cloud are also cached. The voice data cached locally on the electronic device has a storage time limit. For common phrases, the cached voice data is persisted, and the offline data on the electronic device is continuously updated.

[0137] In a human-computer interaction scenario for troubleshooting, taking a set-top box as an example, the testing process of the intelligent customer service system is as follows: After the user presses the remote control of the set-top box and says "I want to complain," they are taken to the complaint page. The user provides the complaint content on the complaint page, such as "Peppa Pig cannot be played." At this time, the set-top box recognizes the statement and uploads it to the cloud. The cloud's multimodal knowledge graph engine parses the movie title as "Peppa Pig" and then checks it on the testing platform based on the geographical location information such as the city where the set-top box is located. If the testing platform tests the content source information at that location and finds it to be normal, it prompts the user to check their home network status for troubleshooting. It should be noted that the recognized statement can be uploaded to the cloud for parsing, or it can be parsed by a multimodal knowledge graph engine on the set-top box side; this application does not specifically limit this.

[0138] In a feasible testing scenario, this application provides an intelligent customer service system that supports real-time large-screen problem troubleshooting. The system consists of five parts: a testing platform, a third-party capability platform, a central control unit, a problem rating module, and a multimodal knowledge graph engine. It can realize dynamic testing of feedback problems. Referring to the six major functions included in the middleware, as well as the aforementioned N1 and N2 interfaces, it can support dynamic fault troubleshooting without hardware modification when connected to the Magic Box.

[0139] In a feasible multimodal interaction scenario, refer to Figure 3 As shown:

[0140] Step 301: The electronic device obtains the interaction information to be identified and the multimodal scene recognition information in the interaction scenario.

[0141] For example, when a user plays content on the terminal, the terminal turns on the microphone, performs automatic speech recognition (ASR) speech recognition, and then performs NLP semantic analysis.

[0142] Step 302: The electronic device locates the entity attribution slot of the interactive information to be identified based on the multimodal scene recognition information.

[0143] Step 303: The electronic device determines whether the entity has a relation template.

[0144] Step 304: The electronic device determines that the entity has a relationship template and retrieves the entity relationship.

[0145] Step 305: The electronic device determines that the entity does not have a relationship template, does not search for entity relationships, and directly obtains the slot.

[0146] Step 306: The electronic device determines whether the sentence intent and key slots have been identified.

[0147] Step 307: The electronic device determines the sentence intent and key slots, enters the question and answer module, and asks follow-up questions based on the entity attributes.

[0148] Step 308: The electronic device determines that the sentence intent and key slots have not been identified, and performs deep recognition based on the hybrid model.

[0149] Step 309: The electronic device retrieves the answer based on the user's reply.

[0150] Step 310: The electronic device scores the problem using rich media based on its various attributes and frequency.

[0151] In the multimodal interaction process of this application, the identified intents are stored on a stack. When the user uses other modal devices, the cloud retrieves the current intent from the intent stack and then executes the corresponding operation. Through human-computer interaction on a large screen, the terminal records various user feedback, including voice and visual feedback, to capture the user's evaluation of the content. During the retrieval process, the engine extracts topics based on the user's browsing history and performs weighted operations on these topics. In ambiguous scenarios, a second query will be made. During the multimodal learning process, the terminal can also flexibly call various capabilities, including capabilities not supported by the terminal itself. In a home LAN environment, it will request the necessary capabilities from the gateway and, after obtaining a capable online device, perform the relevant execution.

[0152] In a feasible multimodal interaction process, users request content on their devices. For smart devices requiring voice playback, in addition to preset offline voice messages, online voice streams from the cloud are also cached. The cached voice messages on the local device have a retention period; for common phrases, the cached voice messages are persisted and continuously updated with offline data on the device side to achieve rapid and human-like responses. For frequently used scenarios, the device side records scene identifiers, including images and voice messages, and uploads them to the cloud to generate a miniature scene composite model. Upon matching a scene, the scene model is used for semantic recognition to meet the personalized needs of different users.

[0153] In a feasible multimodal interaction process, most functions can be implemented in the cloud. Cloud modules include: a multimodal recognition engine module, an entity management module, a conversation control module, a skills module, a call testing module, and a question rating module. Specifically, the voice information processing module converts user voice files into text information, prioritizing matching based on user-uploaded hot keywords from various fields. The multimodal recognition engine module merges or adds entities based on the user's spoken text information and performs attribute retrieval based on graph relationships. The conversation control module distributes the graph parsing results to skill domains. The cloud performs corresponding logical processing based on various skill domains. When the graph retrieval indicates that an entity requires multiple attributes, it triggers multiple rounds of interaction in the cloud to complete the information. The entity management module performs entity fusion and addition based on similarity algorithms. Problem rating module: Since user feedback issues are diverse, different issues will be comprehensively scored based on the complexity of the solution, the number of follow-up questions, and the number of user inquiries. High scores will prompt operations staff to configure video introductions, medium scores will configure text and image introductions, and the simplest level will configure text introductions.

[0154] As can be seen from the above, the multimodal interaction information recognition method provided in this application has the following beneficial effects:

[0155] (1) Parsing users' colloquial language is very difficult, as different words often express the same intention, and the amount of information provided by colloquial language is limited. The question-answering engine needs to actively consult the user to clarify the question scenario. This proposal adopts multimodal intelligent interaction technology to assist in locating user questions.

[0156] (2) During human-computer interaction, the set-top box needs to receive various types of data, including text, voice, images, and visual information. The key to this invention is how to intelligently analyze and combine various types of information during semantic understanding. For example, when text information is displayed on a large screen, it will be prioritized; when the user inputs information on other terminals, it will be incorporated into the semantic understanding engine.

[0157] (3) Innovation in solution feedback mode: The solution feedback in related technologies is all manually configured feedback mode, which cannot exhaustively cover every type of problem. The feedback form is the same across different terminals, and more information cannot be presented. Existing solutions can provide intelligent feedback by comprehensively considering factors such as problem complexity, the functional capacity of the terminal device, and user habits.

[0158] (4) The device also has multimodal learning capabilities for offline voice. For smart devices that require voice broadcasting, in addition to the preset offline voice, the voice stream called from the cloud will also be cached. The voice cached on the local device has a storage time. For common sentences, the cached voice will be persisted and the offline data on the device will be continuously updated.

[0159] (5) The improvement of the customer service system by multimodal processing is also reflected in scene recognition. For frequently used scenarios, the client will record scene identifiers, including images and voice, and upload them to the cloud to generate a miniature scene composite model. After the scene is matched, the scene model will be used for semantic recognition, so that semantic recognition can meet the personalized needs of different users.

[0160] (6) During the speech recognition process, multimodal devices can also flexibly call upon various capabilities, including capabilities that are not supported by this terminal. They will request the necessary capabilities from the gateway in the home LAN environment, and after obtaining the online devices with the required capabilities, they will perform the relevant execution.

[0161] Embodiments of this application provide a multimodal interaction information recognition device, which can be applied to... Figure 1 In a corresponding embodiment, a method for recognizing multimodal interaction information is provided, referring to... Figure 4 As shown, the multimodal interaction information recognition device 400 includes:

[0162] The module 401 is used to obtain the interaction information to be identified in the interaction scenario;

[0163] The module 401 is used to obtain multimodal scene recognition information; wherein, the multimodal scene recognition information is scene information associated with the interaction information to be recognized;

[0164] Processing module 402 is used to locate the target problem hit by the interaction information to be identified based on the interaction information to be identified and the multimodal scene recognition information;

[0165] Module 401 is used to obtain the rich media response method corresponding to the target question;

[0166] Output module 403 is used to output the answer to the target question in a rich media response manner.

[0167] In some embodiments of this application, the obtaining module 401 is used to obtain a sample representation matrix of the entities contained in the interaction information to be identified; obtain scene weighting parameters and scene adjustment matrix corresponding to the multimodal scene recognition information; perform nonlinear transformation on the sample representation matrix, scene weighting parameters and scene adjustment matrix through an activation function to obtain the transformed sample representation features; perform scene weighted summation processing on the transformed sample representation features to obtain the fused scene features; and locate the target problem based on the fused scene features and the sample representation matrix.

[0168] In some embodiments of this application, the processing module 402 is used to select a corresponding scene model from the local scene composite model based on the fused scene features; and to identify the sample representation matrix through the scene model to locate the target problem.

[0169] In some embodiments of this application, the obtaining module 401 is used to obtain the original interaction information in the interaction scenario; call the multimodal knowledge graph to determine the attribute information of the entity association contained in the original interaction information; generate prompt information according to the attribute information and output the prompt information; obtain the interaction information to be identified in response to the prompt information; wherein, the interaction information to be identified includes information supplemented to the prompt information.

[0170] In some embodiments of this application, the obtaining module 401 is used to obtain the entity filling parameters, follow-up question parameters, and entity hit parameters contained in the interactive information to be identified; obtain the solution step complexity coefficient, follow-up question number coefficient, and consultation volume coefficient corresponding to the target question; generate a scoring result for the target question based on the entity filling parameters, follow-up question parameters, entity hit parameters, solution step complexity coefficient, follow-up question number coefficient, and consultation volume coefficient; and determine the rich media response method corresponding to the target question based on the scoring result.

[0171] In some embodiments of this application, the required fill parameters include the number of required fill attributes and the average number of fill attributes; the required follow-up parameters include the number of required follow-up attributes and the average number of follow-up data; and the entity hit parameters include the entity hit frequency and the average question hit frequency.

[0172] In some embodiments of this application, the obtaining module 401 is used to call the southbound interface to interact with the operating system of the electronic device to prompt the system services supported in the interaction scenario; call the first northbound interface to interact with the voice service platform to obtain service data and / or configuration information in the interaction scenario supported by the system services; and call the second northbound interface to interact with third-party application software to obtain application data information in the interaction scenario; wherein, the multimodal scenario recognition information includes service data and / or configuration information, as well as application data information.

[0173] The multimodal interaction information recognition device provided in this application embodiment obtains interaction information to be recognized in an interaction scenario; obtains multimodal scene recognition information, wherein the multimodal scene recognition information is scene information associated with the interaction information to be recognized; locates the target question hit by the interaction information to be recognized based on the interaction information to be recognized and the multimodal scene recognition information, that is, achieves auxiliary location of the target question by combining the multimodal scene recognition information; furthermore, obtains the rich media response mode corresponding to the target question, and outputs the answer to the target question in the rich media response mode. In this way, the corresponding rich media response mode can also be matched for the target question, so as to achieve the purpose of flexibly matching the output mode of the target question, i.e., the answer mode, during the interaction process.

[0174] It should be noted that the descriptions of the same steps and contents as in other embodiments in this embodiment can be found in the descriptions in other embodiments, and will not be repeated here.

[0175] Embodiments of this application provide an electronic device that can be applied to... Figure 5 In a corresponding embodiment, a method for recognizing multimodal interaction information is provided, referring to... Figure 5 As shown, the electronic device 500 includes:

[0176] The processor 501, the memory 502, and the communication bus 503 are provided, wherein the communication bus 503 is used to realize the communication connection between the processor 501 and the memory 502.

[0177] The processor 501 is used to execute a recognition program for multimodal interaction information stored in the memory 502 to perform the following steps:

[0178] Obtain the interaction information to be identified in the interaction scenario;

[0179] Obtain multimodal scene recognition information; wherein, multimodal scene recognition information is scene information associated with the interaction information to be recognized;

[0180] Based on the interaction information to be identified and the multimodal scene recognition information, locate the target problem that the interaction information to be identified hits;

[0181] Obtain the rich media response corresponding to the target question, and output the answer to the target question in the rich media response format.

[0182] In some embodiments of this application, the processor 501 is used to execute a recognition program for multimodal interaction information stored in the memory 502 to implement the following steps:

[0183] Obtain the sample representation matrix of the entities contained in the interaction information to be identified;

[0184] Obtain the scene weighting parameters and scene adjustment matrix corresponding to the multimodal scene recognition information;

[0185] By using an activation function, a nonlinear transformation is performed on the sample representation matrix, scene weighting parameters, and scene adjustment matrix to obtain the transformed sample representation features.

[0186] The transformed sample representation features are subjected to scene-weighted summation to obtain the fused scene features;

[0187] Based on the fused scene features and sample representation matrix, the target problem is located.

[0188] In some embodiments of this application, the processor 501 is used to execute a recognition program for multimodal interaction information stored in the memory 502 to implement the following steps:

[0189] Based on the fused scene characteristics, the corresponding scene model is selected from the local scene composite model;

[0190] The target problem is located by identifying the sample representation matrix through a scene model.

[0191] In some embodiments of this application, the processor 501 is used to execute a recognition program for multimodal interaction information stored in the memory 502 to implement the following steps:

[0192] Obtain the original interaction information in the interaction scenario;

[0193] The multimodal knowledge graph is invoked to determine the attribute information of entity associations contained in the original interaction information;

[0194] Generate and output prompt messages based on attribute information;

[0195] Obtain the interaction information to be identified in response to the prompt; wherein, the interaction information to be identified includes information supplemented to the prompt.

[0196] In some embodiments of this application, the processor 501 is used to execute a recognition program for multimodal interaction information stored in the memory 502 to implement the following steps:

[0197] Obtain the required fill parameters, required follow-up parameters, and entity hit parameters of the entity contained in the interaction information to be identified;

[0198] Obtain the complexity coefficient of the solution steps, the coefficient of the number of follow-up questions, and the coefficient of the number of inquiries corresponding to the target question;

[0199] Based on the entity's required fill parameters, required follow-up question parameters, entity hit parameters, solution step complexity coefficient, follow-up question number coefficient, and consultation volume coefficient, generate the target question's scoring result.

[0200] The rich media response method corresponding to the target question is determined based on the scoring results.

[0201] In some embodiments of this application, the required fill parameters include the number of required fill attributes and the average number of fill attributes; the required follow-up parameters include the number of required follow-up attributes and the average number of follow-up data; and the entity hit parameters include the entity hit frequency and the average question hit frequency.

[0202] In some embodiments of this application, the processor 501 is used to execute a recognition program for multimodal interaction information stored in the memory 502 to implement the following steps:

[0203] It calls the southbound interface to interact with the operating system of the electronic device to indicate the system services supported in the interaction scenario;

[0204] Call the first northbound interface to interact with the voice service platform to obtain business data and / or configuration information for the interaction scenarios supported by the system service;

[0205] The system calls the second northbound interface to interact with third-party application software to obtain application data information in the interaction scenario; among which, the multimodal scenario identification information includes business data and / or configuration information, as well as application data information.

[0206] The electronic device provided in this application embodiment obtains interactive information to be identified in an interactive scenario; obtains multimodal scene recognition information; wherein, the multimodal scene recognition information is scene information associated with the interactive information to be identified; based on the interactive information to be identified and the multimodal scene recognition information, it locates the target question hit by the interactive information to be identified, that is, it achieves auxiliary location of the target question by combining the multimodal scene recognition information; furthermore, it obtains the rich media response mode corresponding to the target question, and outputs the answer to the target question in the rich media response mode. In this way, it is also possible to match the corresponding rich media response mode for the target question, so as to achieve the purpose of flexibly matching the output mode of the target question, i.e., the answer mode, during the interaction process.

[0207] It should be noted that the descriptions of the same steps and contents as in other embodiments in this embodiment can be found in the descriptions in other embodiments, and will not be repeated here.

[0208] Embodiments of this application provide a computer storage medium storing one or more programs, which can be executed by one or more processors to perform the following steps:

[0209] Obtain the interaction information to be identified in the interaction scenario;

[0210] Obtain multimodal scene recognition information; wherein, multimodal scene recognition information is scene information associated with the interaction information to be recognized;

[0211] Based on the interaction information to be identified and the multimodal scene recognition information, locate the target problem that the interaction information to be identified hits;

[0212] Obtain the rich media response corresponding to the target question, and output the answer to the target question in the rich media response format.

[0213] In some embodiments of this application, the one or more programs may be executed by one or more processors to perform the following steps:

[0214] Obtain the sample representation matrix of the entities contained in the interaction information to be identified;

[0215] Obtain the scene weighting parameters and scene adjustment matrix corresponding to the multimodal scene recognition information;

[0216] By using an activation function, a nonlinear transformation is performed on the sample representation matrix, scene weighting parameters, and scene adjustment matrix to obtain the transformed sample representation features.

[0217] The transformed sample representation features are subjected to scene-weighted summation to obtain the fused scene features;

[0218] Based on the fused scene features and sample representation matrix, the target problem is located.

[0219] In some embodiments of this application, the one or more programs may be executed by one or more processors to perform the following steps:

[0220] Based on the fused scene characteristics, the corresponding scene model is selected from the local scene composite model;

[0221] The target problem is located by identifying the sample representation matrix through a scene model.

[0222] In some embodiments of this application, the one or more programs may be executed by one or more processors to perform the following steps:

[0223] Obtain the original interaction information in the interaction scenario;

[0224] The multimodal knowledge graph is invoked to determine the attribute information of entity associations contained in the original interaction information;

[0225] Generate and output prompt messages based on attribute information;

[0226] Obtain the interaction information to be identified in response to the prompt; wherein, the interaction information to be identified includes information supplemented to the prompt.

[0227] In some embodiments of this application, the one or more programs may be executed by one or more processors to perform the following steps:

[0228] Obtain the required fill parameters, required follow-up parameters, and entity hit parameters of the entity contained in the interaction information to be identified;

[0229] Obtain the complexity coefficient of the solution steps, the coefficient of the number of follow-up questions, and the coefficient of the number of inquiries corresponding to the target question;

[0230] Based on the entity's required fill parameters, required follow-up question parameters, entity hit parameters, solution step complexity coefficient, follow-up question number coefficient, and consultation volume coefficient, generate the target question's scoring result.

[0231] The rich media response method corresponding to the target question is determined based on the scoring results.

[0232] In some embodiments of this application, the required fill parameters include the number of required fill attributes and the average number of fill attributes; the required follow-up parameters include the number of required follow-up attributes and the average number of follow-up data; and the entity hit parameters include the entity hit frequency and the average question hit frequency.

[0233] In some embodiments of this application, the one or more programs may be executed by one or more processors to perform the following steps:

[0234] It calls the southbound interface to interact with the operating system of the electronic device to indicate the system services supported in the interaction scenario;

[0235] Call the first northbound interface to interact with the voice service platform to obtain business data and / or configuration information for the interaction scenarios supported by the system service;

[0236] The system calls the second northbound interface to interact with third-party application software to obtain application data information in the interaction scenario; among which, the multimodal scenario identification information includes business data and / or configuration information, as well as application data information.

[0237] It should be noted that the descriptions of the same steps and contents as in other embodiments in this embodiment can be found in the descriptions in other embodiments, and will not be repeated here.

[0238] It should be noted that the aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; it can also be various terminals that include one or any combination of the above-mentioned memory, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0239] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0240] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0241] Furthermore, in the various embodiments of this application, all functional units can be integrated into one processing module, or each unit can be a separate unit, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in a combination of hardware and software functional units. Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0242] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0243] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0244] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0245] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for recognizing multimodal interaction information, characterized in that, The method includes: Obtain the interaction information to be identified in the interaction scenario; Obtain multimodal scene recognition information; wherein, the multimodal scene recognition information is scene information associated with the interaction information to be recognized; Based on the interaction information to be identified and the multimodal scene recognition information, locate the target problem hit by the interaction information to be identified; Obtain the rich media response format corresponding to the target question, and output the answer to the target question in the rich media response format; The step of locating the target problem hit by the interaction information to be identified based on the interaction information to be identified and the multimodal scene identification information includes: determining a sample representation matrix and fused scene features based on the interaction information to be identified and the multimodal scene identification information; selecting a corresponding scene model from a local scene composite model based on the fused scene features; and identifying the sample representation matrix through the scene model to locate the target problem.

2. The method according to claim 1, characterized in that, The step of determining the sample representation matrix and the fused scene features based on the interaction information to be identified and the multimodal scene recognition information includes: Obtain the sample representation matrix of the entities contained in the interaction information to be identified; Obtain the scene weighting parameters and scene adjustment matrix corresponding to the multimodal scene recognition information; By using an activation function, a nonlinear transformation is performed on the sample representation matrix, the scene weighting parameters, and the scene adjustment matrix to obtain the transformed sample representation features; The transformed sample representation features are subjected to scene-weighted summation to obtain the fused scene features.

3. The method according to claim 1 or 2, characterized in that, The process of obtaining the interaction information to be identified in the interaction scenario includes: Obtain the original interaction information in the aforementioned interaction scenario; The multimodal knowledge graph is invoked to determine the attribute information associated with the entities contained in the original interaction information; Generate a prompt message based on the attribute information, and output the prompt message; Obtain the interaction information to be identified in response to the prompt information; wherein, the interaction information to be identified includes information supplemented to the prompt information.

4. The method according to claim 1 or 2, characterized in that, The method for obtaining the rich media response corresponding to the target question includes: Obtain the entity fill parameters, follow-up question parameters, and entity hit parameters contained in the interaction information to be identified; Obtain the solution complexity coefficient, follow-up question number coefficient, and consultation volume coefficient corresponding to the target question; The scoring result of the target question is generated based on the entity's required filling parameters, the required follow-up question parameters, the entity's hit parameters, the solution step complexity coefficient, the follow-up question number coefficient, and the consultation volume coefficient. The rich media response method corresponding to the target question is determined based on the scoring results.

5. The method according to claim 4, characterized in that, The required fill parameters include the number of required fill attributes and the average number of fill attributes; the required follow-up question parameters include the number of required follow-up question attributes and the average number of follow-up question data; and the entity hit parameters include the entity hit frequency and the average question hit frequency.

6. The method according to claim 1 or 2, characterized in that, The acquisition of multimodal scene recognition information includes: The southbound interface is invoked to interact with the operating system of the electronic device to indicate the system services supported in the interaction scenario; The system calls the first northbound interface to interact with the voice service platform in order to obtain business data and / or configuration information under the interaction scenario supported by the system service; The system calls the second northbound interface to interact with third-party application software to obtain application data information in the interaction scenario; wherein, the multimodal scenario identification information includes the business data and / or the configuration information, as well as the application data information.

7. A device for recognizing multimodal interactive information, characterized in that, The device includes: The acquisition module is used to acquire the interaction information to be identified in the interaction scenario; The acquisition module is used to acquire multimodal scene recognition information; wherein, the multimodal scene recognition information is scene information associated with the interaction information to be recognized; The processing module is used to locate the target problem hit by the interaction information to be identified based on the interaction information to be identified and the multimodal scene recognition information; The obtaining module is used to obtain the rich media response method corresponding to the target question; The output module is used to output the answer to the target question in the rich media response mode; The obtaining module is used to determine the sample representation matrix and the fused scene features based on the interaction information to be identified and the multimodal scene recognition information. The processing module is used to select a corresponding scene model from the local scene composite model based on the fused scene features; and to identify the sample representation matrix through the scene model to locate the target problem.

8. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; A processor is configured to execute executable instructions stored in the memory to implement the method for recognizing multimodal interaction information as described in any one of claims 1 to 6.

9. A storage medium, characterized in that, The device stores executable instructions that, when executed, cause a processor to perform the method for recognizing multimodal interaction information as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Man-machine interaction method and device based on multi-modal dialogue state representation

    CN113792196A

  • Intelligent interaction method based on rich media knowledge graph multi-modal sentiment analysis model

    CN114969282A