Multi-modal request processing method and apparatus
Through the multimodal request processing method, receiving and processing multimodal streaming media requests is solved, and the problems of low efficiency and low understanding accuracy of traditional chat robots are achieved, achieving more efficient user interaction.
Patent Information
- Application Number
- PCT/CN2024/124711
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-10-14
- Publication Date
- 2025-06-19
AI Technical Summary
Traditional chatbots are based on multiple rounds of text communication, which is inefficient and difficult to accurately understand user problems, especially text or voice communication based on single mode.
The multimodal request processing method is adopted to receive multimodal streaming media requests (such as text, pictures, cards, voice, and video), and determine whether the request is received through the object detection method, and output the response result.
It improves the accuracy and interaction efficiency of chat robots in understanding user problems, and solves the problem of low efficiency of traditional single-modal communication.
Smart Images

Figure CN2024124711_19062025_PF_FP_ABST
Abstract
Description
Multimodal request processing method and device
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 2023117165871, filed on December 13, 2023, entitled “Multimodal Request Processing Method and Device,” which is hereby incorporated by reference in its entirety. Technical Field
[0003] The present application relates to the field of computer technology, and in particular to a multimodal request processing method and device. Background Art
[0004] With the rise of artificial intelligence, chatbots are appearing in customer service and interactive systems across various industries. However, traditional chatbots rely on multiple rounds of text communication, requiring the chatbot to complete analysis and response rounds. Furthermore, communication based solely on text or voice cannot accurately convey user questions, nor can it help the chatbot understand them. The first issue is the traditional text-based bot's turn-based regulation. In traditional bot interactions, the user enters input once, and the bot then returns an output. This not only hinders the chatbot's understanding of the user's question, but also reduces query efficiency.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a method and apparatus for processing a multimodal request.
[0008] According to one aspect of an embodiment of the present application, a multimodal request processing method is provided, including: receiving a multimodal streaming media request; obtaining the type of request action in the multimodal streaming media request, and determining a target detection method corresponding to the type of request action, wherein the target detection method is used to determine whether the multimodal streaming media request has been received; and using the target detection method to detect the multimodal streaming media request, and if the multimodal streaming media request has been received, outputting a response result of the multimodal streaming media request.
[0009] In one embodiment, the modality corresponding to the multimodal streaming media request is text, picture, card, voice, or video.
[0010] In one embodiment, determining a target detection method corresponding to the type of the requested action includes: when the type of the requested action is an instantaneous action, determining that the multimodal streaming request has been received at the moment the requested action is completed; and when the type of the requested action is a process action, determining the target detection method according to the modality corresponding to the multimodal streaming request, wherein the completion time of the instantaneous action is less than a first preset time, and the completion time of the process action is greater than the first preset time.
[0011] In one embodiment, the instantaneous actions include: clicking, sending pictures, and submitting forms; and the process actions include: text input, voice input, and video input.
[0012] In one embodiment, the target detection method is determined according to the modality corresponding to the multimodal streaming request, including: when the modality corresponding to the multimodal streaming request is text, determining the target detection method is to detect whether the target control is triggered; and when the modality corresponding to the multimodal streaming request is audio, determining the target detection method is to detect whether the received audio content is updated within a second preset time period.
[0013] In one embodiment, when the multimodal streaming request is received, a response result of the multimodal streaming request is output, including: when the type of the request action in the multimodal streaming request is a process action, the multimodal streaming request is continuously identified from the moment the request action in the multimodal streaming request starts until the multimodal streaming request is received, to obtain at least one intermediate request and one final request; generating an intermediate answer corresponding to the intermediate request and a final answer corresponding to the final request; and selecting an optimal answer from the at least one intermediate answer and the final answer as the response result.
[0014] In one embodiment, the multimodal streaming request is continuously identified starting from the request action in the multimodal streaming request until the multimodal streaming request is completely received, and at least one intermediate request and one final request are obtained, including: when the modality corresponding to the multimodal streaming request is text, determining that the request action in the multimodal streaming request is a text input action, and continuously identifying the received text information from the moment the text input action starts, to obtain at least one intermediate request and one final request; when the modality corresponding to the multimodal streaming request is audio, determining that the request action in the multimodal streaming request is a voice input action, and continuously identifying the received text information from the moment the voice input action starts, to obtain at least one intermediate request and one final request; and, when the modality corresponding to the multimodal streaming request is video, determining that the request action in the multimodal streaming request is a video input action, and continuously identifying the received text information from the moment the video input action starts, to obtain at least one intermediate request and one final request.
[0015] In one embodiment, receiving a multimodal streaming request includes: obtaining input permissions at the current moment; when the input permissions only include a target modality, receiving only a streaming request for the target modality; when the input permissions include multiple modalities, receiving streaming requests for multiple modalities, wherein the streaming request for each modality only includes one request action being executed.
[0016] In one embodiment, in the process of outputting the response result of a multimodal streaming request, the method also includes: detecting whether a new modal streaming request is received, and when a new modal streaming request is received, clearing the sending queue for outputting the response result and outputting target information, the target information is used to indicate that the output process of the response result has been interrupted; and, identifying the received new modal streaming request and outputting a new response result.
[0017] In one embodiment, the output response result includes instantaneous output and process output.
[0018] In one embodiment, the instantaneous output includes card rendering, images, or text; and the process output includes text, cards, voice broadcast, video playback, or streaming.
[0019] According to another aspect of an embodiment of the present application, a multimodal request processing device is also provided, including: an acquisition module for receiving a multimodal streaming media request; a determination module for acquiring the type of request action in the multimodal streaming media request and determining a target detection method corresponding to the type of request action, wherein the target detection method is used to determine whether the multimodal streaming media request has been received; and an output module for detecting the multimodal streaming media request using the target detection method, and outputting a response result of the multimodal streaming media request when the multimodal streaming media request has been received.
[0020] In one embodiment, the determination module includes: a determination unit, which is used to determine that the multimodal streaming request has been received and completed at the moment the request action is completed when the type of the request action is an instantaneous action; and, when the type of the request action is a process action, determine the target detection method according to the mode corresponding to the multimodal streaming request, wherein the completion time of the instantaneous action is less than the first preset time, and the completion time of the process action is greater than the first preset time.
[0021] According to another aspect of an embodiment of the present application, a non-volatile storage medium is further provided. The non-volatile storage medium includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the above-mentioned multimodal request processing method.
[0022] According to another aspect of the embodiments of the present application, a computer device is provided, including a memory and a processor, wherein the processor is configured to run a program, wherein the multimodal request processing method is executed when the program is run.
[0023] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.
[0025] FIG1 is a hardware structure block diagram of a computer terminal (or mobile device) for a multimodal request processing method according to an embodiment of the present application;
[0026] FIG2 is a flow chart of a multimodal request processing method according to the present application;
[0027] FIG3 is a schematic diagram of an optional interactive terminal structure according to the present application;
[0028] FIG4 is an optional request action input flow chart according to the present application;
[0029] FIG5 is an optional request action receiving flow chart according to the present application;
[0030] FIG6 is an optional request action screening flow chart according to the present application;
[0031] FIG7 is an optional response result output flow chart according to the present application;
[0032] FIG8 is an optional flow chart of an interactive terminal node according to the present application;
[0033] FIG9 is a timing diagram of an optional interaction process of each module in an interactive terminal according to the present application;
[0034] FIG10 is a control flow chart between modules in an optional interactive terminal according to the present application;
[0035] FIG11 is a control flow chart of various modules in another optional interactive terminal according to the present application;
[0036] FIG12 is a schematic structural diagram of an optional multimodal request processing device according to an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0038] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0039] According to an embodiment of the present application, an embodiment of a multimodal request processing method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0040] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal, a cloud server or a similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a multimodal request processing method. As shown in Figure 1, the computer terminal 10 (or mobile device 10) may include one or more (102a, 102b, ..., 102n are used in the figure to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may also include more or fewer components than those shown in Figure 1, or have a configuration different from that shown in Figure 1.
[0041] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0042] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the multimodal request processing method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned multimodal request processing method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0043] The transmission module 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission module 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0044] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0045] According to an embodiment of the present application, an embodiment of a multimodal request processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0046] First, some nouns or terms that appear in the process of explaining the embodiments of this application are subject to the following explanations:
[0047] TTS: Text To Speech, text to speech;
[0048] ASR: Automatic Speech Recognition, automatic speech recognition;
[0049] RTC: Real-time Communications, real-time communication;
[0050] IVR: Interactive Voice Response, interactive voice question and answer;
[0051] NLP: Natural Language Processing, natural language processing;
[0052] VAD: Voice Activity Detection, voice activity detection;
[0053] PCM: Pulse Code Modulation, pulse code modulation;
[0054] MMD: Multi Modal Data, multimodal data.
[0055] FIG2 is a flowchart of a multimodal request processing method according to an embodiment of the present application. As shown in FIG2 , the method includes the following steps:
[0056] Step S202: receiving a multimodal streaming media request;
[0057] Step S204: obtaining the type of the requested action in the multimodal streaming media request and determining a target detection method corresponding to the type of the requested action, wherein the target detection method is used to determine whether the multimodal streaming media request has been received;
[0058] Step S206: Detect the multimodal streaming media request using a target detection method, and output a response result of the multimodal streaming media request when the multimodal streaming media request is received successfully.
[0059] Through the above steps, it is possible to implement the following methods: receiving a multimodal streaming media request; obtaining the type of request action in the multimodal streaming media request, and determining a target detection method corresponding to the type of request action, wherein the target detection method is used to determine whether the multimodal streaming media request has been received; detecting the multimodal streaming media request using the target detection method, and outputting a response result of the multimodal streaming media request when the multimodal streaming media request has been received. By receiving the multimodal streaming media request and providing a response result when the multimodal streaming media request has been received, the purpose of using multimodal request information to understand the request information and thereby improving the accuracy of understanding the request information is achieved, thereby achieving the technical effect of improving the chat efficiency of the chat robot, and thereby solving the technical problem of low chat efficiency caused by low accuracy of chat robot understanding problems based solely on text communication.
[0060] In step S202, the multimodal streaming media request includes but is not limited to: text, pictures, cards, voice, and video. In actual application scenarios, the multimodal streaming media request is received by the interactive terminal, and the interactive terminal then detects the multimodal streaming media request and obtains a response result.
[0061] Figure 3 shows a schematic diagram of the structure of an interactive terminal. As shown in Figure 3, the interactive terminal includes: an interactive control module, which is used to process multi-modal data, determine input and output times, and perform state control; a multi-modal input conversion module, which includes multiple multi-modal processing models and is used for pre-processing and content conversion of multi-modal data; a streaming robot, which is used to stream output content based on multi-modal data; a multi-modal output processing module, which includes multiple multi-modal processing models and is responsible for content conversion and output; and a terminal, which is used to process the data sent by the server and display and play it, and obtain the local camera and microphone for recording and transmission.
[0062] The above steps S202 to S206 are described in detail below through a specific embodiment.
[0063] In step S204, the specific steps for determining the target detection method corresponding to the type of the requested action are as follows: when the type of the requested action is an instantaneous action, determining that the multimodal streaming request has been received and completed at the moment the requested action is completed; when the type of the requested action is a process action, determining the target detection method according to the mode corresponding to the multimodal streaming request, wherein the completion time of the instantaneous action is less than the first preset time, and the completion time of the process action is greater than the first preset time.
[0064] It should be noted that the first preset duration can be set according to the actual scenario. In some embodiments of the present application, an instantaneous action refers to an action whose process duration can be ignored, while a process action refers to an input whose entire input process time cannot be ignored and may produce some intermediate results (intermediate requests). For example, instantaneous actions include: clicking actions, sending pictures, and submitting forms; process actions include: text input (typing process), voice input, and video input.
[0065] In an optional manner, the state change process of the process action includes, from action input preparation (ready) to action execution (doing) to action completion (done). Figure 4 shows an action input process, including: text stream and audio stream input, wherein the text stream includes: text, card UI event input; the audio stream includes: speaking action input, visual input.
[0066] In some embodiments of the present application, when the multimodal streaming request is received, the response result of the multimodal streaming request is output. Specifically, when the type of the request action in the multimodal streaming request is a process action, the multimodal streaming request is continuously identified from the moment the request action in the multimodal streaming request starts until the multimodal streaming request is received, and at least one intermediate request and one final request are obtained; an intermediate answer corresponding to the intermediate request and a final answer corresponding to the final request are generated; and an optimal answer is selected from the at least one intermediate answer and the final answer as the response result.
[0067] Among them, when the modality corresponding to the multimodal streaming media request is text, the request action in the multimodal streaming media request is determined to be a text input action, and the received text information is continuously recognized from the moment the text input action starts, and at least one intermediate request and one final request are obtained; when the modality corresponding to the multimodal streaming media request is audio, the request action in the multimodal streaming media request is determined to be a voice input action, and the received text information is continuously recognized from the moment the voice input action starts, and at least one intermediate request and one final request are obtained; when the modality corresponding to the multimodal streaming media request is video, the request action in the multimodal streaming media request is determined to be a video input action, and the received text information is continuously recognized from the moment the video input action starts, and at least one intermediate request and one final request are obtained.
[0068] Taking the interactive terminal shown in Figure 3 as an example, the process of the interactive terminal receiving a multimodal streaming media request through a streaming robot is shown in Figure 5. If the current configuration assumes that only an instantaneous action needs to be waited for, the final request can be sent directly to the robot when this action occurs, for example: sending a card UI (User Interface). For process actions, the interactive control module needs to determine the starting point, process, and end point of the action. For example: in voice input, the information of VAD and ASR can be used to determine whether the audio streaming media request in the multimodal mode has been received (whether the user has finished speaking). If it is action detection, the action detection model can be used to determine the start and end of the user action. In the process of receiving a text streaming media request in the multimodal mode (while the user is typing), as long as there is a content update within the second preset time length, it is an intermediate request, and clicking send is the final request.
[0069] The VAD continuously detects audio input (whether the user is speaking), and the ASR continuously translates the audio. When the VAD+ASR detects meaningful audio (audio that can be recognized as text), the user is considered to have started inputting. If multiple modal inputs are being sent simultaneously, the interaction control module must determine when input has ended in all modalities.
[0070] Due to time differences, not all input content has been received when the user input ends. For example, when the audio stream transmission ends (the user's voice ends), the ASR text results have not yet been received. Therefore, the interaction control module needs to maintain the request results and decide whether to wait for all conversion results. This will delay the time when the request is received.
[0071] In some embodiments of the present application, the process of receiving multimodal streaming media requests through the streaming robot requires permission. It can be understood that the moment when the streaming robot receives the multimodal streaming media request is controlled by the interaction control module in order to ensure that audio, video and other data are synchronized in the multimodal interaction.
[0072] In an optional manner, the process of receiving multimodal streaming media requests through a streaming robot is as follows: obtaining input permissions at the current moment; when the input permissions only include the target modality, only receiving streaming media requests of the target modality; when the input permissions include multiple modalities, receiving streaming media requests of multiple modalities, wherein the streaming media request of each modality only includes one request action being executed.
[0073] It is understandable that the multimodal streaming media request received by the interactive terminal is ultimately received by the streaming robot in the interactive terminal and gives a response result.
[0074] As shown in Figure 6, in actual application scenarios, the streaming robot in the interactive terminal does not always allow input in all modalities. The robot needs to inform the interaction control module of the input modalities allowed for the current node. For example, if the current terminal page is a card with only a few selectable buttons, the streaming robot may not accept multimodal streaming requests corresponding to audio and video modes and only allow clicks on the buttons of the current card. After receiving a multimodal streaming request in the audio mode, the interaction control module does not forward it to the streaming robot, or the streaming robot does not process it after receiving it.
[0075] After receiving the intermediate request, it can be cached in the multimodal control module or in the robot's context, which can be adjusted according to the data volume and transmission conditions.
[0076] In some embodiments of the present application, it is possible to roll back to the previous node during the request identification process. For example, when it is determined that the user's input in this round is a supplement to the previous round, even if it has advanced to the next node (outputting the response result), it is necessary to roll back to the previous node / state and receive the request input again.
[0077] It is understood that the intermediate request can be the recognition result of text input during the text input process, such as the recognition result during the typing process (before sending), the result of audio recognition during audio input, or the result of video recognition during video input. The final request is the recognition result after the multimodal streaming request is received. Traditional robots interactively receive input once and then return output. During the input process, the robot is actually in an idle state and only starts thinking after the input is completed. The thinking time largely determines the waiting time. The method proposed in this application is during the process of receiving requests.
[0078] In actual application scenarios, the output response results also include two types: instantaneous output and process output. The action types of instantaneous output and process output are similar to the types of request actions, as shown in Figure 7. For example, instantaneous output includes card rendering, pictures, or text; process output includes text, cards, voice broadcast, video playback, or streaming.
[0079] In an optional manner, whether it can be interrupted can be divided through configuration, where the interruptible output is process output and the non-interruptible output is instantaneous output.
[0080] The specific process is as follows: detect whether a new modal streaming media request is received. If a new modal streaming media request is received, clear the sending queue used to output the response result and output the target information. The target information is used to indicate that the output process of the response result has been interrupted; identify the new modal streaming media request received and output the new response result.
[0081] In actual application scenarios, users often interrupt the robot during an interaction. This is equivalent to a new request being received while the interactive terminal is still delivering content. This interruption can be achieved by clearing the delivery queue to interrupt audio and video playback, while also pushing some connecting words or playing interruption effects.
[0082] It should be noted that, since there is only one stream of output voice and video, the control module needs to maintain which data belongs to the interrupted data, otherwise the interruption may clear out some normal data.
[0083] The terminal can integrate streaming media capabilities through the SDK, access local cameras and microphones for recording, data encoding and push, obtain data, and correctly render and play.
[0084] FIG8 shows a schematic diagram of an interactive terminal node.
[0085] Figure 9 shows a sequential diagram of the interaction process for each module in the interactive terminal. As shown in Figure 9, first, the robot announces the initial welcome message and displays some additional information, such as images or videos. Second, the user begins describing their problem and may upload images, open a camera, or type to assist in describing their problem. During the user input process, the multimodal interaction control module needs to detect the user's status: whether they are not inputting, are in the process of inputting, or have completed input. Third, audio and video streams are continuously input, and the input conversion module also continuously converts information, such as converting speech to questions, extracting video features, and extracting image information. Fourth, a continuous stream of intermediate information is transmitted to the robot, giving it time to think during the user input process. The robot can be viewed as a directed graph, which may have cycles. The streaming information processing is equivalent to requiring all nodes to add cycles pointing to themselves. When the input is marked with an intermediate request, the node's next processing node is still itself; fifth, each time the robot processes an intermediate request, it will generate many intermediate results. As the input becomes richer, the robot continuously corrects its own answers. When the robot receives a request marked as the final input of this round, it needs to determine whether it is necessary to perform further calculations compared to the final request and the previous ones. If there is no need to perform further calculations, the last result will be thrown to the output processing module, and the output processing module will then render the output into various formats.
[0086] Figures 10 and 11 show a control flow chart of a multimodal streaming request between various modules in an interactive terminal. As shown in Figures 10 and 11, a) listens for streaming events and sends streaming media to input processing; b) listens for input processing and detects whether there is an input action, and if so, puts it into the input action list; c) input action list, each mode can currently only have one doing input action, and can have a series of done actions; d) check thread, sees which input actions have been generated, and decides whether to call the robot based on rules; e) robot request thread, requests the robot and caches the robot results; f) check thread, checks whether there is a final result of the robot, and if so, generates an output action; g) output action queue, each mode currently has only one action being output, but there can be actions waiting, and if interruptible, they can be interrupted by other actions; h) output action processing thread, based on the output action, calls the output conversion module, listens for events of the output conversion module, and when the output event callback is received, pushes the stream or sends data to the streaming service; i) critical section, the critical section of each thread consists of the input action list and the output action list.
[0087] It should be noted that the check thread refers to the check thread.
[0088] The multimodal request processing method provided in an embodiment of the present application is also applied to a multimodal request processing device provided in an embodiment of the present application, as shown in Figure 12, including: an acquisition module 110, used to receive a multimodal streaming media request; a determination module 112, used to obtain the type of request action in the multimodal streaming media request and determine the target detection method corresponding to the type of request action, wherein the target detection method is used to determine whether the multimodal streaming media request has been received; an output module 114, used to detect the multimodal streaming media request using the target detection method, and output a response result of the multimodal streaming media request when the multimodal streaming media request has been received.
[0089] The determination module 112 includes: a determination unit, which is used to determine that the multimodal streaming request has been received and completed at the moment the request action is completed when the type of the requested action is an instantaneous action; and, when the type of the requested action is a process action, determine the target detection method according to the mode corresponding to the multimodal streaming request, wherein the completion time of the instantaneous action is less than the first preset time, and the completion time of the process action is greater than the first preset time.
[0090] The determination unit is also used to determine that the target detection method is to detect whether the target control is triggered when the mode corresponding to the multimodal streaming request is text; and to determine that the target detection method is to detect whether the received audio content is updated within a second preset time length when the mode corresponding to the multimodal streaming request is audio.
[0091] The output module 114 includes: an output submodule, which is used to continuously identify the multimodal streaming media request from the moment the request action in the multimodal streaming media request starts until the multimodal streaming media request is received, and obtain at least one intermediate request and one final request; generate an intermediate answer corresponding to the intermediate request and a final answer corresponding to the final request; and select an optimal answer from the at least one intermediate answer and the final answer as a response result.
[0092] The output submodule includes: an output unit, which is used to determine that the request action in the multimodal streaming media request is a text input action when the mode corresponding to the multimodal streaming media request is text, and continuously identify the received text information from the moment the text input action starts, to obtain at least one intermediate request and one final request; when the mode corresponding to the multimodal streaming media request is audio, the request action in the multimodal streaming media request is determined to be a voice input action, and continuously identify the received text information from the moment the voice input action starts, to obtain at least one intermediate request and one final request; when the mode corresponding to the multimodal streaming media request is video, the request action in the multimodal streaming media request is determined to be a video input action, and continuously identify the received text information from the moment the video input action starts, to obtain at least one intermediate request and one final request.
[0093] The acquisition module 110 includes: a receiving submodule for obtaining the input permission at the current moment; when the input permission only includes the target modality, only the streaming media request of the target modality is received; when the input permission includes multiple modalities, the streaming media request of multiple modalities is received, wherein the streaming media request of each modality only includes one request action being executed.
[0094] The output module 114 also includes: a detection submodule, which is used to detect whether a new modal streaming media request is received. When a new modal streaming media request is received, the sending queue for outputting the response result is cleared, and target information is output. The target information is used to indicate that the output process of the response result has been interrupted; the new modal streaming media request received is identified, and a new response result is output.
[0095] According to another aspect of an embodiment of the present application, a non-volatile storage medium is further provided, including a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the above-mentioned multimodal request processing method.
[0096] According to another aspect of an embodiment of the present application, a computer device is further provided, including a memory and a processor, wherein the processor is configured to run a program, wherein the multimodal request processing method is executed when the program is run.
[0097] The above-mentioned computer device executes the above-mentioned multimodal request processing method, which adopts the method of receiving a multimodal streaming media request; obtaining the type of request action in the multimodal streaming media request, and determining the target detection method corresponding to the type of request action, wherein the target detection method is used to determine whether the multimodal streaming media request is received; using the target detection method to detect the multimodal streaming media request, and outputting the response result of the multimodal streaming media request when the multimodal streaming media request is received. By receiving the multimodal streaming media request and giving the response result when the multimodal streaming media request is received, the purpose of using multimodal request information to understand the request information and thus improving the accuracy of understanding the request information is achieved, thereby achieving the technical effect of improving the chat efficiency of the chat robot, and thus solving the technical problem of low chat efficiency caused by low accuracy of chat robot understanding problems based on text communication alone.
[0098] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0099] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0100] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0101] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.
[0102] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0103] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program code.
[0104] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0105] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A multimodal request processing method, comprising: receiving multimodal streaming media requests; Obtaining a type of a request action in the multimodal streaming media request, and determining a target detection method corresponding to the type of the request action, wherein the target detection method is used to determine whether the multimodal streaming media request is received and completed; as well as The multimodal streaming media request is detected using the target detection method, and when the multimodal streaming media request is received completely, a response result of the multimodal streaming media request is output.
2. The method according to claim 1, wherein: The modality corresponding to the multimodal streaming media request is text, picture, card, voice, or video.
3. The method according to claim 1, wherein: The determining of the target detection method corresponding to the type of the requested action includes: In the case where the type of the request action is an instantaneous action, determining that the multimodal streaming media request has been received when the request action is completed; and In the case where the type of the requested action is a process action, the target detection method is determined according to the modality corresponding to the multimodal streaming media request, wherein the completion time of the instantaneous action is less than a first preset time, and the completion time of the process action is greater than the first preset time.
4. The method according to claim 1, wherein: The instantaneous actions include: clicking, sending pictures and submitting forms; the process actions include: text input, voice input and video input.
5. The method according to claim 3, wherein: The determining the target detection method according to the mode corresponding to the multimodal streaming media request includes: In a case where the mode corresponding to the multimodal streaming media request is text, determining that the target detection method is to detect whether a target control is triggered; and In a case where the modality corresponding to the multimodal streaming media request is audio, the target detection method is determined to be detecting whether the received audio content is updated within a second preset time period.
6. The method according to claim 1, wherein: When the multimodal streaming media request is received, outputting a response result of the multimodal streaming media request includes: In the case where the type of the request action in the multimodal streaming media request is a process action, continuously identifying the multimodal streaming media request from the moment when the request action in the multimodal streaming media request starts until the multimodal streaming media request is completely received, to obtain at least one intermediate request and one final request; generating an intermediate answer corresponding to the intermediate request and a final answer corresponding to the final request; and An optimal answer is selected from at least one of the intermediate answers and the final answer as the response result.
7. The method according to claim 6, wherein: The multimodal streaming media request is continuously identified from the request action in the multimodal streaming media request until the multimodal streaming media request is completely received, and at least one intermediate request and one final request are obtained, including: In a case where the modality corresponding to the multimodal streaming media request is text, determining that the request action in the multimodal streaming media request is a text input action, and continuously identifying the received text information from the moment when the text input action starts, to obtain at least one intermediate request and one final request; In a case where the mode corresponding to the multimodal streaming media request is audio, determining that the request action in the multimodal streaming media request is a voice input action, and continuously recognizing the received text information from the moment when the voice input action starts, to obtain at least one intermediate request and one final request; and In the case where the mode corresponding to the multimodal streaming media request is video, the request action in the multimodal streaming media request is determined to be a video input action, and the received text information is continuously identified from the moment the video input action starts to obtain at least one intermediate request and one final request.
8. The method according to claim 1, wherein: Receive multi-modal streaming requests, including: Get the input permission at the current moment; In the case where the input permission only includes the target modality, only receiving a streaming media request of the target modality; and In the case where the input permission includes multiple modalities, streaming media requests of the multiple modalities are received, wherein the streaming media request of each modality includes only one request action being executed.
9. The method according to claim 1, wherein: In the process of outputting the response result of the multimodal streaming media request, the method further includes: Detecting whether a new modal streaming media request is received, and in the case of receiving a new modal streaming media request, clearing a delivery queue for outputting a response result, and outputting target information, wherein the target information is used to indicate that an output process of the response result has been interrupted; and The received new modal streaming media request is identified and a new response result is output.
10. The method according to claim 1, wherein: The output response results include instantaneous output and process output.
11. The method according to claim 10, wherein: The instantaneous output includes card rendering, pictures, or text; and The process output includes text, cards, voice broadcast, video playback or streaming.
12. A multimodal request processing device, comprising: An acquisition module, used for receiving multimodal streaming media requests; A determination module, used to obtain the type of the request action in the multimodal streaming media request, and determine the target detection method corresponding to the type of the request action, wherein the target detection method is used to determine whether the multimodal streaming media request is received; as well as The output module is used to detect the multimodal streaming media request by adopting the target detection method, and output a response result of the multimodal streaming media request when the multimodal streaming media request is received completely.
13. The multimodal request processing device according to claim 12, wherein: The determination module comprises: A determination unit is used to determine, when the type of the request action is an instantaneous action, that the multimodal streaming media request has been received and completed at the moment when the request action is completed; and, when the type of the request action is a process action, determine the target detection method according to the mode corresponding to the multimodal streaming media request, wherein the completion time of the instantaneous action is less than a first preset time, and the completion time of the process action is greater than the first preset time.
14. A non-volatile storage medium, the non-volatile storage medium comprising a stored program, wherein: When the program is running, the device where the non-volatile storage medium is located is controlled to execute the multimodal request processing method described in any one of claims 1 to 11.
15. A computer device comprising a memory and a processor, wherein the processor is used to run a program, wherein: When the program is running, the multimodal request processing method described in any one of claims 1 to 11 is executed.
Citation Information
Patent Citations
Intelligent automated assistant in a media environment
CN107003797A
Information processing method and device and storage medium
CN114201102A
Virtual character driving method, system and equipment based on multi-modal data
CN114840090A
Multi-mode request processing method and device
CN117668196A
Method, system and module for mult-modal data fusion
US20040093215A1