Dialogue data processing method and device, storage medium and computer equipment
By configuring multimodal information entry controls and large-model prediction processing in the chat interaction interface, the flexibility and applicability of the intelligent dialogue system in complex scenarios is solved, and more efficient information interaction and reply accuracy is achieved.
Patent Information
- Application Number
- CN202510434328.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-25
AI Technical Summary
The existing intelligent dialogue system has low flexibility and applicability to input data in complex scenarios, and it is difficult to meet the information expression needs of professional scenarios such as medical consultation, financial services or creative design.
It provides a dialogue data processing method, which displays multimodal information entry controls, including text entry, image entry, audio entry and image drawing entry controls, receives and parses multimodal information, generates reply data, and performs prediction processing through a large model to generate multimodal reply data.
It improves the flexibility of the interactive process and information throughput, reshapes the human-computer interaction cognitive model, and ensures the accuracy and applicability of reply data, especially in professional scenarios such as medical care.
Smart Images

Figure CN120372059A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing and can be applied to the field of digital medicine. In particular, it relates to a method and device for processing dialogue data, a storage medium, and a computer device. Background Art
[0002] In the field of human-computer interaction, with the wide application of intelligent dialogue systems, users' demand for the diversity of input methods is increasing day by day. Traditional dialogue systems mainly rely on text input and are difficult to efficiently process multi-modal information expressions in complex scenarios. The limitation of the input data modality greatly restricts users' expressions.
[0003] Although some existing intelligent dialogue solutions can support the input of data in forms such as audio and images, they still cannot meet the information expression requirements in some professional scenarios, such as medical consultations, financial services, or creative designs, resulting in low flexibility and applicability of the input data in intelligent dialogue scenarios. Summary of the Invention
[0004] In view of this, the present invention provides a method and device for processing dialogue data, a storage medium, and a computer device, mainly aiming to solve the problem of low flexibility and applicability of the input data in existing intelligent dialogue scenarios.
[0005] According to one aspect of the present invention, a method for processing dialogue data is provided, including:
[0006] In response to a click operation on a dialogue initiation control in the user chat interface, display a multi-modal information input control, where the multi-modal information input control includes a text input control, an image input control, an audio input control, and an image drawing input control;
[0007] Receive dialogue input information entered based on any multi-modal information input control, where the dialogue input information includes first dialogue input information entered based on at least one of the text input control, the image input control, and the audio input control and / or second dialogue input information entered based on the image drawing input control;
[0008] Generate dialogue reply data based on the dialogue input information and render the reply data to a dialogue display area to complete a round of dialogue with the user.
[0009] Further, the generating of the dialogue reply data based on the dialogue input information includes:
[0010] Analyze the dialogue input information to obtain question data and modal composition information of the question data;
[0011] Retrieve a large model that matches the modal composition information, where the large model is trained based on multiple training samples of different modal combinations;
[0012] Perform prediction processing on the question data through the large model to generate response data corresponding to the question data.
[0013] Further, in the case where the dialogue entry information includes the first dialogue entry information and the second dialogue entry information, parsing the dialogue entry information to obtain question data and the modal composition information of the question data includes:
[0014] Calculate the dialogue entry interval duration based on the timestamps of the first dialogue entry information and the second dialogue entry information;
[0015] If the data submission interval duration is less than a preset interval duration threshold, parse the first dialogue entry information and the second dialogue entry information to obtain question data and the modal composition information of the question data;
[0016] If the data submission interval duration is greater than or equal to the preset interval duration threshold, parse the operation content entered later in the first dialogue entry information and the second dialogue entry information to obtain question data and the modal composition information of the question data.
[0017] Further, the image drawing entry control is a mind map drawing entry control, and the question data further includes an auxiliary question graph. Parsing the dialogue entry information to obtain question data includes:
[0018] Parse the first dialogue entry information to obtain core question data uploaded through the information entry control corresponding to the modality, where the core question data includes at least one of image data, audio data, and text data;
[0019] Parse the second dialogue entry information to obtain the call data of the visualization elements in the drawing area;
[0020] Generate an auxiliary question graph based on the call data, and the auxiliary question graph is used to guide the response logic to the core question data.
[0021] Further, performing prediction processing on the question data through the large model to generate response data corresponding to the question data includes:
[0022] Extract at least one keyword of the core question data based on an information extraction model that matches the modality of the core question data;
[0023] Traverse each node of the auxiliary question graph based on the keyword to obtain a target node that matches the keyword, and the association relationship between the target node and at least one associated node;
[0024] Use the association relationship between the target node and at least one associated node as reply prompt information, and together with the core question data as input, perform prediction processing through a large model that matches the modality of the core question data to obtain reply data.
[0025] Further, the image drawing input control includes a mind map drawing input control and a drawing board drawing input control, and the second dialogue input information includes mind map dialogue information and picture dialogue information;
[0026] Receive mind map dialogue information input based on the mind map drawing input control, including:
[0027] In response to a click operation on the mind map drawing input control, display a mind map drawing area and display a plurality of mind map drawing logic elements for mind map drawing;
[0028] Receive mind map dialogue information input based on the mind map drawing area and the mind map drawing components;
[0029] Receive picture dialogue information input based on the drawing board drawing input control, including:
[0030] In response to a click operation on the drawing board drawing input control, display a drawing board drawing area and display a plurality of picture drawing sticker elements for drawing board drawing;
[0031] Receive picture dialogue information input based on the drawing board drawing area and the picture drawing sticker elements.
[0032] Further, the chat interaction interface is an online diagnosis and treatment interface of an Internet hospital;
[0033] The image drawing input control is further configured with a multi-angle body schematic diagram and a local body part schematic diagram, so that the user can draw the disease distribution and disease form within the multi-angle body schematic diagram or the local body part schematic diagram;
[0034] The information input control is used to input at least one of voice dialogue data, video dialogue data, and user diagnosis and treatment material images.
[0035] According to another aspect of the present invention, there is provided a dialogue data processing device, including:
[0036] A display module, configured to display a multimodal information input control in response to a click operation on a conversation initiation control in a user chat interface, where the multimodal information input control includes a text input control, an image input control, an audio input control, and an image drawing input control;
[0037] A receiving module, configured to receive conversation input information input based on any multimodal information input control, where the conversation input information includes first conversation input information input based on at least one of the text input control, the image input control, and the audio input control and / or second conversation input information input based on the image drawing input control;
[0038] A generating module, configured to generate conversation reply data based on the conversation input information and render the reply data to a conversation display area to complete a round of conversation with the user.
[0039] Further, the generating module includes:
[0040] An analysis unit, configured to analyze the conversation input information to obtain question data and modal composition information of the question data;
[0041] An extraction unit, configured to extract a large model that matches the modal composition information, where the large model is trained based on multiple training samples of different modal combinations;
[0042] A generation unit, configured to perform prediction processing on the question data through the large model to generate reply data corresponding to the question data.
[0043] Further, in a specific application scenario, the analysis unit is specifically configured to calculate a conversation input interval duration based on timestamps of the first conversation input information and the second conversation input information when the conversation input information includes the first conversation input information and the second conversation input information;
[0044] If the data submission interval duration is less than a preset interval duration threshold, then analyze the first conversation input information and the second conversation input information to obtain question data and modal composition information of the question data;
[0045] If the data submission interval duration is greater than or equal to the preset interval duration threshold, then analyze the later-entered operation content in the first conversation input information and the second conversation input information to obtain question data and modal composition information of the question data.
[0046] Further, in a specific application scenario, the parsing unit is specifically further configured to parse the first conversation input information to obtain core question data uploaded through an information input control corresponding to a modality, where the core question data includes at least one of image data, audio data, and text data;
[0047] Parse the second conversation input information to obtain call data of visualization components in the drawing area;
[0048] Generate an auxiliary question graph based on the call data, where the auxiliary question graph is used to guide the reply logic for the core question data.
[0049] Further, in a specific application scenario, the generating unit is configured to extract at least one keyword of the core question data based on an information extraction model that matches the modality of the core question data;
[0050] Traverse each node of the auxiliary question graph based on the keyword to obtain a target node that matches the keyword, and the association relationship between the target node and at least one associated node;
[0051] Use the association relationship between the target node and at least one associated node as reply prompt information, and use it together with the core question data as input to perform prediction processing through a large model that matches the modality of the core question data to obtain reply data.
[0052] Further, the receiving module includes:
[0053] A mind map drawing area display unit, configured to respond to a click operation on a mind map drawing input control, display a mind map drawing area, and display a plurality of mind map drawing logic components for mind map drawing;
[0054] A mind map drawing operation receiving unit, configured to receive mind map conversation information input based on the mind map drawing area and the mind map drawing components;
[0055] A drawing board drawing area display unit, configured to respond to a click operation on a drawing board drawing input control, display a drawing board drawing area, and display a plurality of drawing sticker components for drawing board drawing;
[0056] A drawing board drawing operation receiving unit, configured to receive picture conversation information input based on the drawing board drawing area and the drawing sticker components.
[0057] Further, in an application scenario, the chat interaction interface is an online diagnosis and treatment interface of an Internet hospital;
[0058] The image drawing and input control is also configured with multi-angle body schematic diagrams and local body part schematic diagrams, so that the user can draw the disease distribution and disease form within the multi-angle body schematic diagram or the local body part schematic diagram;
[0059] The information input control is used to input at least one of voice conversation data, video conversation data, and user diagnosis and treatment material images.
[0060] According to another aspect of the present invention, there is provided a storage medium in which at least one executable instruction is stored, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned conversation data processing method.
[0061] According to still another aspect of the present invention, there is provided a computer device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete mutual communication through the communication bus;
[0062] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned conversation data processing method.
[0063] By means of the above technical solutions, the technical solutions provided by the embodiments of the present invention have at least the following advantages:
[0064] The present invention provides a method and device for processing conversation data. First, in response to a click operation on a conversation initiation control in a user chat interaction interface, a multi-modal information input control is displayed. The multi-modal information input control includes a text input control, an image input control, an audio input control, and an image drawing and input control; conversation input information based on any multi-modal information input control is received, where the conversation input information includes first conversation input information input based on at least one of the text input control, the image input control, and the audio input control and / or second conversation input information input based on the image drawing and input control; conversation reply data is generated according to the conversation input information, and the reply data is rendered to a conversation display area to complete a round of conversation with the user. Compared with the prior art, the embodiments of the present invention configure more-dimensional multi-modal information input channels in the chat interaction interface, which not only improves the flexibility of the interaction process and the information throughput of a single conversation, but also reshapes the cognitive mode of human-computer interaction. By following the multi-channel cognitive characteristics of humans in multi-modal interaction and allocating different modal processing models, the accuracy of generating reply data is ensured, thereby improving the applicability of conversation data processing to professional scenarios such as medical treatment.
[0065] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0067] Figure 1 shows a flowchart of a method for processing dialogue data provided by an embodiment of the present invention;
[0068] Figure 2 shows a flowchart of another method for processing dialogue data provided by an embodiment of the present invention;
[0069] Figure 3 shows a block diagram of a device for processing dialogue data provided by an embodiment of the present invention;
[0070] Figure 4 shows a schematic structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0072] An embodiment of the present invention provides a method for processing dialogue data, as Figure 1 shown, the method includes:
[0073] 101. In response to a click operation on a dialogue initiation control in the user chat interface, display a multi-modal information input control.
[0074] In an embodiment of the present invention, the current execution entity is the front-end server of the human-computer dialogue application program, which can be a local server or a cloud server. The user chat interaction interface is the front-end and visual interface of the human-computer dialogue application program. When the user needs to perform digital medical services such as artificial intelligence chat service, intelligent triage, intelligent health consultation, and psychological counseling chat, the dialogue function is activated by clicking the dialogue initiation control on the user chat interaction interface, and a multi-modal information input control is displayed on the user chat interaction interface, so that the user can input information of the corresponding modality through the corresponding control. The multi-modal information input control includes a text input control, an image input control, an audio input control, and an image drawing input control. The text input control can be a speech bubble, through which the user can input text-form information. The image input control is a selection or input control. The user can select an image stored locally through this control, or drag the image directly to the specified area of this control, or can also collect an image in real time through the camera of the terminal to realize the input of image information. Among them, the image includes pictures and videos. The audio input control is a selection or input control. The user can select an audio file stored locally through this control, or drag the audio file directly to the specified area of this control, or can also record the dialogue voice in real time through the microphone of the terminal device to realize the input of audio information. The image drawing input control is a trigger control for expanding the lower-level image drawing area. The user can draw custom content in the expanded image drawing area to express the content that needs to be input through the drawn content.
[0075] By configuring the multi-modal information input control in the chat interaction interface, the freedom of choice of the user information input method can be greatly expanded, the characteristics of the chat input information can be enriched, and thus the accuracy of the reply can be effectively improved.
[0076] 102. Receive the dialogue input information input based on any multi-modal information input control.
[0077] In the embodiments of the present invention, a user can independently select corresponding controls according to requirements to input information for conversation entry. The information for conversation entry includes first information for conversation entry entered based on at least one of a text entry control, the image entry control, and the audio entry control, and / or second information for conversation entry entered based on the image drawing entry control. The information for conversation entry may include only the first information for conversation entry, or only the second information for conversation entry, or may include both the first information for conversation entry and the second information for conversation entry. Among them, the first information for conversation entry includes one or more of image information, audio information, and text information. The second information for conversation entry is generated based on the drawing operations performed by the user in the lower-level image drawing area expanded by the image drawing entry control. For example, in the scenario of psychological counseling and chatting, the user can enter their favorite songs through the audio entry control and draw some painting content with stickers through the image drawing entry control to express their current emotional state. In the scenario of intelligent triage, the user can describe the symptoms through text entry, enter information on previous medical consultations or main bills and medical records through the image entry control, and mark in the human body schematic diagram displayed in the image drawing area the symptom information that is difficult to describe through text or photos, such as the coverage area of pain and the shape of rashes that have disappeared.
[0078] 103. Generate conversation reply data based on the information for conversation entry and render the reply data to the conversation display area to complete a round of conversation with the user.
[0079] In the embodiments of the present invention, there are many existing implementation methods for conversation reply data. It can be implemented based on a deep learning model (such as a neural network) trained with a large amount of conversation data. It can be based on retrieval-augmented generation technology, combining the retrieval of an external knowledge base with the feature extraction of a neural network to enhance the effect of conversation generation. It can also be based on a finite state machine or a memory network to extract the context information of multi-round conversations and implement the generation of conversation reply data. In summary, generating conversation reply data based on the information for conversation entry can be based on existing conversation generation technologies or a combination of multi-modal feature extraction technologies and conversation generation technologies, and the embodiments of the present invention do not make specific limitations. Among them, the modality of the generated conversation reply data may include a combination of one or more of multiple modalities such as text, audio, image, and drawn image, and the embodiments of the present invention do not make specific limitations.
[0080] It should be noted that the above conversation data processing method can be applied not only to the field of digital medicine, but also to various scenarios involving human-machine conversations such as artificial intelligence chatting and intelligent customer service in the financial field, and the embodiments of the present invention do not make specific limitations.
[0081] In one embodiment of the present invention, for further illustration and limitation, as Figure 2As shown in the figure, the steps generate dialogue reply data based on the dialogue input information, including:
[0082] 201. Analyze the dialogue input information to obtain the question data and the modal composition information of the question data.
[0083] 202. Retrieve a large model that matches the modal composition information.
[0084] 203. Perform a prediction process on the question data through the large model to generate reply data corresponding to the question data.
[0085] In the embodiment of the present invention, after receiving the dialogue input information, based on the control and classification recognition of the input dialogue input information, the data type composition included in the dialogue input information is determined, such as pictures, audio, and drawn images, and the data type composition is used as the modal composition information. And preprocess the dialogue input information, such as word segmentation, audio conversion, and image sampling, to obtain question data that can be recognized by the large model. Perform a prediction process on the question data based on the large model. The large model is pre-trained based on various training samples of different modal combinations. Among them, different modal combinations include text and drawn image combination, image and drawn image combination, audio and drawn image combination, text, image, and drawn image combination, text, audio, and drawn image combination, audio, image, and drawn image combination, text audio, image, and drawn image combination. After completing the pre-training of the large model, an association relationship is established between the modal combination and the corresponding large model, so that when the large model is retrieved, a large model with a matching modal composition can be identified from the multiple trained large models based on the modal composition information of the question data, and used to perform a prediction process on the current question data. Training the large model based on the training samples of different modal combinations can enable the large model to have the ability to fuse and process data with various modal compositions, thereby improving the learning ability of the model and the accuracy of prediction.
[0086] In an embodiment of the present invention, for further illustration and limitation, when the dialogue input information includes the first dialogue input information and the second dialogue input information, the analyzing the dialogue input information to obtain the question data and the modal composition information of the question data includes:
[0087] Calculate the dialogue input interval duration according to the timestamps of the first dialogue input information and the second dialogue input information;
[0088] If the data submission interval duration is less than the preset interval duration threshold, then analyze the first dialogue input information and the second dialogue input information to obtain the question data and the modal composition information of the question data;
[0089] If the data submission interval duration is greater than or equal to the preset interval duration threshold, then parse the operation content entered later in the first dialogue entry information and the second dialogue entry information to obtain the question data and the modal composition information of the question data.
[0090] In an embodiment of the present invention, one composition situation of the dialogue entry information is that the dialogue entry information includes both the first dialogue entry information and the second dialogue entry information. That is to say, the user has entered information through at least one of the text entry control, the image entry control, and the audio entry control, and has also entered information through the image drawing entry control. However, there is a certain operation time interval when the user enters these two pieces of dialogue entry information. For example, the user first enters a mind map through the image drawing entry control, and then enters a piece of music through the audio entry control. At this time, it is necessary to determine whether the two pieces of submitted information need to be combined as question data or used as question data separately. The above controls also generate a time stamp for marking the generation time while submitting information to the current execution entity. Based on the time stamp, the time interval between the entry of different information can be calculated. Through the analysis of the user's historical operation behavior characteristics, it can be obtained that if the interval time between two questions is short, for example, the user enters a new question within a few seconds after receiving a reply data, the possibility that the two questions are related is very high. However, if the interval between the entry of the two pieces of dialogue information is more than ten minutes, it is very likely that the two questions are not related. Therefore, a preset interval duration threshold is configured. If the dialogue entry interval duration is less than the set interval duration threshold, it indicates that the two questions need to be used as question data for prediction processing at the same time; if the dialogue entry interval duration is greater than or equal to the set interval duration threshold, it indicates that the correlation probability of the two questions is small, and only the dialogue information entered later can be used as question data. The set interval duration threshold can be configured according to the user's historical operation behavior data of the current platform, and the embodiment of the present invention does not make specific limitations.
[0091] It should be noted that since the entered content includes the image drawing content input through the image drawing entry control, which is relatively abstract, it is difficult to analyze the relevance of the information entered twice from the dimensions of keyword extraction and similarity matching of the entered information. By analyzing from the dimension of the information entry interval time, the relevance of the content can be identified from the dimension of the user's behavior.
[0092] In an embodiment of the present invention, for further explanation and limitation, parsing the dialogue entry information to obtain the question data includes:
[0093] Parse the first dialogue entry information to obtain the core question data uploaded through the information entry control of the corresponding modality;
[0094] Parse the second dialogue entry information to obtain the call data of the visualization element in the drawing area;
[0095] Generate an auxiliary question graph based on the call data.
[0096] In an embodiment of the present invention, the question data further includes an auxiliary question graph and core question data. The auxiliary question graph is used to guide the reply logic for the core question data, and the core question data includes at least one of image data, audio data, and text data. When the image drawing input control is a mind map drawing input control, the mind map drawing input control is used for the user to input the thinking logic, and at least one of the text input control, the image input control, and the audio input control is used for the user to input the question content. The auxiliary question graph includes nodes and the connection relationships between the nodes, all of which are generated by visual elements. The visual elements include node options, logic options, connection element options, etc. For example, the user draws a mind map in the drawing area through the drag-and-drop operation of the visual elements, and then constructs a graph structure according to the logical relationship between the nodes of the mind map to generate an auxiliary question graph.
[0097] In an embodiment of the present invention, for further illustration and limitation, the prediction processing of the question data by the large model to generate the reply data corresponding to the question data includes:
[0098] Extract at least one keyword of the core question data based on an information extraction model that matches the modality of the core question data;
[0099] Traverse each node of the auxiliary question graph based on the keyword to obtain the target node that matches the keyword and the association relationship between the target node and at least one associated node;
[0100] Use the association relationship between the target node and at least one associated node as the reply prompt information, and together with the core question data as the input, perform prediction processing through the large model that matches the modality of the core question data to obtain the reply data.
[0101] In the embodiments of the present invention, the information extraction model is used to extract keywords of the information characterized by the core question data. It can be a natural language neural network model such as BERT. Based on the keywords, each node of the auxiliary question graph is traversed to find the node that matches the current keyword. Then, a large model is used to make a prediction based on the association relationship between the core question data and the associated nodes to obtain the response data. The large model can be GPT-4. For example, the image keywords are "bitten furniture, dog, damage marks", and the voice keywords are "wrecking the house, correcting, training". The target node matched is the destructive behavior, and the associated nodes are "increasing exercise amount, providing toys, gradually leaving for training". Then, the response data generated is "It is recommended to consume energy by increasing outdoor activities, use a slow feeder toy to relieve anxiety, and conduct progressive separation training". Suppose the core question is "How to trim a cat's nails" (text modality). Another example, the core question is "How to trim a cat's nails". The extracted keywords are "cat, trimming nails, method, safety". The auxiliary question graph includes nodes of pet care knowledge. The target nodes matched include the method of trimming nails (matching the keyword "trimming nails"), and the associated nodes include the required tools (association relationship: need), safe posture (association relationship: requirement), and hemostasis measures (association relationship: prevention). The response data includes that the method of trimming nails requires nail clippers and styptic powder, the safe posture requires holding the cat firmly and avoiding cutting the blood line, and the hemostasis measures include pressing with styptic powder and observing the wound.
[0102] It should be noted that the user can choose to first establish a thinking framework through mind map drawing. After converting the key nodes into drawing elements, the details can be further designed in the drawing board area. The mind map drawing input control can directly draw content or import a mind map template and then modify the template content. The embodiments of the present invention do not make specific limitations. By adding the auxiliary question graph, a positive guidance can be given to the response process, so that the thinking direction of the response data more conforms to the user's expectation of obtaining the answer, thereby improving the accuracy of response generation.
[0103] In an embodiment of the present invention, for further explanation and limitation, the mind map dialogue information received based on the mind map drawing input control includes:
[0104] In response to the click operation of the mind map drawing input control, the mind map drawing area is displayed, and multiple mind map drawing logic elements for mind map drawing are displayed;
[0105] The mind map dialogue information received based on the mind map drawing area and the mind map drawing components;
[0106] The picture dialogue information received based on the drawing board drawing input control includes:
[0107] In response to a click operation on the drawing input control of the drawing board, display the drawing board area and display multiple drawing sticker elements for drawing on the drawing board;
[0108] Receive the picture conversation information input based on the drawing board area and the drawing sticker elements.
[0109] In an embodiment of the present invention, the image drawing input control includes a mind map drawing input control and a drawing board drawing input control. Correspondingly, the second conversation input information includes mind map conversation information and picture conversation information. When the user clicks on the mind map drawing input control or the drawing board drawing input control, the current execution entity responds to the operation and immediately expands the mind map drawing area, which serves as the core canvas for mind map creation. The mind map drawing area adopts an infinite canvas design, supporting the user to freely expand the drawing space by dragging. The area is built-in with intelligent alignment auxiliary lines to help the user quickly construct a structured mind map framework. The background supports multi-level grid display, which can be completely hidden to keep the interface clean and fresh, or the grid density can be adjusted as needed to assist in precise positioning. A variety of logic element libraries are integrated in the sidebar of the mind map drawing area, including: basic shape elements: standard graphics such as circles, rectangles, and rhombuses; connection line elements: various connection styles such as straight lines, curves, and broken lines; logic relation symbols: special symbol elements containing logical operators such as "AND", "OR", and "NOT"; priority markers: visual markers providing levels of urgency and importance; status indicators: dynamic indication icons such as progress status and verification status. And all elements support one-key dragging to the canvas, and visual parameters such as color, line width, and filling style can be adjusted in real time through the property panel, and an intelligent adsorption connection relationship can be established between elements.
[0110] The mind map control is arranged side by side with the drawing board drawing input control, and different colors / icons are used for function differentiation. After clicking on this control, the system switches to the picture drawing mode and displays the full-function drawing board area. A sticker element library is set in the sidebar of the drawing board area, including: scene stickers: scene elements such as indoor furniture and outdoor landscapes; character stickers: silhouette of characters in various poses; decorative elements: annotation elements such as dialog boxes, explosion symbols, and arrows; emoji: a collection of dynamic / static emoji stickers; professional diagrams: professional field-specific diagrams such as circuit symbols and biological structures. All sticker elements are in vector graphic format, supporting lossless scaling and color replacement. The user can quickly locate the required elements through keyword search, and the stickers support transformation operations such as rotation, mirroring, and transparency adjustment.
[0111] In an embodiment of the present invention, for further illustration and limitation, the chat interaction interface is the online diagnosis and treatment interface of an Internet hospital;
[0112] The image drawing input control is also configured with a multi-angle body diagram and a local body part diagram, so that the user can draw the disease distribution and disease morphology in the multi-angle body diagram or the local body part diagram;
[0113] The information entry control is used to enter at least one of voice conversation data, video conversation data, and user medical data images.
[0114] The chat interaction interface is an interface adapted to the diagnosis and treatment scenario. In the embodiment of the present invention, the chat interaction interface specifically refers to the online diagnosis and treatment interface of the Internet hospital. The interface layout is optimized according to the particularity of the medical dialogue: an electronic medical record quick viewing window is set, and the historical diagnosis and treatment records are displayed in conjunction with the dialogue display area.
[0115] The medical insurance information authentication module is integrated to complete the patient identity authentication before the conversation is initiated. The prescription quick toolbar is configured, including the classification of common drugs and the dosage calculator. The image drawing input control is enhanced for medical scenarios: Multi-angle body diagram: Provides a 3D rotatable human body model, supports front / back / side view switching, built-in bone, muscle, organ layered display function, local body part diagram: Divides the head and neck, chest and abdomen, limbs and other 8 major parts according to medical standards, and provides anatomical structure details for each part. Symptom annotation tool: Equipped with a special symptom morphology brush, supports drawing rash distribution, lump shape, pain area, etc., and provides medical standard map comparison function. Dynamic symptom demonstration: Allows patients to draw symptom development trajectories, supports timeline annotation and symptom evolution animation generation. Match and analyze the hand-drawn symptom distribution map with the medical map to generate structured medical treatment prompts. In the reply dimension, you can personalize the reply template and generate customized explanations based on the patient's age, gender, and previous medical history. And visualize the reply component to automatically generate visual elements such as symptom explanation maps and medication guidance diagrams.
[0116] The present invention provides a method for processing dialogue data. First, in response to a click operation on a dialogue initiation control in the user's chat interaction interface, a multi-modal information input control is displayed. The multi-modal information input control includes a text input control, an image input control, an audio input control, and an image drawing input control. Dialogue input information is received based on any of the multi-modal information input controls. The dialogue input information includes first dialogue input information input based on at least one of the text input control, the image input control, and the audio input control and / or second dialogue input information input based on the image drawing input control. Dialogue reply data is generated based on the dialogue input information, and the reply data is rendered to a dialogue display area to complete a round of dialogue with the user. Compared with the prior art, in the embodiments of the present invention, by configuring more-dimensional multi-modal information input channels in the chat interaction interface, not only the flexibility of the interaction process and the information throughput of a single dialogue are improved, but also the cognitive mode of human-computer interaction is reshaped. By following the multi-channel cognitive characteristics of humans through multi-modal interaction and allocating different modal processing models, the accuracy of generating reply data is ensured, thereby improving the applicability of dialogue data processing to professional scenarios such as medical treatment.
[0117] Further, as an implementation of the method described above Figure 1 The embodiments of the present invention provide a dialogue data processing device, as Figure 3 shown. The device includes:
[0118] A display module 31, configured to display a multi-modal information input control in response to a click operation on a dialogue initiation control in the user's chat interaction interface. The multi-modal information input control includes a text input control, an image input control, an audio input control, and an image drawing input control;
[0119] A receiving module 32, configured to receive dialogue input information based on any of the multi-modal information input controls. The dialogue input information includes first dialogue input information input based on at least one of the text input control, the image input control, and the audio input control and / or second dialogue input information input based on the image drawing input control;
[0120] A generating module 33, configured to generate dialogue reply data based on the dialogue input information and render the reply data to a dialogue display area to complete a round of dialogue with the user.
[0121] Further, the generating module 33 includes:
[0122] An analysis unit, configured to analyze the dialogue input information to obtain question data and modal composition information of the question data;
[0123] A retrieval unit, configured to retrieve a large model that matches the modal composition information, where the large model is trained based on multiple training samples of different modal combinations;
[0124] A generation unit, configured to perform prediction processing on the question data through the large model to generate response data corresponding to the question data.
[0125] Further, in a specific application scenario, the parsing unit is specifically configured to calculate the dialogue entry interval duration according to the timestamps of the first dialogue entry information and the second dialogue entry information when the dialogue entry information includes the first dialogue entry information and the second dialogue entry information;
[0126] If the data submission interval duration is less than a preset interval duration threshold, then parse the first dialogue entry information and the second dialogue entry information to obtain question data and the modal composition information of the question data;
[0127] If the data submission interval duration is greater than or equal to the preset interval duration threshold, then parse the operation content entered later in the first dialogue entry information and the second dialogue entry information to obtain question data and the modal composition information of the question data.
[0128] Further, in a specific application scenario, the parsing unit is specifically further configured to parse the first dialogue entry information to obtain core question data uploaded through an information entry control of the corresponding modality, where the core question data includes at least one of image data, audio data, and text data;
[0129] Parse the second dialogue entry information to obtain call data of visual elements in the drawing area;
[0130] Generate an auxiliary question graph based on the call data, where the auxiliary question graph is used to guide the response logic to the core question data.
[0131] Further, in a specific application scenario, the generation unit is configured to extract at least one keyword of the core question data based on an information extraction model that matches the modality of the core question data;
[0132] Traverse each node of the auxiliary question graph based on the keyword to obtain a target node that matches the keyword and the association relationship between the target node and at least one associated node;
[0133] Use the association relationship between the target node and at least one associated node as response prompt information, and use it together with the core question data as input, and perform prediction processing through a large model that matches the modality of the core question data to obtain response data.
[0134] Further, the receiving module includes:
[0135] A mind map drawing area display unit, configured to respond to a click operation on a mind map drawing input control, display a mind map drawing area, and display a plurality of mind map drawing logic elements for mind map drawing;
[0136] A mind map drawing operation receiving unit, configured to receive mind map conversation information input based on the mind map drawing area and the mind map drawing components;
[0137] A drawing board drawing area display unit, configured to respond to a click operation on a drawing board drawing input control, display a drawing board drawing area, and display a plurality of drawing sticker elements for drawing board drawing;
[0138] A drawing board drawing operation receiving unit, configured to receive picture conversation information input based on the drawing board drawing area and the drawing sticker elements.
[0139] Further, in an application scenario, the chat interaction interface is an online diagnosis and treatment interface of an Internet hospital;
[0140] The image drawing input control is further configured with a multi-angle body schematic diagram and a local body part schematic diagram, so that a user can draw the distribution and form of a disease within the multi-angle body schematic diagram or the local body part schematic diagram;
[0141] The information input control is configured to input at least one of voice conversation data, video conversation data, and user diagnosis and treatment material images.
[0142] The present invention provides a dialogue data processing device, which first responds to a click operation on a dialogue initiation control in a user chat interaction interface, displays a multi-modal information input control, and the multi-modal information input control includes a text input control, an image input control, an audio input control, and an image drawing input control; receives dialogue input information input based on any multi-modal information input control, wherein the dialogue input information includes first dialogue input information input based on at least one of the text input control, the image input control, and the audio input control and / or second dialogue input information input based on the image drawing input control; generates dialogue reply data according to the dialogue input information, and renders the reply data to a dialogue display area to complete a round of dialogue with the user. Compared with the prior art, the embodiment of the present invention configures more-dimensional multi-modal information input channels in the chat interaction interface, which not only improves the flexibility of the interaction process and the information throughput of a single dialogue, but also reshapes the cognitive mode of human-computer interaction. By following the multi-channel cognitive characteristics of humans in multi-modal interaction and allocating different modal processing models, the accuracy of the generated reply data is ensured, thereby improving the applicability of dialogue data processing to professional scenarios such as medical treatment.
[0143] According to an embodiment of the present invention, a storage medium is provided. The storage medium stores at least one executable instruction, and the computer executable instruction can execute the dialogue data processing method in any of the above method embodiments.
[0144] Figure 4 The structural schematic diagram of a computer device provided according to an embodiment of the present invention is shown. The specific implementation of the computer device in the specific embodiment of the present invention is not limited.
[0145] As Figure 4 shown, the computer device may include: a processor 402, a communications interface 404, a memory 406, and a communication bus 408.
[0146] Among them: The processor 402, the communications interface 404, and the memory 406 communicate with each other through the communication bus 408.
[0147] The communications interface 404 is used to communicate with network elements of other devices such as clients or other servers.
[0148] The processor 402 is used to execute the program 410, and specifically can execute the relevant steps in the above dialogue data processing method embodiment.
[0149] Specifically, the program 410 may include program code, and the program code includes computer operation instructions.
[0150] The processor 402 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the computer device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0151] The memory 406 is used to store the program 410. The memory 406 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0152] The program 410 is specifically used to cause the processor 402 to perform the following operations:
[0153] In response to a click operation on a conversation initiation control in the user chat interface, a multimodal information entry control is displayed. The multimodal information entry control includes a text entry control, an image entry control, an audio entry control, and an image drawing entry control;
[0154] Receive conversation entry information entered based on any of the multimodal information entry controls. Among them, the conversation entry information includes first conversation entry information entered based on at least one of the text entry control, the image entry control, and the audio entry control and / or second conversation entry information entered based on the image drawing entry control;
[0155] Generate conversation reply data based on the conversation entry information and render the reply data to the conversation display area to complete a round of conversation with the user
[0156] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple of them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0157] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for processing dialogue data, characterized in that, Including: In response to a click operation on a conversation initiation control in the user chat interface, a multimodal information input control is displayed, and the multimodal information input control includes a text input control, an image input control, an audio input control, and an image drawing input control; Receiving conversation input information entered based on any multimodal information input control, where the conversation input information includes first conversation input information entered based on at least one of the text input control, the image input control, and the audio input control and / or second conversation input information entered based on the image drawing input control; Generating conversation reply data based on the conversation input information and rendering the reply data to a conversation display area to complete a round of conversation with the user.
2. The method according to claim 1, wherein The generating conversation reply data based on the conversation input information includes: Parsing the conversation input information to obtain question data and modal composition information of the question data; Invoking a large model that matches the modal composition information, where the large model is trained based on multiple training samples of different modal combinations; Performing a prediction process on the question data through the large model to generate reply data corresponding to the question data.
3. The method according to claim 2, wherein In the case where the conversation input information includes the first conversation input information and the second conversation input information, the parsing the conversation input information to obtain question data and modal composition information of the question data includes: Calculating the conversation input interval duration based on the timestamps of the first conversation input information and the second conversation input information; If the data submission interval duration is less than a preset interval duration threshold, parsing the first conversation input information and the second conversation input information to obtain question data and modal composition information of the question data; If the data submission interval duration is greater than or equal to the preset interval duration threshold, parsing the operation content entered later in the first conversation input information and the second conversation input information to obtain question data and modal composition information of the question data.
4. The method according to claim 3, wherein The image drawing input control is a mind map drawing input control, and the question data further includes an auxiliary question graph. Parsing the conversation input information to obtain question data includes: Parsing the first conversation input information to obtain core question data uploaded through the information input control of the corresponding modality, where the core question data includes at least one of image data, audio data, and text data; Parsing the second conversation input information to obtain call data of visual elements in the drawing area; generating an auxiliary question graph based on the call data, and the auxiliary question graph is used to guide the reply logic for the core question data.
5. The method according to claim 4, wherein The performing a prediction process on the question data through the large model to generate reply data corresponding to the question data includes: Extracting at least one keyword of the core question data based on an information extraction model that matches the modality of the core question data; Traversing each node of the auxiliary question graph based on the keyword to obtain a target node that matches the keyword and the association relationship between the target node and at least one associated node; Use the association relationship between the target node and at least one associated node as reply prompt information, and use it together with the core question data as input, and perform prediction processing through a large model that matches the modality of the core question data to obtain reply data.
6. The method according to claim 1, wherein The image drawing input control includes a mind map drawing input control and a drawing board drawing input control, and the second conversation input information includes mind map conversation information and picture conversation information; Receiving mind map conversation information input based on the mind map drawing input control, including: In response to a click operation on the mind map drawing input control, display a mind map drawing area and display a plurality of mind map drawing logic elements for mind map drawing; Receive mind map conversation information input based on the mind map drawing area and the mind map drawing components; Receiving picture conversation information input based on the drawing board drawing input control, including: In response to a click operation on the drawing board drawing input control, display a drawing board drawing area and display a plurality of picture drawing sticker elements for drawing board drawing; Receive picture conversation information input based on the drawing board drawing area and the picture drawing sticker elements.
7. The method according to claim 1, characterized in that, The chat interaction interface is an online diagnosis and treatment interface of an Internet hospital; The image drawing input control is further configured with a multi-angle body schematic diagram and a local body part schematic diagram, so that the user can draw the disease distribution and disease form within the multi-angle body schematic diagram or the local body part schematic diagram; The information input control is used to input at least one of voice conversation data, video conversation data, and user diagnosis and treatment material images.
8. A dialogue data processing device, characterized in that, Including: A display module, configured to display a multi-modal information input control in response to a click operation on a conversation initiation control in the user chat interaction interface, where the multi-modal information input control includes a text input control, an image input control, an audio input control, and an image drawing input control; A receiving module, configured to receive conversation input information input based on any multi-modal information input control, where the conversation input information includes first conversation input information input based on at least one of the text input control, the image input control, and the audio input control and / or second conversation input information input based on the image drawing input control; A generating module, configured to generate conversation reply data according to the conversation input information, and render the reply data to a conversation display area to complete a round of conversation with the user.
9. A storage medium, in which at least one executable instruction is stored, and the executable instruction causes the processor to perform operations corresponding to the conversation data processing method according to any one of claims 1-7.
10. A computer device, comprising: A processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the conversation data processing method according to any one of claims 1-7.