Response information generation method and device, storage medium and electronic equipment
By extracting and converting semantic features of multimodal input files, the problem of low efficiency in generating reply information in question-answering systems is solved, and accurate understanding of user query intentions and efficient generation of reply information are achieved.
Patent Information
- Application Number
- CN202510685272.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-12
AI Technical Summary
Existing question-answering systems rely on specific prompt words when processing complex queries and are unable to accurately understand the user's query intent, resulting in low efficiency in generating response information.
By receiving multimodal input files, extracting semantic features, converting information query intent, and generating reply information based on preset business information, multimodal feature fusion and semantic understanding technology are used to improve the accuracy of information query intent.
It realizes the perception of users' real query needs, improves the efficiency and matching degree of response information generation, and ensures the accuracy of information query intentions.
Smart Images

Figure CN120633834A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method and device for generating reply information, a storage medium, and an electronic device. Background Art
[0002] With the widespread adoption and development of artificial intelligence (AI) technology, it has been gradually applied to question-answering systems. Users enter questions into the system, and the system provides responses based on the user's input. Current question-answering systems employ a prompt-based response approach. This approach requires users to trigger the system to perform a corresponding action using specific prompts. Users embed their query requirements into the corresponding prompt structure according to the prompt format, thereby constructing a corresponding query. The system then identifies the prompts in the query and understands the user's query intent only when the prompts match the query. This response approach relies heavily on prompts. If the pre-set prompts are not matched, the system cannot perceive the user's true query intent. Complex queries require more nuanced input files in different modalities, such as text input describing the query requirements and event screenshots describing the status of the event to which the query belongs. In these cases, it is difficult for the system to accurately construct prompts to understand the user's query intent, which in turn affects the accuracy of the response information it outputs. Summary of the Invention
[0003] The present application provides a method and device for generating reply information, a storage medium and an electronic device, so as to at least solve the problem of low efficiency in generating reply information in related technologies.
[0004] The present application provides a method for generating reply information, comprising: receiving a target query request input by a target account when requesting to query information within a target business field, wherein the target query request includes a multimodal input file, and the multimodal input file is used to characterize the information query conditions of the target query request from multiple dimensions; performing feature extraction on the multimodal input file to obtain semantic features of the input file of each modality, wherein the semantic features are used to indicate the information query requirements described by the corresponding input file; converting the current information query intention of the target query request according to the multiple semantic features; generating target reply information of the target query request according to the information query intention and preset business information of the target business field; and sending the target reply information to the target account.
[0005] The present application also provides a reply information generation device, including: a receiving module, used to receive a target query request input by a target account when requesting to query information within a target business field, wherein the target query request includes a multimodal input file, and the multimodal input file is used to characterize the information query conditions of the target query request from multiple dimensions; an extraction module, used to perform feature extraction on the multimodal input file to obtain semantic features of the input file of each modality, wherein the semantic features are used to indicate the information query requirements described by the corresponding input file; a conversion module, used to convert the current information query intention of the target query request according to multiple semantic features; a generation module, used to generate target reply information of the target query request according to the information query intention and preset business information of the target business field; and a sending module, used to send the target reply information to the target account.
[0006] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned methods for generating reply information when executing the computer program.
[0007] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned methods for generating reply information are implemented.
[0008] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned methods for generating reply information when the computer program is executed by a processor.
[0009] Through this application, the target account includes a multimodal input file in the input target query request, and the multimodal input file carries the user's real query needs. Then, by extracting features from the multimodal input file, the semantic features of each modal input file are obtained, thereby realizing the perception of the information query needs described by the input file, and then converting the information query intention indicated by the target query request through multiple semantic features, realizing the conversion of the real information query intention of the target query request through the semantic understanding of the multimodal input file, ensuring the accuracy of the information query intention, and then generating target reply information based on the more detailed query intention and the preset business information of the target business field, thereby improving the matching degree between the target reply information and the target query request. Therefore, it can solve the technical problem of low efficiency in generating reply information in related technologies, and achieve the technical effect of improving the efficiency in generating reply information. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0011] Figure 1 This is a hardware structure diagram of the method for generating reply information according to an embodiment of the present application;
[0012] Figure 2 is a flowchart of a method for generating reply information according to an embodiment of the present application;
[0013] Figure 3 This is a diagram of the question-answering system architecture according to an embodiment of the present application;
[0014] Figure 4 This is a structural block diagram of the generation of reply information according to an embodiment of the present application. DETAILED DESCRIPTION
[0015] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0016] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0017] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0018] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the reply information generation method depends, the specific application environment architecture or specific hardware architecture is described here.
[0019] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1This is a hardware structure diagram of the method for generating reply information according to an embodiment of the present application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data, wherein the above-mentioned server device may also include a transmission device 106 for communication functions and an input and output device 108. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0020] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the startup method of the operating system in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the server device via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0021] The transmission device 106 is used to receive or send data via a network. A specific example of the aforementioned network may include a wireless network provided by a communication provider of the server device. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0022] An embodiment of the present application provides a method for generating reply information, and the method is described in detail in conjunction with the execution flow of the method for generating reply information.
[0023] In this embodiment, a method for generating reply information is provided. Figure 2 is a flow chart of a method for generating reply information according to an embodiment of the present application, such as Figure 2As shown, the method includes the following steps:
[0024] Step S202: receiving a target query request input by a target account when requesting to query information within a target business domain, wherein the target query request includes a multimodal input file, and the multimodal input file is used to represent the information query conditions of the target query request from multiple dimensions;
[0025] Step S204: extracting features from the multimodal input files to obtain semantic features of the input files in each modality, wherein the semantic features are used to indicate information query requirements described by the corresponding input files;
[0026] Step S206, converting the current information query intent of the target query request according to the plurality of semantic features;
[0027] Step S208: generating target response information for the target query request according to the information query intention and the preset business information of the target business field;
[0028] Step S210: Send the target reply information to the target account.
[0029] Through the above steps, the target account includes a multimodal input file in the input target query request, and the multimodal input file carries the user's real query needs. Then, by extracting features from the multimodal input file, the semantic features of each modal input file are obtained, and the perception of the information query needs described by the input file is realized. Then, the information query intention indicated by the target query request is converted from multiple semantic features, and the real information query intention of the target query request is converted by semantic understanding of the multimodal input file, thereby ensuring the accuracy of the information query intention. Then, the target reply information is generated according to the information query intention and the preset business information of the target business field, thereby improving the matching degree between the target reply information and the target query request. Therefore, the technical problem of low efficiency in generating reply information in related technologies can be solved, and the technical effect of improving the efficiency in generating reply information can be achieved.
[0030] The above-mentioned reply information generation method can be applied to, but is not limited to, an information question-and-answer system. The target account inputs a target query request including a multimodal input file on the question-and-answer system, and the question-and-answer system generates target reply information for the target query request by running the above-mentioned reply information generation method.
[0031] In the embodiment provided in step S202, the modality of the input file is used to indicate the carrier of the input file. The file carriers of input files of different modalities are different. Multimodal input files can be, but are not limited to, files in the form of text, audio, video, image, sensor data, etc.
[0032] In the embodiments of the present application, users have increasingly higher demands for responses from the question-and-answer system and are no longer satisfied with the traditional single-answer retrieval response method. The questions input by users are more professional, the question description information required to describe the questions to be answered is more loaded, the description information involves more information dimensions, and the description information may be carried by files of multiple different modalities. For example, when requesting the question-and-answer system to perform fault diagnosis based on the operating status of the server, the multimodal input files that need to be input may include but are not limited to system log screenshots, data collected by sensors, multimedia images of the system operating status, voice instructions input by users, text instructions, etc. This solution does not limit this.
[0033] In the embodiment provided in step S204, the purpose of the feature extraction operation on the input file is to perceive the query requirements described in the corresponding input file. Since the file carriers of input files of different modalities are different, the feature extraction is performed using the feature extraction method corresponding to the modality to which the input file belongs. For example, when the input file is a text file, the text is processed to remove special characters or stop words in the text, and then the processed text is subjected to semantic recognition and feature extraction to obtain the semantic features of the text. When the input file is a multimedia image file, the image feature extraction is performed on the image to obtain the semantic features corresponding to the image. When the input file is an audio file, the audio file is converted into text to obtain the text file corresponding to the audio file, and the text file is subjected to semantic recognition and feature extraction to obtain the semantic features of the text.
[0034] In the embodiment provided in step S206, the information query intent can be converted by, but is not limited to, aligning and fusing multiple semantic features. The purpose of aligning and fusing multiple semantic features is to map semantic features of different modalities to a unified feature space, and to represent the entire query intent of the target account in the target query request through the fused information query intent.
[0035] Optionally, in an embodiment of the present application, the information query intent can be, but is not limited to, obtained by identifying and converting multiple semantic features through an intent conversion model, wherein the intent conversion model records the conversion relationship between semantic features and query intent, and by inputting multiple semantic features into the intent conversion model, the content output by the intent conversion model is used as the information query intent. This solution does not limit this.
[0036] In the embodiment provided in step S208, the preset business information is business information configured within the target business domain, and the preset business information is used to provide a basis for replying to query requests within the target business domain, that is, the target reply information is generated based on the information content included in the preset business information. The method of generating the target reply information based on the information query intent and the preset business information may be, but is not limited to, querying candidate business information related to the information query intent from the preset business information based on the information query intent, and generating corresponding target reply information for the target query request in accordance with the reply method indicated by the information query intent using the candidate business information as information basis.
[0037] Optionally, in an embodiment of the present application, in order to improve the query efficiency of business information, the preset business information of the target business field can be stored in the form of a knowledge graph library, wherein the knowledge graph library includes multiple entity points and connection edges connecting the entity points, wherein the multiple entity points correspond one-to-one to the sub-businesses included in the target business field, and each entity point is used to represent the sub-business information of the corresponding sub-business. The connection edge is used to indicate the transfer relationship between the sub-businesses corresponding to the entity point. The information query intention may, but is not limited to, include the reference entity point and the corresponding reference connection edge involved in the target query request, and the reference entity point and the reference connection edge are searched in the indication graph library for the entity point corresponding to the reference entity point and the reference connection edge to obtain the answer. When there is no existing answer in the knowledge graph library, a generation model is used to generate a natural language answer, and the generation model is based on the existing knowledge graph library structure (such as generation based on knowledge prompts). Incremental updates are supported. When new data arrives, the newly added entity points or connection edges are automatically identified and the knowledge graph library is updated. User feedback-driven updates are supported: after the user marks the answer as wrong, the system automatically repairs the knowledge graph library. Support for multi-version control: retain multiple versions of the knowledge graph library, support backtracking and auditing. Piece-driven update: listen to code repository changes through Git Webhook to trigger incremental knowledge extraction. Stream processing: use Apache Flink to analyze log streams in real time, detect new error patterns and update the graph. For example, associate the error code "Nova 500" in the log with the "resource competition" node in the knowledge graph, and map it to the voice description "virtual machine startup failure"). In this embodiment, the construction method in the knowledge graph library can be to extract entity points from information in the target business field, and obtain the relationship between the entity point and other entity points in the knowledge graph library, so as to store the entity point in the knowledge graph library and construct a connection edge between the entity point and the existing entity points in the graph library (eliminating duplicate entities, merging the same information from different sources, and ensuring the consistency and accuracy of the knowledge graph) for easy retrieval, and use domain adaptive models (such as fine-tuned BERT-NER) to extract entity relationships from unstructured text (such as logs, community discussions).
[0038] As an optional implementation manner, converting the current information query intent of the target query request according to the multiple semantic features includes:
[0039] Performing feature fusion on the plurality of semantic features to obtain fused multimodal semantic features;
[0040] The multimodal semantic features are semantically parsed to obtain the information query intent of the target query request.
[0041] Optionally, in an embodiment of the present application, a multimodal semantic feature may be converted by, but is not limited to, aligning and fusing multiple semantic features. The purpose of aligning and fusing multiple semantic features is to map the semantic features of different modalities to a unified feature space. The method of fusing multimodal semantic features may be, but is not limited to, introducing a learnable gating mechanism using a dynamic gating network, dynamically adjusting the contribution weights of the semantic features corresponding to different modalities to the query intent representing the target query request based on multiple semantic features input into the dynamic gating network, and then using the contribution weights to perform a weighted summation of multiple semantic features to determine the fused semantic feature as the information query intent.
[0042] Optionally, in an embodiment of the present application, the method of converting multimodal semantic features according to multiple semantic features can also be based on the combination of generative adversarial networks (GANs) to achieve special and effective multimodal fusion, where the generator maps the multimodal input (such as text, image, and voice) into a unified vector, and the discriminator determines whether the vector can reconstruct the original modal data. Through adversarial training, the vector is forced to retain the core information of all modalities, and even if a certain modality is missing (such as only text input), a sound multimodal vector can still be generated. The specific implementation steps are as follows:
[0043] 1. Define the generator;
[0044] Input: Semantic vectors of different modalities (such as text vector v t , image vector v i , speech vector v s ).
[0045] Goal: fuse these vectors into a unified multimodal vector v mm .
[0046] operate:
[0047] Feature extraction: Semantic vectors are extracted for each modality input.
[0048] Fusion module: Design a fusion module to fuse vectors of different modalities. This can be done by weighted summation, concatenation, and passing through a fully connected layer.
[0049] For example, weighted sum:
[0050] v min =α t v t +α i v i +α s v s
[0051] Among them, the weight α t ,αi ,α s Can be learned through training.
[0052] Or concatenate and pass through the fully connected layer: v min =ReLU(W[v t ;v i ;v s ]+b)
[0053] Output: The generator outputs a unified multimodal vector v mm .
[0054] 2. Define the Discriminator;
[0055] Input: multimodal vector v generated by the generator mm .
[0056] Objective: Determine whether the vector can reconstruct the original modal data.
[0057] operate:
[0058] Reconstruction module: Design a reconstruction module and try to mm Reconstruct the original modality data (such as text, image, speech).
[0059] For example, for text reconstruction, a decoder can be used:
[0060]
[0061] For image reconstruction, a generative network can be used:
[0062]
[0063] For speech reconstruction, an acoustic model can be used:
[0064]
[0065] Loss calculation: Calculate the difference between the reconstructed data and the original data as the loss function of the discriminator.
[0066] For example, using mean squared error (MSE):
[0067]
[0068] Output: The discriminator outputs a probability value indicating whether the vector can reconstruct the original modality data.
[0069] 3. Adversarial training;
[0070] Objective: Through adversarial training of the generator and the discriminator, the multimodal vectors generated by the generator can retain the core information of all modalities.
[0071] operate:
[0072] Generator loss: The goal of the generator is to make the discriminator mistakenly judge the generated vector as real data. The loss function of the generator can be defined as:
[0073]
[0074] Among them, D(vmm) is the output probability of the discriminator.
[0075] Discriminator loss: The goal of the discriminator is to distinguish between generated vectors and real data. The loss function of the discriminator can be defined as:
[0076]
[0077] Among them, Pdata is the distribution of real data, and PG is the distribution of data generated by the generator.
[0078] Training process:
[0079] Alternately train the generator and discriminator:
[0080] The generator is fixed and the discriminator is trained to better distinguish between real data and generated data.
[0081] The discriminator is fixed and the generator is trained so that the vectors it generates can better fool the discriminator.
[0082] Through multiple iterations, the vector generated by the generator can pass the test of the discriminator while retaining the core information of all modalities.
[0083] 4. Handle the missing mode situation;
[0084] Goal: Even if one modality is missing (e.g., only text input), the generator can still produce a robust multimodal vector.
[0085] operate:
[0086] Modality completion: During training, the inputs of certain modalities are randomly discarded, forcing the generator to learn how to generate a complete multimodal vector relying only on other modalities.
[0087] Regularization: A regularization term is added to the generator’s loss function to ensure that the generated vector remains stable and consistent even in the absence of modality.
[0088] For example, adding reconstruction loss:
[0089]
[0090] Where λ is the regularization coefficient, is the reconstruction loss.
[0091] Through these steps, the generator learns how to map multimodal inputs into a unified vector while preserving the core information of all modalities. The discriminator supervises the generator by determining whether the vector can reconstruct the original modal data. Through adversarial training, the model can handle the missing modalities and generate robust multimodal vectors.
[0092] Optionally, in an embodiment of the present application, the information query intention can be, but is not limited to, obtained based on the fused multimodal semantic features, using a deep learning model to perform semantic understanding of the multimodal semantic features, using a pre-trained language model to encode the multimodal semantic features, learn the semantic information therein, understand the user's intention and semantics, and identify the information query intention corresponding to the multimodal semantic features through analysis of the language information. In this embodiment, to improve the reliability of information intent, dynamic context modeling can also be used to determine the target account's current information query intent. This can be achieved by using a pre-trained language model to encode multimodal semantic features, learn the semantic information therein, understand the user's intent and semantics, analyze the language information, identify reference information query intent corresponding to the multimodal semantic features, and obtain multiple initial information query intents corresponding to multiple information query requests of the target account before the current moment. The multiple initial information query intents are arranged in chronological order, and the target account's intent transformation information is converted from the multiple initial information query intents. The intent transformation information is used to characterize how the target account's information query intent changes over time. The reference information query intent is then modified based on the intent transformation amount indicated by the intent transformation information to obtain the final information query intent. This method ensures the accuracy and reliability of the information query intent by modeling the current information query intent based on the target account's query intent in adjacent historical time periods.
[0093] Through the above content, by fusing multiple semantic features, multimodal semantic features are obtained, so that the semantic features included in the multimodal input file can be perceived, and then semantic analysis is performed based on the multimodal semantic features, so as to perceive the information query intention corresponding to the target query request, realize the conversion of information query intention according to the semantic features of the multimodal input file, and improve the accuracy of the information query intention.
[0094] As an optional implementation manner, the performing feature fusion on the plurality of semantic features to obtain a fused multimodal semantic feature includes one of the following:
[0095] Performing text recognition on the file content carried by the multimodal input files to obtain text information carried by each of the input files; converting association information of the input files of each modality based on the information content of the text information, wherein the association information is used to indicate a dependency relationship between the information content of the input file of the current modality and the information content of the input files of other modalities in the multimodal input files; assigning a target weight parameter to the multimodal input files based on the dependency relationship, wherein the target weight parameter is used to indicate the influence of the corresponding input file on the current query intent of the target account; performing a weighted sum calculation on the corresponding semantic features using the target weight parameter to obtain the multimodal semantic features;
[0096] By inputting multiple semantic features into the target generation model, the target vector features corresponding to each semantic feature output by the target generation model are obtained, wherein the multiple target vector features are vector features in the target vector space, and the target generation model records the conversion relationship between the vector features in the vector space corresponding to the multimodal input file and the vector features in the target vector space; the multiple target vector features are input into the target fusion model to obtain the multimodal semantic features output by the target fusion model, wherein the target fusion model records the fusion relationship between multiple vector features.
[0097] Optionally, in an embodiment of the present application, the association information of the input file can be converted in the following manner, but is not limited to: extracting target keywords from the text information, wherein the target keywords are used to characterize the core content described by the corresponding input file, calculating the correlation between the target keywords in each target input file included in the multiple input files and the target keywords in the reference input file, and obtaining a correlation parameter, wherein the association information includes the correlation parameter, and the reference input file is a file other than the reference input file in the multiple input files. Then, according to the dependency relationship, the target weight parameter can be assigned to the input files of multiple modalities by calculating the sum of the correlation parameters of all the target keywords included in each input file, obtaining the target correlation parameter of each input file, and assigning the corresponding target weight parameter according to the high or low target correlation parameters of the input files.
[0098] Through the above content, by assigning a corresponding target weight parameter to each semantic feature based on the dependency relationship between the input files, multiple semantic features are weighted and merged according to the dependency relationship between the input files, thereby merging features based on the influence of the input files on the query semantics, making the semantic content expressed by the merged multimodal semantic features more realistic and reliable. By using the target generation model to perform feature processing on multiple semantic features, semantic features of multiple different feature dimensions are converted to the same semantic feature dimension. Then, the dataset uses a fusion model to fuse the multiple converted target vector features to obtain multimodal semantic features, thereby ensuring the accuracy of the multimodal semantic features.
[0099] As an optional implementation manner, performing semantic parsing on the multimodal semantic features to obtain the information query intent of the target query request includes:
[0100] Inputting the multimodal semantic features into a target semantic generation model, wherein the target semantic generation model records the conversion relationship between the semantic features and the query intent;
[0101] Obtain the information query intention output by the target semantic generation model.
[0102] As an optional implementation manner, the generating of target reply information of the target query request according to the information query intention and the preset business information of the target business field includes:
[0103] Searching the target knowledge graph for candidate entities having a target association relationship with the target entity and target relationship information, wherein the information query intent includes the target entity and the target relationship information, the target entity is used to indicate the sub-business to which the query information requested by the target query request belongs, the target relationship information is used to indicate the transfer relationship between the target entities, the target knowledge graph records multiple entities and relationship information for indicating the transfer relationship between the multiple entities, and the entity is used to indicate a sub-business included in the target business field;
[0104] Obtaining candidate business information of the sub-business corresponding to the candidate entity in the preset business information;
[0105] The target reply information is generated according to the candidate service information.
[0106] Optionally, in an embodiment of the present application, the candidate business information provides information support for the generated target reply information. When the degree of association between the target entity and the candidate entity is greater than a certain threshold, the candidate business information can be considered to be the content requested by the target query request, and the candidate business information can be directly used as the target reply information. When the degree of association between the target entity and the candidate entity is less than or equal to a certain threshold, the candidate business information can be considered to be related to the target query request, but it is not yet a standard reply to the target query question. Therefore, a target reply information can be generated based on the content of the candidate business information and the information query intention. When there is no target reply information that answers the target query question in the target business field, the target indication map can be updated according to the currently generated target reply information, and the target reply information can be updated to the preset business information included in the target business field, thereby enriching the information content in the information library of the target business field.
[0107] Through the above content, by constructing a knowledge graph, each sub-business is stored as an entity in the knowledge graph, and the transfer relationship between sub-businesses is constructed into the relationship information between entities, thereby improving the maintenance efficiency of business information and information query efficiency in the target business field.
[0108] As an optional implementation manner, generating the target reply information according to the candidate service information includes:
[0109] When the correlation between the target entity and the candidate entity is less than or equal to the target correlation, the candidate business information and the information query intention are input into the target language model to obtain the target reply information output by the target language model, wherein the target language model records the business information and the transfer relationship between the query intention and the reply information.
[0110] Optionally, in an embodiment of the present application, when the degree of association between the target entity and the candidate entity is less than or equal to the target degree of association, it can be considered that the candidate business information is not the standard answer to the target query request, but there is a certain degree of correlation between the candidate business information and the target query request, so the content of the candidate business information can be used to generate reply content for the target query request.
[0111] Through the above content, when the correlation between the target entity and the candidate entity is less than or equal to the target correlation, the target semantic model is used to perform identification processing based on the candidate business information and information query intention, so as to realize the use of the candidate business information as the information basis for replying the target query request, thereby generating more accurate target reply information.
[0112] As an optional implementation manner, generating the target reply information according to the candidate service information includes:
[0113] In a case where the degree of association between the target entity and the candidate entity is greater than a target degree of association, the candidate business information is determined as the target reply information.
[0114] Figure 3 This is a diagram of the question-answering system architecture according to an embodiment of the present application. Figure 3 As shown in the figure, the question-answering system adopts a layered architecture, and the core modules are as follows:
[0115] Multimodal perception layer:
[0116] Supports text, voice, image, video, and sensor data (such as IoT devices) input, and uniformly encodes them into semantic vectors through a cross-modal alignment network;
[0117] Integrate emotion recognition module to analyze user emotional state (such as urgency, confusion) in real time to optimize response strategy.
[0118] Intelligent understanding layer:
[0119] Generate context-aware intent vectors based on a large model (integrating BERT and GPT-4 architecture);
[0120] A bidirectional attention mechanism is used to associate historical conversations with current questions, and multi-granularity intent parsing: user questions are decomposed into atomic intents ("comparing performance") and an intent dependency graph is constructed.
[0121] Dynamic Knowledge Engine:
[0122] Knowledge graph construction: Based on the Neo4j graph database, it includes a three-layer structure of entities, relationships, and attributes;
[0123] Real-time data stream processing: Using Apache Flink to access external data sources (such as news APIs and academic databases) in real time, triggering incremental updates to the knowledge graph (with a latency of <30 seconds).
[0124] Cross-domain reasoning module: Use graph neural network (GNN) to realize the association reasoning of knowledge nodes.
[0125] Decision generation layer:
[0126] Multi-strategy answer generation: Combine rule templates, generative models (such as T5) and retrieval-enhanced generation to select the optimal output form according to the scenario.
[0127] Closed-loop optimization layer:
[0128] Feedback perception module: collects explicit feedback (user ratings) and implicit signals (conversation turns, answer adoption rate, etc.);
[0129] Incremental learning engine: Based on the Elastic Weight Curing (EWC) algorithm, only key parameters are updated to avoid catastrophic forgetting.
[0130] As shown in the figure, the process of the intelligent question answering method is as follows:
[0131] Step 1: Multimodal input parsing and alignment;
[0132] Speech input: converted to text through end-to-end speech recognition (ASR), and pitch and speech rate features are extracted;
[0133] Image input: Generate descriptive text through a joint visual-semantic embedding model;
[0134] Text, speech, and image features are fused into a unified semantic vector via the CMAN network.
[0135] Step 2: Deep semantic understanding and intent decomposition;
[0136] Context encoding: Encode the current question and historical conversations into time-series vectors and capture long-term dependencies through LSTM networks;
[0137] Use large models to generate context-aware semantic representations and identify user intent (e.g., consultation, complaint, suggestion) to improve classification accuracy;
[0138] Step 3: Dynamic knowledge retrieval and causal reasoning;
[0139] Parallel search of local knowledge graphs, external databases (encyclopedias, etc.) and real-time data streams;
[0140] Integrate multi-source results through a confidence-weighted fusion algorithm, exclude low-quality answers with confidence < 90%, and only retain results with confidence >= 90%.
[0141] Knowledge reasoning: Generate reasoning paths based on GNN and give the correct path;
[0142] Step 4: Personalized answer generation and multimodal output;
[0143] User profile matching: Generate customized answers based on user historical behavior and search type;
[0144] Multimodal output: Generate answers in formats suitable for different terminals through TTS (text-to-speech) and data visualization engines;
[0145] Step 5: Feedback-driven system self-optimization;
[0146] Collect user behavior data (such as answer adoption rate, question depth, and user ratings);
[0147] Incremental learning is triggered daily, and model parameters are updated every morning using a gradient clipping + elastic weight solidification strategy to ensure that historical knowledge is not overwritten.
[0148] The knowledge graph dynamically adjusts node priorities based on popularity, and frequently accessed nodes are cached to the edge.
[0149] The embodiments of the present application have the following beneficial effects: a) Real-time knowledge: the delay from data collection to knowledge graph update is less than 1 minute, supporting immediate response to emergencies; b) Multimodal collaboration: the accuracy of cross-modal question and answer is improved, supporting 10+ input / output forms, and can more accurately give users the answers they want; c) Adaptive ability: through a feedback loop, the system automatically optimizes the model every month, and the error rate decreases; d) Improved precision: the accuracy of question and answer is improved compared to traditional systems, and the ability to handle complex problems is increased by 3 times, which greatly reduces the costs brought by human maintenance.
[0150] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0151] The embodiment of the present application also provides a device for generating reply information. Figure 4 This is a structural block diagram of the generation of a reply message according to an embodiment of the present application, such as Figure 4 As shown, the device includes:
[0152] a receiving module, configured to receive a target query request input by a target account when requesting to query information within a target business domain, wherein the target query request includes a multimodal input file, and the multimodal input file is used to represent the information query conditions of the target query request from multiple dimensions;
[0153] An extraction module, configured to perform feature extraction on the multimodal input files to obtain semantic features of the input files of each modality, wherein the semantic features are used to indicate the information query requirements described by the corresponding input files;
[0154] A conversion module, configured to convert the current information query intent of the target query request according to the plurality of semantic features;
[0155] A generating module, configured to generate target reply information of the target query request according to the information query intention and the preset business information of the target business field;
[0156] A sending module is used to send the target reply information to the target account.
[0157] Through the above device, the target account includes a multimodal input file in the input target query request, and the multimodal input file carries the user's real query needs, and then the multimodal input file is subjected to feature extraction to obtain the semantic features of each modal input file, thereby realizing the perception of the information query needs described by the input file, and then the information query intention indicated by the target query request is converted by converting multiple semantic features, and the real information query intention of the target query request is converted by semantic understanding of the multimodal input file, thereby ensuring the accuracy of the information query intention, and then generating target reply information based on more detailed query intentions and preset business information of the target business field, thereby improving the matching degree between the target reply information and the target query request, and therefore, it can solve the technical problem of low efficiency in generating reply information in related technologies, and achieve the technical effect of improving the efficiency in generating reply information.
[0158] Optionally, the conversion module includes:
[0159] A fusion unit, configured to fuse the plurality of semantic features to obtain a fused multimodal semantic feature;
[0160] The parsing unit is used to perform semantic parsing on the multimodal semantic features to obtain the information query intent of the target query request.
[0161] Optionally, the fusion unit is configured to perform one of the following operations:
[0162] Performing text recognition on the file content carried by the multimodal input files to obtain text information carried by each of the input files; converting association information of the input files of each modality based on the information content of the text information, wherein the association information is used to indicate a dependency relationship between the information content of the input file of the current modality and the information content of the input files of other modalities in the multimodal input files; assigning a target weight parameter to the multimodal input files based on the dependency relationship, wherein the target weight parameter is used to indicate the influence of the corresponding input file on the current query intent of the target account; performing a weighted sum calculation on the corresponding semantic features using the target weight parameter to obtain the multimodal semantic features;
[0163] By inputting multiple semantic features into the target generation model, the target vector features corresponding to each semantic feature output by the target generation model are obtained, wherein the multiple target vector features are vector features in the target vector space, and the target generation model records the conversion relationship between the vector features in the vector space corresponding to the multimodal input file and the vector features in the target vector space; the multiple target vector features are input into the target fusion model to obtain the multimodal semantic features output by the target fusion model, wherein the target fusion model records the fusion relationship between multiple vector features.
[0164] Optionally, the parsing unit is used to:
[0165] Inputting the multimodal semantic features into a target semantic generation model, wherein the target semantic generation model records the conversion relationship between the semantic features and the query intent;
[0166] Obtain the information query intention output by the target semantic generation model.
[0167] Optionally, the generating module includes:
[0168] A search unit is configured to search a target knowledge graph for a candidate entity having a target association relationship with a target entity and target relationship information, wherein the information query intent includes the target entity and the target relationship information, the target entity is used to indicate the sub-business to which the query information requested by the target query request belongs, the target relationship information is used to indicate a transfer relationship between the target entities, the target knowledge graph records a plurality of entities and relationship information indicating transfer relationships between the plurality of entities, and the entity is used to indicate a sub-business included in the target business field;
[0169] an acquiring unit, configured to acquire candidate business information of the sub-business corresponding to the candidate entity in the preset business information;
[0170] A generating unit is configured to generate the target reply information according to the candidate service information.
[0171] Optionally, the generating unit is configured to:
[0172] When the correlation between the target entity and the candidate entity is less than or equal to the target correlation, the candidate business information and the information query intention are input into the target language model to obtain the target reply information output by the target language model, wherein the target language model records the business information and the transfer relationship between the query intention and the reply information.
[0173] Optionally, the generating unit is configured to:
[0174] In a case where the degree of association between the target entity and the candidate entity is greater than a target degree of association, the candidate business information is determined as the target reply information.
[0175] For the description of the features in the embodiment corresponding to the reply information generation device, please refer to the relevant description of the embodiment corresponding to the reply information generation method, which will not be repeated here.
[0176] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the method for generating reply information.
[0177] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned embodiments of the method for generating reply information when running.
[0178] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0179] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned embodiments of the method for generating reply information are implemented.
[0180] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned embodiments of the method for generating reply information.
[0181] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0182] The above is a detailed introduction to a method and device for generating reply information, a storage medium, and an electronic device provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for generating a reply message, characterized in that: include: receiving a target query request input by a target account when requesting to query information within a target business domain, wherein the target query request includes a multimodal input file, and the multimodal input file is used to represent information query conditions of the target query request from multiple dimensions; Performing feature extraction on the multimodal input files to obtain semantic features of the input files in each modality, wherein the semantic features are used to indicate the information query requirements described by the corresponding input files; Converting the current information query intent of the target query request according to the plurality of semantic features; Generate target response information for the target query request according to the information query intention and the preset business information of the target business field; The target reply information is sent to the target account.
2. The method according to claim 1, characterized in that The converting the current information query intent of the target query request according to the plurality of semantic features includes: Performing feature fusion on the plurality of semantic features to obtain fused multimodal semantic features; The multimodal semantic features are semantically parsed to obtain the information query intent of the target query request.
3. The method according to claim 2, characterized in that The performing feature fusion on the plurality of semantic features to obtain a fused multimodal semantic feature includes one of the following: Performing text recognition on the file content carried by the multimodal input files to obtain text information carried by each of the input files; converting association information of the input files of each modality based on the information content of the text information, wherein the association information is used to indicate a dependency relationship between the information content of the input file of the current modality and the information content of the input files of other modalities in the multimodal input files; assigning a target weight parameter to the multimodal input files based on the dependency relationship, wherein the target weight parameter is used to indicate the influence of the corresponding input file on the current query intent of the target account; performing a weighted sum calculation on the corresponding semantic features using the target weight parameter to obtain the multimodal semantic features; By inputting multiple semantic features into the target generation model, the target vector features corresponding to each semantic feature output by the target generation model are obtained, wherein the multiple target vector features are vector features in the target vector space, and the target generation model records the conversion relationship between the vector features in the vector space corresponding to the multimodal input file and the vector features in the target vector space; the multiple target vector features are input into the target fusion model to obtain the multimodal semantic features output by the target fusion model, wherein the target fusion model records the fusion relationship between multiple vector features.
4. The method according to claim 2, characterized in that The performing semantic parsing on the multimodal semantic features to obtain the information query intent of the target query request includes: Inputting the multimodal semantic features into a target semantic generation model, wherein the target semantic generation model records the conversion relationship between the semantic features and the query intent; Obtain the information query intention output by the target semantic generation model.
5. The method according to claim 1, wherein The generating of target response information of the target query request according to the information query intention and the preset business information of the target business field includes: Searching the target knowledge graph for candidate entities having a target association relationship with the target entity and target relationship information, wherein the information query intent includes the target entity and the target relationship information, the target entity is used to indicate the sub-business to which the query information requested by the target query request belongs, the target relationship information is used to indicate the transfer relationship between the target entities, the target knowledge graph records multiple entities and relationship information for indicating the transfer relationship between the multiple entities, and the entity is used to indicate a sub-business included in the target business field; Obtaining candidate business information of the sub-business corresponding to the candidate entity in the preset business information; The target reply information is generated according to the candidate service information.
6. The method according to claim 5, characterized in that The generating the target reply information according to the candidate service information includes: When the correlation between the target entity and the candidate entity is less than or equal to the target correlation, the candidate business information and the information query intention are input into the target language model to obtain the target reply information output by the target language model, wherein the target language model records the business information and the transfer relationship between the query intention and the reply information.
7. The method according to claim 5, characterized in that The generating the target reply information according to the candidate service information includes: In a case where the degree of association between the target entity and the candidate entity is greater than a target degree of association, the candidate business information is determined as the target reply information.
8. A device for generating reply information, characterized in that: include: a receiving module, configured to receive a target query request input by a target account when requesting to query information within a target business domain, wherein the target query request includes a multimodal input file, and the multimodal input file is used to represent the information query conditions of the target query request from multiple dimensions; An extraction module, configured to perform feature extraction on the multimodal input files to obtain semantic features of the input files of each modality, wherein the semantic features are used to indicate the information query requirements described by the corresponding input files; A conversion module, configured to convert the current information query intent of the target query request according to the plurality of semantic features; A generating module, configured to generate target reply information of the target query request according to the information query intention and the preset business information of the target business field; A sending module is used to send the target reply information to the target account.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method for generating reply information as claimed in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method for generating reply information according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Distribution network fault recovery method and system based on artificial intelligence
CN117458432A
Multi-modal metadata retrieval enhancement generation method and system
CN118626662A
Intelligent question and answer method and system, electronic equipment and readable storage medium
CN119046440A
Knotarization intelligent question and answer customer service method and system based on knowledge graph
CN119938816A