Intelligent question and answer method and device and electronic equipment

By adopting a global Q&A model in the intelligent Q&A system, combining the local Q&A model and the characteristic data of the knowledge database, and updating the model parameters, the problem of time-consuming and low accuracy of the intelligent Q&A system caused by the messy information of the knowledge database is solved, and efficient and accurate intelligent Q&A and effective utilization of multi-source data is achieved.

CN120216624APending Publication Date: 2025-06-27CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411472719.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The internal knowledge database information of the enterprise is messy, which leads to a long time-consuming and low accuracy of the intelligent question-and-answer system, which is unable to effectively integrate multi-source databases, affecting employee work efficiency and customer perception.

Method used

The global question and answer model is adopted to receive the question information of the target object and analyze and process the answer information generated. The model parameters are updated based on the parameters of multiple local question and answer models and the target feature vectors. The target feature vector is determined by the first feature data output by the local question and answer model and the second feature data in the knowledge database.

Benefits of technology

It improves the efficiency and accuracy of intelligent question-and-answer questions and answers, optimizes resource utilization, protects data privacy, solves the problem of messy information in the knowledge database, and realizes the effective integration and utilization of multi-source data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216624A_ABST
    Figure CN120216624A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent question and answer method and device and electronic equipment. The method comprises the following steps: receiving question information of a target object; the question information is analyzed and processed through a global question and answer model, answer information corresponding to the question information is generated, model parameters of the global question and answer model are updated according to model parameters of the multiple local question and answer models and the target feature vector, and answer information is generated; the target feature vector is determined according to first feature data output by a plurality of local question and answer models and second feature data in a plurality of knowledge databases, and the first feature data is a question category obtained after the local question and answer models classify historical question and answer information through a base classifier; the second feature data are multi-modal features corresponding to the first feature data. According to the method and the device, the technical problems of relatively long time consumption and relatively low accuracy of intelligent question answering caused by disordered knowledge database information in related technologies are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language technology, and in particular, to an intelligent question-answering method, device, and electronic device. Background Art

[0002] Enterprise internal knowledge databases are important resources in an organization, which help with knowledge sharing, reduce duplicate work, improve work efficiency, promote teamwork and innovation, protect important information from being lost, and effectively manage and utilize the knowledge assets within the organization.

[0003] However, due to problems such as the vast amount of information in each department of the enterprise, inconsistent data types, and high privacy requirements, it is difficult to integrate different knowledge databases, and a high-quality multi-source database cannot be provided for the intelligent question-answering system. At the same time, since the intelligent question-answering systems in related technologies generally come from the enterprise's own data and do not directly provide more useful external information, more accurate answers cannot be provided during the question-answering process with users, seriously affecting the work efficiency of employees and even the customer perception. In addition, the enterprise internal knowledge databases need to be manually maintained and updated, specific keywords are required to search for answers, and answers even need to be found in other electronic or paper documents, accompanied by the risk of privacy leakage.

[0004] For the above problems, no effective solutions have been proposed yet. Summary of the Invention

[0005] Embodiments of this application provide an intelligent question-answering method, device, and electronic device to at least solve the technical problems of long time consumption and low accuracy in intelligent question-answering caused by the messy information in the knowledge databases in related technologies.

[0006] According to one aspect of the embodiments of this application, an intelligent question-answering method is provided, including: receiving question information of a target object; analyzing and processing the question information through a global question-answering model to generate answer information corresponding to the question information, where the model parameters of the global question-answering model are updated based on the model parameters of multiple local question-answering models and a target feature vector, the target feature vector is determined based on first feature data output by multiple local question-answering models and second feature data in multiple knowledge databases, the first feature data is the question category after classifying historical question-answering information by a base classifier of the local question-answering model, and the second feature data is the multi-modal feature corresponding to the first feature data.

[0007] Optionally, the target feature vector is determined based on the first feature data output by multiple local question-answering models and the second feature data in multiple knowledge databases, including: classifying historical question-answering information through a base classifier in the local question-answering model to obtain the first feature data, where the historical question-answering information includes historical question information and historical answer information; determining the second feature data in the target database based on the first feature data, where the target database is the knowledge database among the multiple knowledge databases that contains the first feature data, and there is an association relationship between the first feature data and the second feature data; determining the target feature vector based on the second feature data.

[0008] Optionally, before classifying the historical question-answering information through the base classifier in the local question-answering model, the method further includes: determining the number of base classifiers; determining the question category of the historical question-answering information, where the question category includes at least one of the following: fact category, definition category, process category, comparison category, and suggestion category; determining an encoding matrix based on the number of base classifiers and the question category, where the encoding information in the encoding matrix corresponds one-to-one with the question category.

[0009] Optionally, classifying the historical question-answering information through the base classifier in the local question-answering model to obtain the first feature data includes: determining the output vector of the historical question-answering information through the base classifier; determining the Euclidean distance of the historical question-answering information based on the output vector and the row vector of the encoding matrix; determining the cosine similarity of the historical question-answering information based on the output vector and the row vector of the encoding matrix; determining the first feature data based on the Euclidean distance and the cosine similarity.

[0010] Optionally, determining the target feature vector based on the second feature data includes: removing the noise data of the text feature and the image feature in the second feature data through discrete wavelet transform to obtain the first text feature vector and the first image feature vector; reducing the spatial dimensions of the first text feature vector and the first image feature vector through a hash function to obtain the second text feature vector and the second image feature vector; fusing the second text feature vector and the second image feature vector through fast Fourier transform to obtain the target feature vector.

[0011] Optionally, the model parameters of the global question-answering model are updated based on the model parameters of multiple local question-answering models and the target feature vector, including: determining the model parameters of the global question-answering model, where the model parameters of the global question-answering model are the average values of the model parameters of multiple local question-answering models; determining the clustering result of the local models based on the model parameters of multiple local models and the target feature vector, where the clustering result is used to represent the similarity degree between each local question-answering model and the global question-answering model; determining the weight coefficients of multiple local question-answering models based on the clustering result; updating the model parameters of the global question-answering model based on the weight coefficients.

[0012] Optionally, determining the clustering result of the local models based on the model parameters of multiple local models and the target feature vector includes: determining the clustering centers of multiple local question-and-answer models based on the target feature vector, where the clustering centers are used to represent the average values of the model parameters of multiple local question-and-answer models in the clustering clusters, and the clustering clusters are used to represent the set of model parameters of multiple local question-and-answer models; determining the Euclidean distance between the model parameters of the global question-and-answer model and the clustering centers, and assigning the model parameters of the global question-and-answer model to the target clustering cluster, where the target clustering cluster is used to represent the clustering cluster where the clustering center corresponding to the minimum Euclidean distance is located; and updating the clustering centers of the target clustering cluster to obtain the clustering result.

[0013] According to another aspect of the embodiments of the present application, there is also provided an intelligent question-and-answer device, including: a receiving module, configured to receive the question information of the target object; a generating module, configured to analyze and process the question information through a global question-and-answer model to generate answer information corresponding to the question information, where the model parameters of the global question-and-answer model are updated based on the model parameters of multiple local question-and-answer models and the target feature vector, the target feature vector is determined based on the first feature data output by multiple local question-and-answer models and the second feature data in multiple knowledge databases, the first feature data is the question category after the local question-and-answer model classifies the historical question-and-answer information through a base classifier, and the second feature data is the multimodal feature corresponding to the first feature data.

[0014] According to yet another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, where the memory is configured to store program instructions; the processor is connected to the memory and is configured to execute to implement the above-mentioned intelligent question-and-answer method.

[0015] According to still another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium, where the non-volatile storage medium includes a stored computer program, and the device where the non-volatile storage medium is located executes the above-mentioned intelligent question-and-answer method by running the computer program.

[0016] According to still another aspect of the embodiments of the present application, there is also provided a computer program product, including computer instructions, and when the computer instructions are executed by a processor, the above-mentioned intelligent question-and-answer method is implemented.

[0017] In an embodiment of the present application, by receiving the question information of the target object; analyzing and processing the question information through a global question-answering model to generate answer information corresponding to the question information, wherein the model parameters of the global question-answering model are updated based on the model parameters of multiple local question-answering models and the target feature vector, the target feature vector is determined based on the first feature data output by multiple local question-answering models and the second feature data in multiple knowledge databases, the first feature data is the question category after the local question-answering model classifies the historical question-and-answer information through a base classifier, and the second feature data is the multi-modal feature corresponding to the first feature data, the purpose of improving the efficiency and accuracy of intelligent question-answering within the enterprise is achieved, thereby realizing the technical effects of optimizing resource utilization, protecting data privacy, and solving the messy information in the knowledge database, and further solving the technical problem that the intelligent question-answering takes a long time and has low accuracy due to the messy information in the knowledge database in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0019] Figure 1 is a hardware structure diagram of a computer terminal for implementing an intelligent question-answering method according to an embodiment of the present application;

[0020] Figure 2 is a flowchart of an intelligent question-answering method according to an embodiment of the present application;

[0021] Figure 3 is a framework structure diagram of an intelligent question-answering system according to an embodiment of the present application;

[0022] Figure 4 is a structure diagram of an intelligent question-answering device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0024] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] In order to solve the problem of low efficiency of intelligent question answering in related technologies, an embodiment of this application provides an intelligent question answering method, which can run on Figure 1 the computer terminal shown below, and the computer terminal will be described below.

[0026] The embodiment of the intelligent question answering method provided by the embodiment of this application can be executed on a mobile terminal, a computer terminal or a similar computing device. Figure 1 The following shows a hardware structure block diagram of a computer terminal for implementing the intelligent question answering method. As Figure 1 shown, the computer terminal 10 may include one or more processors (the processors may include, but are not limited to, processing devices such as a microprocessor MCU or a field programmable gate array FPGA, shown as 102a, 102b,..., 102n in the figure), a memory 104 for storing data, and a transmission module 106 for communication functions connected by wired and / or wireless networks. In addition, it may further include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0027] It should be noted that one or more of the above-mentioned processors and / or other data processing circuits can generally be referred to as "data processing circuits" herein. This data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or incorporated in whole or in part into any one of the other components in the computer terminal 10. As involved in the embodiments of the present application, this data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0028] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the intelligent question-answering method in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned intelligent question-answering method. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely set relative to the processor, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0029] The transmission module 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission module 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0030] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10.

[0031] It should be noted here that in some alternative embodiments, the above-mentioned Figure 1 shown computer terminal can include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance and is intended to show the types of components that can exist in the above-mentioned computer terminal.

[0032] Under the above operating environment, an embodiment of the intelligent question-answering method provided by the present application is described. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0033] Figure 2 is a flowchart of an intelligent question-answering method according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:

[0034] Step S202, receiving the question information of the target object.

[0035] In the above step S202, the intelligent question-answering system can receive the user's question information through the question-answering interface where the target object (user) is located or other input methods. The question information can be text information, voice information, image information, and complex query information including several of them.

[0036] Step S204, analyzing and processing the question information through a global question-answering model to generate answer information corresponding to the question information. Among them, the model parameters of the global question-answering model are updated based on the model parameters of multiple local question-answering models and the target feature vector. The target feature vector is determined based on the first feature data output by multiple local question-answering models and the second feature data in multiple knowledge databases. The first feature data is the question category after the local question-answering model classifies the historical question-answering information through a base classifier, and the second feature data is the multimodal feature corresponding to the first feature data.

[0037] In the above step S204, by deeply analyzing the user's question information through the optimized global question-answering model, accurate and relevant answer information can be generated. Among them, the global question-answering model is trained through a federated transfer learning framework and combined with the parameters and multimodal features of multiple local question-answering models, and can integrate the information of different data sources and knowledge databases to analyze and answer various different questions. Based on the question-answering analysis of this global question-answering model, it can not only improve the efficiency and accuracy of intelligent question-answering, but also protect data privacy and achieve the effective integration and utilization of multi-source data.

[0038] Through the above steps S202 to S204, the purpose of improving the efficiency and accuracy of intelligent question-answering is achieved, thereby realizing the technical effects of optimizing resource utilization, protecting data privacy, and solving the messy information in the knowledge database, and further solving the technical problems of long time consumption and low accuracy of intelligent question-answering caused by the messy information in the knowledge database in the related technology. The following is a detailed description.

[0039] Optionally, the target feature vector is determined based on the first feature data output by multiple local Q&A models and the second feature data in multiple knowledge databases, including: classifying historical Q&A information through a base classifier in the local Q&A model to obtain the first feature data, where the historical Q&A information includes historical question information and historical answer information; determining the second feature data in the target database based on the first feature data, where the target database is the knowledge database among the multiple knowledge databases that contains the first feature data, and there is an association relationship between the first feature data and the second feature data; determining the target feature vector based on the second feature data.

[0040] In an embodiment of the present application, the first feature data (i.e., the question category to which the historical Q&A information belongs after classification) output by multiple base classifiers in the local Q&A model, and the second feature data in the target knowledge database, where the target knowledge database is the knowledge database among all the enterprise internal knowledge databases that contains the first feature data, can provide a comprehensive feature representation for the historical question information, that is, the above-mentioned target feature vector.

[0041] It should be noted that there is a certain association relationship between the first feature data and the second feature data, that is, by combining the category to which the historical Q&A information belongs and the multi-modal features (such as text features, image features, etc.) corresponding to this question category in the relevant knowledge database, it is possible to support the local Q&A model to generate more accurate answers, thereby providing a good optimization basis for the subsequent global question model.

[0042] In an embodiment of the present application, before classifying the historical Q&A information through the base classifier in the local Q&A model, it is also necessary to train the base classifier in the local Q&A model, including: determining the number of base classifiers; determining the question category of the historical Q&A information, where the question category includes at least one of the following: fact class, definition class, process class, comparison class, and suggestion class; determining an encoding matrix based on the number of base classifiers and the question category, where the sum encoding information in the encoding matrix corresponds to the question category one by one. The specific process can be as follows:

[0043] First, the ECOC (Error-Correcting Output Codes) can be used to enhance the system's tolerance to input noise. For example, define an encoding matrix S m×n , where each element value of the encoding matrix S takes values of {-1, 0, 1}, and m×n represents that m sample categories are divided n times. Specifically, m represents the number of categories to be classified, including but not limited to fact class, definition class, process class, comparison class, and suggestion class; n represents the number of base classifiers, which is 7 here.

[0044] Further, a coding sequence of length n can be constructed using a one-to-one coding method, as shown in Table 1. Among them, the set of base classifiers is h1, h2, ..., h7. When the value of the classifier h k is "1" or "-1", it means that the problem is classified as a positive classification sample or a negative classification sample; when the value of the classifier h k is "0", it means that the sample of this type of problem is not used.

[0045] Table 1 ECOC Error Correction Output Coding Matrix

[0046]

[0047] Specifically, according to common problem features, historical Q&A information can be divided into the following 5 categories:

[0048] 1. Fact-based questions: such as "What day of the week is today?", usually have clear answers.

[0049] 2. Definition-based questions: such as "What is machine learning?", usually require explanations or definitions.

[0050] 3. Process-based questions: such as "How to make a cake?", usually require detailed steps or methods.

[0051] 4. Comparison-based questions: such as "What are the differences between Python and Java?", usually require comparisons or contrasts.

[0052] 5. Suggestion-based questions: such as "Which mobile phone should I choose?", usually involve user preference information.

[0053] Secondly, the standard error backpropagation algorithm can be used to train the 7 base classifiers so that they can accurately identify the problem categories of historical Q&A information. Among them, each classifier is responsible for judging whether the input historical question information belongs to one of the five predefined categories. For example, use the error backpropagation algorithm to construct a neural network (such as an L-layer perceptron network), including an input layer Layer1, a hidden layer Layer2, and an output layer Layer3, to finally determine what kind of question it is; then use the standard backpropagation to minimize the cost function, so as to train the optimal classifier.

[0054] Optionally, classify the historical Q&A information through the base classifiers in the local Q&A model to obtain the first feature data, including: determining the output vector of the historical Q&A information through the base classifier; determining the Euclidean distance of the historical Q&A information based on the output vector and the row vector of the coding matrix; determining the cosine similarity of the historical Q&A information based on the output vector and the row vector of the coding matrix; determining the first feature data based on the Euclidean distance and the cosine similarity.

[0055] In the embodiments of the present application, the first feature data, i.e., the question category, can be extracted from the historical Q&A information through the base classifier, and then can be used to determine the second feature data in the target database, thereby providing a basis for the multimodal fusion of the historical Q&A information and the generation of the final answer. The specific process can be as follows:

[0056] First, input the samples of the historical Q&A information into each base classifier to obtain an output vector:

[0057] H(x) = (h1(x), h2(x), h3(x), h4(x), h5(x), h6(x), h7(x))

[0058] In the formula, x represents the sample of the historical Q&A information, H(x) represents the output vector of the historical Q&A information, and h1 - h7 represent 7 base classifiers.

[0059] Secondly, calculate the Euclidean distance D(y) between the output vector H(x) of the historical Q&A information and each row vector in the encoding matrix S. The specific expression is as follows:

[0060]

[0061] In the formula, h i (x) represents the output of the i-th base classifier for the sample x, S yi represents the element in the y-th row and i-th column of the encoding matrix S, and n represents the number of columns of the encoding matrix, that is, the number of base classifiers.

[0062] At the same time, calculate the cosine similarity Cosine_Similarity(y) between the output vector H(x) of the historical Q&A information and each row vector in the encoding matrix S. The specific expression is as follows:

[0063]

[0064] In the formula, y represents the row index of the encoding matrix, · represents the dot product operation, ||H(x)|| and ||S y || respectively represent the norms of the output vector H(x) and the row vector S y .

[0065] Finally, combine the above Euclidean distance and cosine similarity and make a decision by the method of weighted average. Assume that α and β are the weights of the Euclidean distance and cosine similarity respectively, then the comprehensive metric formula is as follows:

[0066] D combined (y) = α·D(y) - β·Cosine_Similarity(y)

[0067] In the formula, D combined(y) represents the comprehensive metric, α represents the weight coefficient of the Euclidean distance, and β represents the weight coefficient of the cosine similarity.

[0068] Select the y with the minimum comprehensive metric as the predicted output problem category:

[0069]

[0070] In the formula, represents the output problem category, that is, the first feature data, and argmin represents the parameter value used to select the best classification result.

[0071] Optionally, determine the target feature vector based on the second feature data, including: removing the noise data of the text feature and the image feature in the second feature data through discrete wavelet transform to obtain the first text feature vector and the first image feature vector; reducing the spatial dimensions of the first text feature vector and the first image feature vector through a hash function to obtain the second text feature vector and the second image feature vector; fusing the second text feature vector and the second image feature vector through fast Fourier transform to obtain the target feature vector.

[0072] In the embodiments of the present application, fusing the same problem category and the multi-modal features in the corresponding knowledge database can obtain a target feature vector that can represent the accurate problem answer, thereby improving the accuracy of problem classification and answering. This process involves multiple links such as feature extraction, denoising, dimensionality reduction, and feature fusion to generate a comprehensive and accurate problem feature representation. The specific process can be as follows:

[0073] First, extract the image features in the historical Q&A information through a pre-trained convolutional neural network to obtain the image feature vector and extract the text features in the historical Q&A information through a pre-trained text encoder to obtain the text feature vector

[0074] Secondly, perform discrete wavelet transform denoising on the extracted image features and text features respectively. The specific expressions are as follows:

[0075]

[0076] In the formula, x img represents the original image feature vector, x txt represents the original text feature vector, DWT represents the discrete wavelet transform operation, represents the low-frequency component after DWT is applied to the image features, represents the high-frequency component after DWT is applied to the image features, represents the low-frequency component after DWT is applied to the text features, Indicates the high-frequency components after applying DWT to text features.

[0077] Subsequently, the high-frequency components are removed, and only the low-frequency components are retained to obtain the first image feature vector and the first text feature vector. The specific expressions are as follows:

[0078]

[0079] In the formula, x i ′ mg represents the denoised image feature vector, that is, the first image feature vector; x t ′ xt represents the denoised text feature vector, that is, the first text feature vector.

[0080] It should be noted that after this step, a compact bilinear pooling operation can also be used to fuse the first image feature vector x′ img and the first text feature vector x′ txt into a unified feature representation for subsequent hashing processing.

[0081] Furthermore, a hash function is used to map the denoised first image feature vector x′ img and the first text feature vector x′ txt to a low-dimensional space respectively. The specific expressions are as follows:

[0082]

[0083] In the formula, x′ img_h represents the hashed image feature vector, that is, the second image feature vector; x′ txt_h represents the hashed text feature vector, that is, the second text feature vector, h img represents the hash function of the first image feature vector, h txt represents the hash function of the first text feature vector, Sketch represents the hash function, which is used to map high-dimensional features to a low-dimensional space.

[0084]

[0085] In the formula, a img , b img , a txt , b txt represent constants, which can be randomly selected from a uniform distribution to uniformly map high-dimensional features to a low-dimensional space. Assume a img , b img , a txt , b txt ~Uniform(0, 1), a img= 0.589, b img = 0.127, a txt = 0.345, b txt = 0.678; floor represents rounding down, mod represents the modulo operation, d′ is the dimension of the target low-dimensional space, and j represents the dimension index. The specific hashed feature vector is as follows:

[0086]

[0087] In the formula, S(i) is the sign function S(i) = (-1) rand(i) , and the sign function is used to assign a sign (+1 or -1) to each feature element. rand(i) is a random number that generates 0 or 1. Among them, S img (i) represents the sign function of the image feature vector, and S txt (j) represents the sign function of the text feature vector. d_img represents the dimension of the first image feature vector x′ img , and d_txt represents the dimension of the first text feature vector x′ txt . x′ img (i) represents the value of the i-th dimension in the first image feature vector x′ img , and x′ txt (j) represents the value of the j-th dimension in the first text feature vector x′ txt .

[0088] Finally, perform a fast Fourier transform on the hashed second image feature x′ img_h and the second text feature x′ txt_h . The specific expression is as follows:

[0089]

[0090] In the formula, represents the fast Fourier transform operation, represents the fast Fourier transform result of the second image feature x i ′ mg_h , represents the fast Fourier transform result of the second text feature x t ′ xt_h .

[0091] Perform a dot product convolution in the frequency domain to obtain the element-wise product of the features:

[0092]

[0093] In the formula, H Z′ represents the result of the dot product convolution of the frequency domain representations of the image and text features, and ⊙ represents the dot product convolution operation.

[0094] Apply the inverse Fourier transform to the convolution result to obtain the fused low-dimensional feature representation:

[0095]

[0096] In the formula, represents the inverse Fourier transform operation, and Z ′ represents the dot product convolution result H Z′ The low-dimensional feature representation after the inverse Fourier transform, that is, the above-mentioned target feature vector.

[0097] Finally, classify the fused target feature vector through a fully connected layer to obtain the output question answer.

[0098] Optionally, the model parameters of the global question answering model are updated based on the model parameters of multiple local question answering models and the target feature vector, including: determining the model parameters of the global question answering model, where the model parameters of the global question answering model are the average values of the model parameters of multiple local question answering models; determining the clustering results of the local models based on the model parameters of multiple local models and the target feature vector, where the clustering results are used to represent the similarity degree between each local question answering model and the global question answering model; determining the weight coefficients of multiple local question answering models based on the clustering results; updating the model parameters of the global question answering model based on the weight coefficients.

[0099] Among them, determining the clustering results of the local models based on the model parameters of multiple local models and the target feature vector includes: determining the clustering centers of multiple local question answering models based on the target feature vector, where the clustering centers are used to represent the average values of the model parameters of multiple local question answering models in the clustering clusters, and the clustering clusters are used to represent the set of model parameters of multiple local question answering models; determining the Euclidean distance between the model parameters of the global question answering model and the clustering centers, and allocating the model parameters of the global question answering model to the target clustering cluster, where the target clustering cluster is used to represent the clustering cluster where the clustering center corresponding to the minimum Euclidean distance is located; updating the clustering centers of the target clustering cluster to obtain the clustering results.

[0100] In the embodiments of the present application, the parameter update and optimization of the global question answering model aim to ensure that the global question answering model can accurately reflect the knowledge of all local question answering models. By constructing a federated transfer learning framework to jointly train multi-party data (that is, constructing a global question answering model by integrating the knowledge of multiple local question answering models), considering the similarity degree and contribution of each local question answering model, it can effectively improve the performance of the global question answering model and protect data privacy, realizing the effective fusion and utilization of multi-source data. The specific steps can be as follows:

[0101] S1: Initialize the model parameters of the global question answering model.

[0102] Suppose there are k local question-answering models, then initialize the model parameters of the global question-answering model as the average of the model parameters of all local question-answering models. The specific expression is as follows:

[0103]

[0104] In the formula, θ g represents the model parameters of the global question-answering model, k represents the total number of local question-answering models, and θ k represents the model parameters of the k-th local question-answering model.

[0105] S2: Migration clustering.

[0106] First, for each parameter in the model parameters of the global question-answering model, calculate its Euclidean distance from the clustering center of the local question-answering models. Here, the clustering center is used to represent the average value of the model parameters of multiple local question-answering models in the clustering cluster, and the clustering cluster is used to represent the set of model parameters of multiple local question-answering models. The specific expression is as follows:

[0107]

[0108] In the formula, β k represents the square of the Euclidean distance of the k-th local question-answering model, represents the i-th parameter in the model parameters of the global question-answering model, represents the clustering center of the k-th local question-answering model, represents the Euclidean distance between the model parameters of the global question-answering model and the clustering center, and ||.||2 represents the Euclidean norm.

[0109] Secondly, according to the distance metric, assign the i-th parameter in the model parameters of the global question-answering model to the clustering cluster corresponding to the clustering center of the local question-answering model with the closest distance, that is, the above target clustering cluster. The specific formula is as follows:

[0110]

[0111] In the formula, represents the target clustering cluster, and argmin i represents the assignment operation.

[0112] According to the new clustering result, update the clustering center of each current cluster. For example, for each i-th cluster, calculate the mean of all its samples as the new clustering center. The specific expression is as follows:

[0113]

[0114] In the formula, represents the clustering center of the current i-th cluster, Denote the clustering center of the \(i\)-th cluster after update, and \(\alpha\) represents the learning rate, which is used to control the amplitude of the update.

[0115] Repeat the above until the preset number of iterations is reached.

[0116] S3: Update the model parameters of the global Q&A model.

[0117] First, update the weight coefficients of each local Q&A model according to the above clustering results. The specific expression is as follows:

[0118]

[0119] In the formula, denotes the initial weight coefficient of the \(k\)-th local Q&A model, denotes the updated weight coefficient of the \(k\)-th local Q&A model. The SoftMax function is used to convert the weight coefficient into a probability distribution; exp represents the exponential function, which is used to calculate the exponential term in the SoftMax function.

[0120] Furthermore, update the model parameters of the global Q&A model according to the updated weight coefficients of the local Q&A model and the model parameters of all local Q&A models. The specific expression is as follows:

[0121]

[0122] In the formula, denotes the model parameters of the updated global Q&A model, and \(\theta\) k denotes the model parameters of the \(k\)-th local Q&A model.

[0123] Figure 3 is a framework structure diagram of an intelligent Q&A system according to an embodiment of the present application. As Figure 3 shown, the intelligent Q&A system includes multiple participating parties, their respective local (Q&A) models, and a global (Q&A) model. Among them, multiple participating parties (participating party 1, participating party 2, participating party 3, …, participating party K in the figure) represent different entities participating in federated learning, such as different enterprises, organizations, or data holders. Each participating party has its own local model (such as Figure 3 local model 1, local model 2, local model 3, …, local model K in), and this local model captures the characteristics and patterns of its own dataset and is used to process Q&A tasks related to its own data. The global model is trained through federated learning, that is, without directly sharing data, it updates its own model parameters using the parameters of all local models. This intelligent Q&A system allows multiple participating parties to jointly contribute to a more powerful global model while protecting their own data privacy, enabling the system to process more diverse data and improve the ability to understand and answer complex questions.

[0124] In the embodiments of the present application, based on existing artificial intelligence technologies such as federated learning, transfer learning, and natural language processing, an intelligent question answering method (based on federated transfer learning) is proposed, which makes full use of multi-modal multi-source data within an enterprise, solves the problem of missing feature data in a single data source, and thus enhances the data processing ability and knowledge richness of the system. At the same time, the characteristics of transfer learning are used to greatly reduce the training time of the intelligent question answering model, forming a perfect intelligent question system. Through the integration of accurate question classification and a targeted knowledge database, the question answering time of users can be effectively shortened and the accuracy of answers can be increased. In addition, through the federated learning framework, while protecting the privacy of enterprise data, the characteristics of multi-party data can be integrated, solving the problem of data islands, so that the system can obtain more useful external information.

[0125] According to the embodiments of the present application, an intelligent question answering device is provided. It should be noted that the intelligent question answering device in the embodiments of the present application can be used to execute the intelligent question answering method provided in the embodiments of the present application. The following introduces the intelligent question answering device provided in the embodiments of the present application.

[0126] Figure 4 is a structural diagram of an intelligent question answering device provided according to the embodiments of the present application. As Figure 4 shown, the device includes:

[0127] A receiving module 40, configured to receive the question information of the target object;

[0128] A generating module 42, configured to analyze and process the question information through a global question answering model to generate answer information corresponding to the question information, where the model parameters of the global question answering model are updated according to the model parameters of multiple local question answering models and a target feature vector, and the target feature vector is determined according to the first feature data output by multiple local question answering models and the second feature data in multiple knowledge databases. The first feature data is the question category after the local question answering model classifies the historical question answering information through a base classifier, and the second feature data is the multi-modal feature corresponding to the first feature data.

[0129] Through the receiving module 40 and the generating module 42 in the above intelligent question answering device, the purpose of improving the efficiency and accuracy of intelligent question answering is achieved, thereby realizing the technical effects of optimizing resource utilization, protecting data privacy, and solving the messy information in the knowledge database, and further solving the technical problems of long time consumption and low accuracy of intelligent question answering caused by the messy information in the knowledge database in the related art.

[0130] In the intelligent question-answering device provided by the embodiment of the present application, a determination module 44 is further included. The determination module is configured to classify historical question-answering information through a base classifier in a local question-answering model to obtain first feature data, where the historical question-answering information includes historical question information and historical answer information; determine second feature data in a target database according to the first feature data, where the target database is a knowledge database among multiple knowledge databases that contains the first feature data, and there is an association relationship between the first feature data and the second feature data; and determine a target feature vector according to the second feature data.

[0131] In the intelligent question-answering device provided by the embodiment of the present application, the determination module is further configured to determine the number of base classifiers; determine the question category of the historical question-answering information, where the question category includes at least one of the following: fact category, definition category, process category, comparison category, and suggestion category; and determine an encoding matrix according to the number of base classifiers and the question category, where the sum encoding information in the encoding matrix corresponds one-to-one with the question category.

[0132] In the intelligent question-answering device provided by the embodiment of the present application, the determination module is further configured to determine an output vector of the historical question-answering information through the base classifier; determine the Euclidean distance of the historical question-answering information according to the output vector and the row vector of the encoding matrix; determine the cosine similarity of the historical question-answering information according to the output vector and the row vector of the encoding matrix; and determine the first feature data according to the Euclidean distance and the cosine similarity.

[0133] In the intelligent question-answering device provided by the embodiment of the present application, the determination module is further configured to remove the noise data of the text feature and the image feature in the second feature data through discrete wavelet transform to obtain a first text feature vector and a first image feature vector; reduce the spatial dimensions of the first text feature vector and the first image feature vector through a hash function to obtain a second text feature vector and a second image feature vector; and fuse the second text feature vector and the second image feature vector through fast Fourier transform to obtain a target feature vector.

[0134] In the intelligent question-answering device provided by the embodiment of the present application, the determination module is further configured to determine the model parameters of the global question-answering model, where the model parameters of the global question-answering model are the average values of the model parameters of multiple local question-answering models; determine the clustering result of the local models according to the model parameters of the multiple local models and the target feature vector, where the clustering result is used to represent the similarity degree between each local question-answering model and the global question-answering model; determine the weight coefficients of the multiple local question-answering models according to the clustering result; and update the model parameters of the global question-answering model according to the weight coefficients.

[0135] In the intelligent question-answering device provided in the embodiments of the present application, the determination module is further configured to determine the clustering centers of multiple local question-answering models based on the target feature vectors, where the clustering centers are used to represent the average values of the model parameters of the multiple local question-answering models in the clustering clusters, and the clustering clusters are used to represent the set of model parameters of the multiple local question-answering models; determine the Euclidean distance between the model parameters of the global question-answering model and the clustering centers, and assign the model parameters of the global question-answering model to the target clustering cluster, where the target clustering cluster is used to represent the clustering cluster where the clustering center corresponding to the minimum Euclidean distance is located; update the clustering center of the target clustering cluster to obtain the clustering result.

[0136] The embodiments of the present application further provide an electronic device, including: a memory and a processor, where the memory is used to store program instructions; the processor is connected to the memory and is used to execute the above-mentioned intelligent question-answering method.

[0137] It should be noted that the above-mentioned electronic device is used to execute Figure 2 the intelligent question-answering method shown, so the relevant explanations in the above-mentioned intelligent question-answering method also apply to this electronic device, and will not be elaborated here.

[0138] The embodiments of the present application further provide a non-volatile storage medium, which includes a stored computer program, where the device where the non-volatile storage medium is located executes the above-mentioned intelligent question-answering method by running the computer program.

[0139] It should be noted that the above-mentioned non-volatile storage medium is used to execute Figure 2 the intelligent question-answering method shown, so the relevant explanations in the above-mentioned intelligent question-answering method also apply to this non-volatile storage medium, and will not be elaborated here.

[0140] The embodiments of the present application further provide a computer program product, including computer instructions, which implement the above-mentioned intelligent question-answering method when executed by a processor.

[0141] It should be noted that the above-mentioned computer program product is used to execute Figure 2 the intelligent question-answering method shown, so the relevant explanations in the above-mentioned intelligent question-answering method also apply to this computer program product, and will not be elaborated here.

[0142] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0143] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0144] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0145] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0147] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0148] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. An intelligent question-answering method, characterized in that: include: Receive question information from the target object; The question information is analyzed and processed by a global question-answering model to generate answer information corresponding to the question information, wherein model parameters of the global question-answering model are updated based on model parameters of multiple local question-answering models and a target feature vector, and the target feature vector is determined based on first feature data output by the multiple local question-answering models and second feature data in multiple knowledge databases, wherein the first feature data is a question category after the local question-answering model classifies historical question-answering information through a base classifier, and the second feature data is a multimodal feature corresponding to the first feature data.

2. The method according to claim 1, characterized in that The target feature vector is determined based on the first feature data output by the multiple local question-answering models and the second feature data in the multiple knowledge databases, including: Classifying the historical question and answer information by a base classifier in the local question and answer model to obtain the first feature data, wherein the historical question and answer information includes historical question information and historical answer information; Determine the second feature data in a target database according to the first feature data, wherein the target database is a knowledge database including the first feature data among the multiple knowledge databases, and the first feature data is associated with the second feature data; The target feature vector is determined according to the second feature data.

3. The method according to claim 2, characterized in that Before classifying the historical question-answer information by a base classifier in the local question-answer model, the method further includes: Determining the number of base classifiers; Determining the question category of the historical question and answer information, wherein the question category includes at least one of the following: fact category, definition category, process category, comparison category, and suggestion category; An encoding matrix is ​​determined according to the number of the base classifiers and the problem category, wherein the encoding information in the encoding matrix corresponds one-to-one to the problem category.

4. The method according to claim 3, characterized in that The historical question and answer information is classified by a base classifier in the local question and answer model to obtain the first feature data, including: Determining an output vector of the historical question-answer information by a base classifier; Determining the Euclidean distance of the historical question and answer information according to the output vector and the row vector of the encoding matrix; Determining the cosine similarity of the historical question and answer information according to the output vector and the row vector of the encoding matrix; The first feature data is determined according to the Euclidean distance and the cosine similarity.

5. The method according to claim 2, characterized in that: Determining the target feature vector according to the second feature data includes: Remove noise data of text features and image features in the second feature data by discrete wavelet transform to obtain a first text feature vector and a first image feature vector; Reducing the spatial dimensions of the first text feature vector and the first image feature vector by using a hash function to obtain a second text feature vector and a second image feature vector; The second text feature vector and the second image feature vector are fused by fast Fourier transform to obtain the target feature vector.

6. The method according to claim 1, characterized in that The model parameters of the global question answering model are updated according to the model parameters of the multiple local question answering models and the target feature vector, including: Determining model parameters of the global question-answering model, wherein the model parameters of the global question-answering model are average values ​​of the model parameters of the multiple local question-answering models; Determine a clustering result of the local model according to the model parameters of the multiple local models and the target feature vector, wherein the clustering result is used to indicate the degree of similarity between each local question-answering model and the global question-answering model; Determining weight coefficients of the multiple local question-answering models according to the clustering result; The model parameters of the global question-answering model are updated according to the weight coefficients.

7. The method according to claim 6, characterized in that Determining a clustering result of the local model according to the model parameters of the multiple local models and the target feature vector includes: Determine cluster centers of the multiple local question-answering models according to the target feature vector, wherein the cluster centers are used to represent average values ​​of model parameters of the multiple local question-answering models in cluster clusters, and the cluster clusters are used to represent a set of model parameters of the multiple local question-answering models; Determine the Euclidean distance between the model parameters of the global question answering model and the cluster center, and assign the model parameters of the global question answering model to a target cluster, wherein the target cluster is used to represent the cluster where the cluster center corresponding to the minimum value of the Euclidean distance is located; The cluster center of the target cluster is updated to obtain the clustering result.

8. An intelligent question-answering device, characterized in that: include: A receiving module, used for receiving question information from a target object; A generation module is used to analyze and process the question information through a global question and answer model to generate answer information corresponding to the question information, wherein the model parameters of the global question and answer model are updated based on the model parameters of multiple local question and answer models and a target feature vector, and the target feature vector is determined based on first feature data output by the multiple local question and answer models and second feature data in multiple knowledge databases, the first feature data is the question category after the local question and answer model classifies historical question and answer information through a base classifier, and the second feature data is a multimodal feature corresponding to the first feature data.

9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is used to store program instructions; The processor is connected to the memory and is used to execute the intelligent question and answer method described in any one of claims 1 to 7.

10. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the intelligent question-answering method described in any one of claims 1 to 7 by running the computer program.

11. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the intelligent question-answering method described in any one of claims 1 to 7 is implemented.