Multi-task Vector Retrieval Method and Device for Intelligent Conversation

Through a multi-task vector retrieval method for intelligent dialogue, the user identity and intention are obtained using vector alignment and generation models, the problems of inaccurate dialogue content and slow response speed in the existing intelligent dialogue system are solved, and high-quality and efficient intelligent dialogue are achieved.

CN119691093BActive Publication Date: 2025-07-01LONGSHINE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510209043.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-01
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The existing intelligent dialogue system relies on generative pre-trained transformers (GPTs) to generate dialogues, which have problems with insufficient knowledge static and generalization capabilities, resulting in inaccurate response content and slow response speed, making it difficult to provide high-quality dialogue output in dynamic, domain-specific or information-intensive scenarios.

Method used

The multi-task vector search method for intelligent dialogue is adopted to generate the current dialogue vector through the Embed model, and the trained insert vector alignment model and fine-tuned vector generation model are used to obtain the vector representation of user identity and user intention respectively. These vector representations as inputs generate high-quality conversation content through a generative pre-training transformer.

Benefits of technology

It improves the accuracy and response speed of conversation content, can accurately grasp the user's identity and intentions in dynamic scenarios, and provides high-quality intelligent dialogue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119691093B_ABST
    Figure CN119691093B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-task vector retrieval method and device for intelligent dialogue, which relates to the field of artificial intelligence technology. The method first converts the current dialogue language input by the user into a current dialogue vector through an Embed model; then, through a trained plug-in vector alignment model, a first vector for characterizing the current user identity is obtained, and based on the current dialogue language input by the user, a second vector for characterizing the user intention is obtained through a fine-tuning vector generation model. Finally, through a generative pre-trained transformer, the target dialogue language corresponding to the current dialogue language is output. The multi-task vector retrieval method for intelligent dialogue provided by the present invention can accurately and quickly output the target dialogue language, realizing high-quality intelligent dialogue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a multi-task vector retrieval method and device for intelligent dialogue. Background Art

[0002] In recent years, with the development of large-scale pre-trained language models (LLMs), natural language generation tasks based on deep learning have made great progress. At the same time, intelligent dialogue systems based on artificial intelligence have emerged and been applied to various industries, such as online intelligent customer service, offline intelligent furniture, and voice assistants.

[0003] However, existing intelligent dialogue systems only rely on language models such as Generative Pre-trained Transformer (GPT) to generate conversations. The conversation content generated by GPT has problems such as static knowledge and insufficient generalization ability, which in turn lead to inaccurate response content and slow response speed, and it is difficult to provide high-quality dialogue output in dynamic, domain-specific, or information-intensive scenarios. Summary of the Invention

[0004] The present invention provides a multi-task vector retrieval method and device for intelligent dialogue to solve the technical problems of inaccurate conversation content and slow conversation response speed in the prior art.

[0005] The present invention provides a multi-task vector retrieval method for intelligent dialogue, including the following steps:

[0006] Obtain the current conversation language input by the user;

[0007] Based on the current conversation language, generate a current conversation vector through an Embed model;

[0008] Based on the current conversation vector, obtain a first vector through a trained plug-in vector alignment model, and based on the current conversation language, obtain a second vector through a trained fine-tuning vector generation model; the first vector is used to represent the current user identity, and the second vector is used to represent the current user intention; the plug-in vector alignment model is trained based on first historical sample data; the fine-tuning vector generation model is trained based on second historical sample data;

[0009] Based on the first vector and the second vector, obtain a target conversation language corresponding to the current conversation language through a Generative Pre-trained Transformer.

[0010] A multi-task vector retrieval method for intelligent dialogue according to the present invention, the training steps of the plug-in vector alignment model include:

[0011] Obtain the first historical dialogue record and the corresponding historical user identity of the first historical dialogue record;

[0012] Based on the first historical dialogue record, generate a first historical dialogue vector through the Embed model, and, based on the historical user identity, generate a first historical target vector through the Embed model;

[0013] Use the first historical dialogue vector as the first historical training sample and the first historical target vector as the first historical sample label to construct the first historical sample data;

[0014] Based on the first historical sample data, perform supervised training on the plug-in vector alignment model to obtain a trained plug-in vector alignment model.

[0015] A multi-task vector retrieval method for intelligent dialogue according to the present invention, the plug-in vector alignment model includes any one of a multi-layer perceptron, an error backpropagation network model, and a radial basis function neural network model.

[0016] A multi-task vector retrieval method for intelligent dialogue according to the present invention, the training steps of the fine-tuning vector generation model include:

[0017] Obtain the second historical dialogue record;

[0018] Based on the second historical dialogue record, obtain the historical user intention through coreference resolution;

[0019] Based on the historical user intention, generate a second historical target vector through the Embed model;

[0020] Use the second historical dialogue record as the second historical training sample and the second historical target vector as the second historical sample label to construct the second historical sample data;

[0021] Based on the second historical sample data, perform supervised training on the BERT model to obtain a trained BERT model;

[0022] Use the trained BERT model as the fine-tuning vector generation model.

[0023] A multi-task vector retrieval method for intelligent dialogue according to the present invention, the loss function of the plug-in vector alignment model includes any one of a mean square error loss function and a cosine similarity loss function.

[0024] A multi-task vector retrieval method for intelligent dialogue provided by the present invention, the loss function of the fine-tuning vector generation model includes any one of the mean square error loss function and the cosine similarity loss function.

[0025] The present invention also provides a multi-task vector retrieval device for intelligent dialogue, including the following modules:

[0026] An acquisition module, configured to acquire the current dialogue language input by the user;

[0027] A first generation module, configured to generate a current dialogue vector through an Embed model based on the current dialogue language;

[0028] A second generation module, configured to obtain a first vector through a trained plug-in vector alignment model based on the current dialogue vector, and obtain a second vector through a trained fine-tuning vector generation model based on the current dialogue language; the first vector is used to represent the current user identity, and the second vector is used to represent the current user intention; the plug-in vector alignment model is trained based on the first historical sample data; the fine-tuning vector generation model is trained based on the second historical sample data;

[0029] A dialogue module, configured to obtain a target dialogue language corresponding to the current dialogue language through a generative pre-trained transformer based on the first vector and the second vector.

[0030] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, and when the processor executes the computer program, it implements the multi-task vector retrieval method for intelligent dialogue as described in any one of the above.

[0031] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the multi-task vector retrieval method for intelligent dialogue as described in any one of the above.

[0032] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the multi-task vector retrieval method for intelligent dialogue as described in any one of the above.

[0033] The multi-task vector retrieval method for intelligent conversation provided by the present invention first converts the current conversation language input by the user into a current conversation vector through an Embed model; then, through a trained plug-in vector alignment model, a first vector for characterizing the current user identity is obtained, realizing that in the task scenario of obtaining the user identity, the user input can be directly converted into a vector, and then vector retrieval can be performed through vector alignment, improving the accuracy and response speed of the vector retrieval process, so as to accurately and quickly obtain the user's identity information; and, based on the current conversation language input by the user, a second vector for characterizing the user intention is obtained through a fine-tuning vector generation model, thus realizing that in the task scenario of obtaining the user intention, key information can be directly extracted according to the current conversation language input by the user, improving the accuracy of vector retrieval and accurately and quickly obtaining the user intention; finally, through a generative pre-trained transformer, the target conversation language corresponding to the current conversation language is output. Since the generative pre-trained transformer receives the first vector and the second vector after processing the current conversation language input by the user, there is no need for further language processing and large-scale vector retrieval, so the target conversation language can be accurately and quickly output, realizing high-quality intelligent conversation. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a schematic flowchart of a multi-task vector retrieval method for intelligent conversation provided by the present invention.

[0036] Figure 2 It is a schematic flowchart of vector retrieval by a plug-in vector alignment model provided by the present invention.

[0037] Figure 3 It is a schematic flowchart of vector retrieval by a fine-tuning vector generation model provided by the present invention.

[0038] Figure 4 It is a schematic flowchart of the training process of a plug-in vector alignment model provided by the present invention.

[0039] Figure 5 It is a schematic structural diagram of a multi-task vector retrieval device for intelligent conversation provided by the present invention.

[0040] Figure 6 It is a schematic structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] Existing intelligent dialogue systems only rely on language models such as Generative Pre-trained Transformer (GPT) to generate conversations.

[0043] During the process of generating conversations, keyword retrieval, vector retrieval, hybrid retrieval, etc. can be performed through Retrieval-Augmented Generation (RAG) technology. These methods are fast and simple to implement, and are suitable for static documents or task scenarios with clear keywords, but lack the understanding of deep semantics and are difficult to handle synonyms or queries with complex semantics. Therefore, the conversation content generated by GPT has problems of knowledge staticization and insufficient generalization ability, which in turn leads to inaccurate response content and slow response speed, and it is difficult to provide high-quality conversation output in dynamic, domain-specific or information-intensive scenarios.

[0044] To solve the above problems, the present invention proposes a multi-task vector retrieval method and device for intelligent dialogue. The following will be combined with Figures 1 to 6 Describe the multi-task vector retrieval method and device for intelligent dialogue of the present invention.

[0045] Figure 1 is a schematic flowchart of a multi-task vector retrieval method for intelligent dialogue provided by the present invention. As Figure 1 shown, the method includes the following steps:

[0046] Step 101: Obtain the current conversation language input by the user.

[0047] Specifically, the specific content of the current conversation language input by the user can be asking questions, making suggestions, and putting forward requirements, etc.

[0048] Step 102: Generate a current conversation vector through an Embed model based on the current conversation language.

[0049] Specifically, the Embed model is an existing model that maps high-dimensional data to a low-dimensional space. The Embed model can map data to a low-dimensional vector space by learning the internal structure and semantic information of the data, so that similar data points are close to each other in the vector space.

[0050] In the embodiment of the present invention, the Embed model maps high-dimensional data (i.e., the current dialogue language) to a low-dimensional space, generates an embedding vector corresponding to the current dialogue language, that is, the current dialogue vector, and uses the current dialogue vector as the input of the plug-in vector alignment model.

[0051] Step 103: Based on the current dialogue vector, obtain a first vector through the trained plug-in vector alignment model, and based on the current dialogue language, obtain a second vector through the trained fine-tuning vector generation model; the first vector is used to represent the current user identity, and the second vector is used to represent the current user intention; the plug-in vector alignment model is trained based on the first historical sample data; the fine-tuning vector generation model is trained based on the second historical sample data.

[0052] Specifically, in an intelligent dialogue scenario, in order to improve the quality of the dialogue (such as accurately and quickly responding to user needs, etc.), it is necessary to accurately and quickly obtain the user identity and user intention according to the current dialogue language input by the user.

[0053] In the task scenario of obtaining the user identity, input the current dialogue vector into the trained plug-in vector alignment model, perform vector retrieval by means of vector alignment, obtain the user's identity information, and thus accurately and quickly output the first vector used to represent the current user identity.

[0054] The plug-in vector alignment model is trained based on the first historical sample data. During the training process, the plug-in vector alignment model can accurately match the actual identity of the user according to the dialogue language by learning the correlation features between the first historical dialogue record and the corresponding historical user identity.

[0055] Optionally, the loss function of the plug-in vector alignment model includes any one of the mean square error loss function and the cosine similarity loss function.

[0056] Figure 2 It is a schematic flow diagram of vector retrieval by a plug-in vector alignment model provided by the present invention, as Figure 2 shown.

[0057] After converting the current conversation language input by the user into a current conversation vector through the Embed model, through the retrieval-augmented generation technology, the current conversation vector is fitted and aligned with the vectors in the vector database, and the similarity between the current conversation vector and the vectors in the vector database is calculated through a preset loss function (such as mean squared error or cosine similarity, etc.), so as to output vectors with a similarity higher than the preset threshold, which match the user identity. Among them, the vector database is constructed by converting a variety of languages representing different user identities prepared in advance into corresponding vectors through the Embed model.

[0058] Since in the process of vector alignment, only the similarity between vectors needs to be calculated to achieve vector fitting, neither semantic relationships need to be understood nor natural language-vector interactions are required, thus greatly reducing the computational amount, simplifying the vector retrieval process, improving the response speed and accuracy, and accurately and quickly outputting the first vector used to represent the current user identity.

[0059] In the task scenario of obtaining the user intention, the current conversation language input by the user is input into the trained fine-tuned vector generation model, the key semantic information is extracted according to the current conversation language input by the user, and vector retrieval is performed, so as to accurately and quickly output the second vector used to represent the current user intention.

[0060] The fine-tuned vector generation model is trained based on the second historical sample data. During the training process, the fine-tuned vector generation model can accurately extract the actual intention of the user according to the conversation language by learning the correlation features between the second historical conversation record and the corresponding historical user intention.

[0061] Optionally, the loss function of the fine-tuned vector generation model includes any one of the mean squared error loss function and the cosine similarity loss function.

[0062] Figure 3 It is a schematic flow diagram of vector retrieval by a fine-tuned vector generation model provided by the present invention, as Figure 3 shown.

[0063] The current conversation language input by the user is input into the fine-tuned vector generation model, the key semantic information is extracted through coreference resolution, and the text data is encoded into dense vectors. These vectors contain the semantics and context information of the text. The similarity between these vectors can be calculated through a preset loss function (such as mean squared error or cosine similarity, etc.) to achieve more accurate information retrieval, accurately extract the user intention, and thus accurately and quickly output the second vector used to represent the current user intention.

[0064] In the embodiments of the present invention, the task of obtaining the user identity is performed by a pre-trained plug-in vector alignment model, and the task of obtaining the user intention is performed by a pre-trained fine-tuning vector generation model. According to the current conversation language input by the user, different tasks are performed in parallel by the two models, which improves the processing efficiency in multi-task scenarios. At the same time, through efficient vector retrieval, the user identity and user intention are obtained accurately and quickly, thus providing a basis for subsequent high-quality conversations.

[0065] Step 104: Based on the first vector and the second vector, a target conversation language corresponding to the current conversation language is obtained through a generative pre-trained transformer.

[0066] Specifically, the first vector and the second vector are jointly input into the generative pre-trained transformer, and the generative pre-trained transformer can quickly output a target conversation language corresponding to the current conversation language, realizing high-quality intelligent conversations.

[0067] Among them, the generative pre-trained transformer (GPT) is an existing language model based on artificial intelligence technology. GPT learns the statistical laws of language through pre-training on a large-scale corpus and can generate coherent and natural text. The GPT technology is based on the Transformer architecture in deep learning, and through unsupervised learning for pre-training and fine-tuning on specific tasks, it can achieve efficient and accurate language processing.

[0068] However, during the intelligent conversation process, GPT cannot accurately grasp the user identity and user intention, and there are problems such as knowledge staticization and insufficient generalization ability in the generated conversation content, which in turn leads to inaccurate response content and slow response speed.

[0069] The embodiments of the present invention improve from the input end of GPT. The task of obtaining the user identity is performed by a plug-in vector alignment model, and the task of obtaining the user intention is performed by a fine-tuning vector generation model, enabling GPT to accurately and timely grasp the changing user identity and user intention in dynamic conversation tasks. In addition, since the input received by GPT is the first vector representing the user identity and the second vector representing the user intention, there is no need to perform language processing and large-scale vector retrieval, so that a target conversation language corresponding to the current conversation language can be accurately and quickly output to have a high-quality intelligent conversation with the user.

[0070] The multi-task vector retrieval method for intelligent conversation provided by the present invention first converts the current conversation language input by the user into a current conversation vector through an Embed model; then, through a trained plug-in vector alignment model, a first vector for characterizing the current user identity is obtained, realizing that in the task scenario of obtaining the user identity, the user input can be directly converted into a vector, and then vector retrieval can be performed by means of vector alignment, improving the accuracy and response speed of the vector retrieval process, so as to accurately and quickly obtain the user's identity information; and, based on the current conversation language input by the user, a second vector for characterizing the user intention is obtained through a fine-tuning vector generation model, so as to realize that in the task scenario of obtaining the user intention, key information can be directly extracted according to the current conversation language input by the user, improving the accuracy of vector retrieval and accurately and quickly obtaining the user intention; finally, through a generative pre-trained transformer, the target conversation language corresponding to the current conversation language is output. Since the generative pre-trained transformer receives the first vector and the second vector processed from the current conversation language input by the user and does not need to perform language processing and large-scale vector retrieval, the target conversation language can be accurately and quickly output, realizing high-quality intelligent conversation.

[0071] Optionally, the training steps of the plug-in vector alignment model include:

[0072] Obtain the first historical conversation record and the corresponding historical user identity;

[0073] Based on the first historical conversation record, generate a first historical conversation vector through the Embed model, and, based on the historical user identity, generate a first historical target vector through the Embed model;

[0074] Construct the first historical sample data with the first historical conversation vector as the first historical training sample and the first historical target vector as the first historical sample label;

[0075] Based on the first historical sample data, perform supervised training on the plug-in vector alignment model to obtain a trained plug-in vector alignment model.

[0076] Specifically, Figure 4 is a schematic diagram of the training process of a plug-in vector alignment model provided by the present invention, as Figure 4 shown.

[0077] First, obtain the first historical conversation record and the corresponding historical user identity, that is, obtain the historical conversation records of a large number of users and the corresponding user identities in these historical conversation records. In this process, the historical conversation records of users with different identities on different chat platforms can be counted. For example, the historical conversation records of users with various identities such as primary and secondary school students, university students, primary and secondary school teachers, university teachers, engineers in various industries, self-employed individuals, farmers, and workers on public chat platforms (such as Weibo comment areas and movie review areas).

[0078] Then, convert the first historical conversation record into a first historical conversation vector through the Embed model, and convert the corresponding historical user identity into a first historical target vector through the Embed model.

[0079] Then, use all the first historical conversation vectors as the first historical training samples, and all the first historical target vectors as the first historical sample labels to construct the first historical sample data.

[0080] Finally, input the first historical sample data into the plug-in vector alignment model for supervised training. Through repeated iteration, update the model parameters to obtain a trained plug-in vector alignment model.

[0081] In the embodiment of the present invention, using all the first historical conversation vectors as the first historical training samples and all the first historical target vectors as the first historical sample labels makes the training data contain no natural language and only includes vectors. Thus, the trained plug-in vector alignment model can directly perform vector fitting alignment by calculating the similarity between vectors only, achieving vector fitting. It neither needs to understand semantic relationships nor requires interaction between natural language and vectors, reducing the computational amount, simplifying the vector retrieval process, and improving the response speed and accuracy.

[0082] Optionally, the plug-in vector alignment model includes any one of a multi-layer perceptron, an error backpropagation network model, and a radial basis function neural network model.

[0083] Specifically, since the plug-in vector alignment model neither needs to perform interaction between natural language and vectors nor needs to understand complex semantic relationships, and only needs to perform alignment fitting between vectors through similarity calculation, it can be a multi-layer feedforward neural network with a simple structure and a small number of parameters to perform the task of obtaining user intentions. For example, any one of a multi-layer perceptron, an error backpropagation network model, and a radial basis function neural network model.

[0084] In the plug-in vector alignment model in the embodiments of the present invention, through lightweight model design, not only the computational complexity of the model is reduced, the compatibility and adaptability to the deployment environment are improved, but also the response speed is further enhanced. At the same time, by simplifying the vector retrieval process and performing vector alignment, the accuracy of vector retrieval is improved.

[0085] Optionally, the training steps of the fine-tuning vector generation model include:

[0086] Obtain the second historical conversation record;

[0087] Based on the second historical conversation record, obtain the historical user intention through coreference resolution;

[0088] Based on the historical user intention, generate a second historical target vector through the Embed model;

[0089] Use the second historical conversation record as the second historical training sample and the second historical target vector as the second historical sample label to construct the second historical sample data;

[0090] Based on the second historical sample data, perform supervised training on the BERT model to obtain a trained BERT model;

[0091] Use the trained BERT model as the fine-tuning vector generation model.

[0092] Specifically, when training the fine-tuning vector generation model, it is first necessary to obtain the second historical conversation record, that is, to obtain a large number of historical conversation records of different users. In this process, the historical conversation records of users with different identities on different chat platforms can be counted, such as historical conversation records of primary and secondary school students, college students, primary and secondary school teachers, college teachers, various industry engineers, self-employed individuals, farmers, and workers, etc., on public chat platforms (such as Weibo comment areas and movie review areas, etc.).

[0093] Then, based on the second historical conversation record, accurately understand the entity reference in the user's question through coreference resolution to obtain the historical user intention.

[0094] For example, in some long conversations, in order to avoid repetition, pronouns, appellations, and abbreviations are usually used in the conversation to refer to the full name of the entity mentioned earlier. For example, at the beginning, it will be written as "a certain university", and later it will be abbreviated as "a certain big", and in subsequent conversations, "this university" and "it" will also be used to refer to the "a certain university" mentioned earlier. This phenomenon is the coreference phenomenon. Through coreference resolution, different descriptions of the same entity can be merged together, key information can be extracted from the long conversation, the user intention can be captured, and the long conversation can be converted into more concise language.

[0095] Then, the historical user intention is converted into a second historical target vector through the Embed model, that is, the concise language corresponding to the user intention is converted into the corresponding vector representation.

[0096] Then, all the second historical conversation records are used as the second historical training samples, and all the second historical target vectors are used as the second historical sample labels to construct the second historical sample data.

[0097] Then, the second historical sample data is input into the BERT model for supervised training. Through repeated iteration, the model parameters are updated to obtain a trained BERT model;

[0098] Finally, the trained BERT model is used as the fine-tuning vector generation model.

[0099] In the embodiment of the present invention, the Bidirectional Encoder Representations from Transformers (BERT) model can utilize the context information of words before and after simultaneously through bidirectional encoding, improving the semantic understanding ability. By training the BERT model with the second historical sample data, the model learns the correlation features between the second historical conversation records and the corresponding historical user intentions, and uses the trained BERT model as the fine-tuning vector generation model. Thus, the fine-tuning vector generation model can extract key semantic information according to the dialogue language input by the user, accurately extract the actual intention of the user, and thus accurately and quickly output the vector representation for characterizing the user intention, improving the accuracy of vector retrieval.

[0100] Next, the multi-task vector retrieval device for intelligent dialogue provided by the present invention will be described. The multi-task vector retrieval device for intelligent dialogue described below can be correspondingly referred to the multi-task vector retrieval method for intelligent dialogue described above.

[0101] Based on any of the above embodiments, Figure 5 is a schematic structural diagram of a multi-task vector retrieval device for intelligent dialogue provided by the present invention, as Figure 5 shown. The embodiment of the present invention provides a multi-task vector retrieval device for intelligent dialogue, including an acquisition module 501, a first generation module 502, a second generation module 503, and a dialogue module 504, where:

[0102] The acquisition module 501 is used to acquire the current dialogue language input by the user; the first generation module 502 is used to generate a current dialogue vector through the Embed model based on the current dialogue language; the second generation module 503 is used to obtain a first vector through the trained plug-in vector alignment model based on the current dialogue vector, and obtain a second vector through the trained fine-tuning vector generation model based on the current dialogue language; the first vector is used to represent the current user identity, and the second vector is used to represent the current user intention; the plug-in vector alignment model is trained based on the first historical sample data; the fine-tuning vector generation model is trained based on the second historical sample data; the dialogue module 504 is used to obtain the target dialogue language corresponding to the current dialogue language through the generative pre-trained transformer based on the first vector and the second vector.

[0103] The multi-task vector retrieval device for intelligent dialogue provided by the present invention first converts the current dialogue language input by the user into a current dialogue vector through the Embed model; then obtains a first vector used to represent the current user identity through the trained plug-in vector alignment model, realizing that in the task scenario of obtaining the user identity, the user input is directly converted into a vector, and then vector retrieval can be performed through vector alignment, improving the accuracy and response speed of the vector retrieval process, so as to accurately and quickly obtain the user's identity information; and, based on the current dialogue language input by the user, obtains a second vector used to represent the user intention through the fine-tuning vector generation model, so as to realize that in the task scenario of obtaining the user intention, key information is directly extracted according to the current dialogue language input by the user, improving the accuracy of vector retrieval and accurately and quickly obtaining the user intention; finally, the generative pre-trained transformer outputs the target dialogue language corresponding to the current dialogue language. Since the generative pre-trained transformer receives the first vector and the second vector after processing the current dialogue language input by the user, there is no need to perform language processing and large-scale vector retrieval, so the target dialogue language can be accurately and quickly output, realizing high-quality intelligent dialogue.

[0104] Figure 6 An example of the physical structure diagram of an electronic device is shown as Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the multi-task vector retrieval method for intelligent dialogue, and the method includes:

[0105] Obtain the current conversation language input by the user;

[0106] Based on the current conversation language, generate a current conversation vector through the Embed model;

[0107] Based on the current conversation vector, obtain a first vector through a trained plug-in vector alignment model, and based on the current conversation language, obtain a second vector through a trained fine-tuning vector generation model; the first vector is used to represent the current user identity, and the second vector is used to represent the current user intention; the plug-in vector alignment model is trained based on the first historical sample data; the fine-tuning vector generation model is trained based on the second historical sample data;

[0108] Based on the first vector and the second vector, obtain a target conversation language corresponding to the current conversation language through a generative pre-trained transformer.

[0109] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0110] On the other hand, the present invention also provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer can execute the multi-task vector retrieval method for intelligent conversation provided by the above-mentioned various methods, and this method includes:

[0111] Obtain the current conversation language input by the user;

[0112] Based on the current conversation language, generate a current conversation vector through the Embed model;

[0113] Based on the current dialogue vector, a first vector is obtained through a trained plug-in vector alignment model, and based on the current dialogue language, a second vector is obtained through a trained fine-tuning vector generation model; the first vector is used to represent the current user identity, and the second vector is used to represent the current user intention; the plug-in vector alignment model is trained based on first historical sample data; the fine-tuning vector generation model is trained based on second historical sample data;

[0114] Based on the first vector and the second vector, a target dialogue language corresponding to the current dialogue language is obtained through a generative pre-trained transformer.

[0115] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the multi-task vector retrieval method for intelligent dialogue provided by the above-mentioned various methods. The method includes:

[0116] Obtain the current dialogue language input by the user;

[0117] Based on the current dialogue language, generate a current dialogue vector through an Embed model;

[0118] Based on the current dialogue vector, a first vector is obtained through a trained plug-in vector alignment model, and based on the current dialogue language, a second vector is obtained through a trained fine-tuning vector generation model; the first vector is used to represent the current user identity, and the second vector is used to represent the current user intention; the plug-in vector alignment model is trained based on first historical sample data; the fine-tuning vector generation model is trained based on second historical sample data;

[0119] Based on the first vector and the second vector, a target dialogue language corresponding to the current dialogue language is obtained through a generative pre-trained transformer.

[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0122] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0123] It should be further noted that in the present invention, terms such as "target", "first", "second", etc. are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same category, and do not limit the number of objects. For example, the first object can be one or multiple.

[0124] "Determining B based on A" in the embodiments of the present application means that the factor A should be considered when determining B. It is not limited to "determining B only based on A", but also includes: "determining B based on A and C", "determining B based on A, C, and E", "determining C based on A and further determining B based on C", etc. In addition, it may also include using A as a condition for determining B. For example, "when A meets the first condition, use the first method to determine B"; for another example, "when A meets the second condition, determine B"; for yet another example, "when A meets the third condition, determine B based on the first parameter", etc. Of course, it may also be using A as a condition for the factor of determining B. For example, "when A meets the first condition, use the first method to determine C and further determine B based on C", etc.

[0125] The term "a plurality of" in the present invention means two or more, and other quantifiers are similar.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-task vector retrieval method for intelligent dialogue, characterized in that: include: Get the current dialog language entered by the user; Based on the current conversation language, generate a current conversation vector through an Embed model; Based on the current conversation vector, a first vector is obtained by using a trained insertion vector alignment model, and based on the current conversation language, a second vector is obtained by using a trained fine-tuning vector generation model; the first vector is used to represent the current user identity, and the second vector is used to represent the current user intention; The insertion-type vector alignment model is obtained by training based on the first historical sample data; the fine-tuning vector generation model is obtained by training based on the second historical sample data; Based on the first vector and the second vector, obtaining a target dialogue language corresponding to the current dialogue language through a generative pre-trained transformer; The training steps of the plug-in vector alignment model include: Obtaining a first historical conversation record and a historical user identity corresponding to the first historical conversation record; Based on the first historical conversation record, a first historical conversation vector is generated through an Embed model, and based on the historical user identity, a first historical target vector is generated through an Embed model; Using the first historical conversation vector as a first historical training sample and the first historical target vector as a first historical sample label, constructing first historical sample data; Based on the first historical sample data, supervised training is performed on the pluggable vector alignment model to obtain a trained pluggable vector alignment model; The training steps of the fine-tuning vector generation model include: Get the second historical conversation record; Based on the second historical conversation record, obtaining the historical user intention through coreference resolution; Based on the historical user intention, generating a second historical target vector through an Embed model; Using the second historical conversation record as a second historical training sample and the second historical target vector as a second historical sample label, constructing second historical sample data; Based on the second historical sample data, supervised training is performed on the BERT model to obtain a trained BERT model; Use the trained BERT model as a fine-tuned vector generation model.

2. The multi-task vector retrieval method for intelligent dialogue according to claim 1, characterized in that: The plug-in vector alignment model includes any one of a multi-layer perceptron, an error back propagation network model, and a radial basis function neural network model.

3. The multi-task vector retrieval method for intelligent dialogue according to claim 1, characterized in that: The loss function of the plug-in vector alignment model includes any one of a mean square error loss function and a cosine similarity loss function.

4. The multi-task vector retrieval method for intelligent dialogue according to claim 1, characterized in that: The loss function of the fine-tuning vector generation model includes any one of a mean square error loss function and a cosine similarity.

5. A multi-task vector retrieval device for intelligent dialogue, characterized in that: include: The acquisition module is used to obtain the current dialogue language input by the user; A first generating module, configured to generate a current dialogue vector through an Embed model based on the current dialogue language; A second generation module is configured to obtain a first vector based on the current dialogue vector by using a trained insertion-type vector alignment model, and to obtain a second vector based on the current dialogue language by using a trained fine-tuning vector generation model; the first vector is used to represent the current user identity, and the second vector is used to represent the current user intention; The insertion-type vector alignment model is obtained by training based on the first historical sample data; the fine-tuning vector generation model is obtained by training based on the second historical sample data; a dialogue module, configured to obtain a target dialogue language corresponding to the current dialogue language through a generative pre-trained transformer based on the first vector and the second vector; The training steps of the plug-in vector alignment model include: Obtaining a first historical conversation record and a historical user identity corresponding to the first historical conversation record; Based on the first historical conversation record, a first historical conversation vector is generated through an Embed model, and based on the historical user identity, a first historical target vector is generated through an Embed model; Using the first historical conversation vector as a first historical training sample and the first historical target vector as a first historical sample label, constructing first historical sample data; Based on the first historical sample data, supervised training is performed on the pluggable vector alignment model to obtain a trained pluggable vector alignment model; The training steps of the fine-tuning vector generation model include: Get the second historical conversation record; Based on the second historical conversation record, obtaining the historical user intention through coreference resolution; Based on the historical user intention, generating a second historical target vector through an Embed model; Using the second historical conversation record as a second historical training sample and the second historical target vector as a second historical sample label, constructing second historical sample data; Based on the second historical sample data, supervised training is performed on the BERT model to obtain a trained BERT model; Use the trained BERT model as a fine-tuned vector generation model.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the multi-task vector retrieval method for intelligent dialogue according to any one of claims 1 to 4 is implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-task vector retrieval method for intelligent dialogue as claimed in any one of claims 1 to 4 is implemented.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the multi-task vector retrieval method for intelligent dialogue as claimed in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Vertical class search method and device, search system and storage medium

    CN117290563A

  • Task-based dialogue method based on vector retrieval and large language model

    CN117829295A