Method, device, and computer-readable storage medium for obtaining title representation
Obtaining the problem representation in the online learning platform through generative models solves the problem that the problem representation of the problem in the prior art is not accurate enough, and improving the efficiency and accuracy of clustering and recommendation.
Patent Information
- Application Number
- CN202210641187.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-08
AI Technical Summary
In the clustering of questions and recommendations of similar questions, it is difficult to obtain more accurate and informative questions, which affects push efficiency and accuracy.
By obtaining the question description information and solution information, input the question representation generation model, semantic features are obtained using the mask language model layer, the text encoding layer and the pooling layer obtains classification and structural features, and the feature merging layer generates fusion features as the question representation.
It improves the efficiency of the generation of question representation, enhances the applicability of question representation, can express the question more fully, and improves the effect of clustering and recommendations of similar questions.
Smart Images

Figure CN115129849B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, and computer-readable storage medium for obtaining a question representation. Background Art
[0002] With the rapid development of online education, various online learning platforms can provide rich learning resources to information push objects such as students. Students can do exercises through the questions pushed by the online learning platform to master the knowledge content of learning. Generally, when an online learning platform or the like pushes resources to an information push object, it can obtain the question representations of each question to be pushed, and then perform question clustering and question similarity matching based on the question representations of each question to obtain recommended questions matched with each question, and push the recommended questions to students to provide online learning services. However, in question clustering and question similarity matching, how to obtain more accurate or information-rich question representations corresponding to each question is related to the efficiency or accuracy of question pushing, etc. How to obtain the question representations of each question has become one of the technical problems to be solved urgently. Summary of the Invention
[0003] The embodiments of this application provide a method, device, and computer-readable storage medium for obtaining a question representation, which can improve the generation efficiency of the question representation and enhance the applicability of the question representation.
[0004] In a first aspect, the embodiments of this application provide a method for obtaining a question representation, and the method includes:
[0005] Obtain the question description information and question answer information included in the target question, where the question description information includes the stem information and / or option information, and the question answer information includes the answer information and / or analysis information, and input the target question, the question description information, and the question answer information into a question representation generation model;
[0006] Through the masked language model layer in the question representation generation model, obtain the semantic features of the target question based on the input target question;
[0007] Obtain a first vector representation corresponding to the input question description information through the first text encoding layer and the first pooling layer in the question representation generation model, and obtain the question classification feature of the target question based on the first vector representation through the question classification layer in the question representation generation model;
[0008] Obtain a second vector representation corresponding to the input question answer information through the second text encoding layer and the second pooling layer, and obtain the question structure composition feature of the target question based on the first vector representation and the second vector representation;
[0009] The above-mentioned title representation generates the feature merging layer in the model. Based on the above semantic features, the above title classification features, and the features composed of the above title structures, the fused features of the above target title are generated as the title representation of the above target title. The above title representation is used for title clustering of target applications and / or recommendation of similar titles.
[0010] In a possible implementation, the above-mentioned masked language model layer includes a third text encoding layer and a masked classification layer; before the above-mentioned masked language model layer in the above-mentioned title representation generation model obtains the semantic features of the above-mentioned target title based on the input above-mentioned target title, the above method further includes:
[0011] Replace one or more target words in the above-mentioned target title with one or more masked labels, and carry the above one or more masked labels in the above-mentioned target title and input it into the above-mentioned title representation generation model.
[0012] The above-mentioned masked language model layer in the above-mentioned title representation generation model obtains the semantic features of the above-mentioned target title based on the input above-mentioned target title, including:
[0013] Obtain the word vectors corresponding to the above one or more masked labels through the above-mentioned third text encoding model to obtain the word vectors of the above one or more target words.
[0014] The above-mentioned prediction target words corresponding to the above one or more masked labels obtained by the above-mentioned masked classification layer based on the above word vectors are used as the semantic features of the above-mentioned target title.
[0015] In a possible implementation, the above-mentioned title representation generation model includes at least one of the above-mentioned title classification layers. The above-mentioned first text encoding layer and the first pooling layer in the above-mentioned title representation generation model obtain the first vector representation corresponding to the input above-mentioned title description information, and the above-mentioned title classification layer in the above-mentioned title representation generation model obtains the title classification features of the above-mentioned target title based on the above first vector representation, including:
[0016] Obtain the word vectors corresponding to each word in the above-mentioned title description information through the above-mentioned first text encoding model in the above-mentioned title representation generation model, and perform summation on the sequence dimension of the word vectors corresponding to each word through the above-mentioned first pooling layer in the above-mentioned title representation generation model to obtain the first vector representation corresponding to the above-mentioned title description information.
[0017] Obtain any classification corresponding to the above-mentioned target title through any of the above-mentioned title classification layers in the above-mentioned title representation generation model, obtain each classification obtained through each of the above-mentioned title classification layers, and obtain the title classification features corresponding to the above-mentioned title description information based on the above each classification.
[0018] In a possible implementation, the obtaining of the second vector representation corresponding to the above-mentioned problem solution information of the input through the second text encoding layer and the second pooling layer in the above-mentioned problem representation generation model includes:
[0019] Obtaining word vectors corresponding to each word in the above-mentioned problem solution information through the second text encoding model in the above-mentioned problem representation generation model, and performing summation on the sequence dimension of the word vectors corresponding to each word through the second pooling layer in the above-mentioned problem representation generation model to obtain the above-mentioned second vector representation.
[0020] In a possible implementation, before the above-mentioned obtaining of the problem description information and the problem solution information included in the above-mentioned target problem, the above-mentioned method further includes:
[0021] Obtaining a first loss function for the above-mentioned problem representation generation model to generate the semantic feature corresponding to the above-mentioned semantic feature based on multiple sample problems and the above-mentioned masked language model layer, obtaining a second loss function for the above-mentioned problem representation generation model to generate the problem classification feature corresponding to the above-mentioned problem classification feature based on the above-mentioned multiple sample problems and the above-mentioned first text encoding layer, the above-mentioned first pooling layer, and the above-mentioned problem classification layer, and obtaining a third loss function for the above-mentioned problem representation generation model to generate the problem structure composition feature corresponding to the above-mentioned problem structure composition feature based on the above-mentioned multiple sample problems and the above-mentioned first text encoding layer, the above-mentioned first pooling layer, the above-mentioned second text encoding layer, and the above-mentioned second pooling layer;
[0022] Performing weighted summation on the above-mentioned first loss function, the above-mentioned second loss function, and the above-mentioned third loss function to obtain a target loss function, and training the above-mentioned problem representation generation model based on the above-mentioned target loss function and the above-mentioned multiple sample problems, so that the above-mentioned masked language model layer of the above-mentioned problem representation generation model obtains the ability to obtain the semantic feature of any input target problem for any input target problem, so that the above-mentioned first text encoding layer, the first pooling layer, and the above-mentioned problem classification layer obtain the ability to obtain the problem classification feature of any input target problem for the problem description information of any input target problem, and so that the above-mentioned first text encoding layer, the above-mentioned first pooling layer, the above-mentioned second text encoding layer, and the above-mentioned second pooling layer obtain the ability to obtain the problem structure composition feature of any input target problem for the above-mentioned problem description information and the above-mentioned problem solution information of any input target problem.
[0023] In a possible implementation, each of the multiple sample problems at least includes sample problem description information and sample problem solution information, and the obtaining of the third loss function based on the above-mentioned multiple sample problems and the above-mentioned first text encoding layer, the above-mentioned first pooling layer, the above-mentioned second text encoding layer, and the above-mentioned second pooling layer includes:
[0024] Set the above sample question description information and the above sample question answer information in any sample question as the first training sample of any sample question, and pair the above sample question description information in any sample question with the remaining sample information in the above multiple sample questions in pairs to form the second training sample of any sample question. The above remaining sample information is other sample question answer information included in the above multiple sample questions except the above sample question answer information of any sample question.
[0025] Train the first text encoding layer, the first pooling layer, the second text encoding layer, and the second pooling layer based on the first training sample and the second training sample of each of the above sample questions to obtain the above third loss function.
[0026] In a possible implementation manner, after generating the fusion feature of the target question as the question representation of the target question based on the semantic feature, the question classification feature, and the question structure composition feature through the feature merging layer in the above question representation generation model, the method further includes:
[0027] Obtain the cosine similarity between the above fusion feature and the candidate recommendation features corresponding to each candidate recommendation question in multiple candidate recommendation questions, obtain the target recommendation feature from the multiple candidate recommendation features based on the cosine similarity between the above fusion feature and each candidate recommendation feature, and use the candidate recommendation question associated with the above target recommendation feature as the first candidate question;
[0028] Obtain a second candidate question with a text similarity to the above target question not less than a set threshold from the above multiple candidate recommendation questions through text similarity matching, and send the similar questions of the above target question obtained based on the above first candidate question and the above second candidate question to the target push object.
[0029] In a second aspect, an embodiment of the present application provides an apparatus for obtaining a question representation, and the apparatus includes:
[0030] An acquisition module, configured to obtain the question description information and the question answer information included in the above target question when receiving the target question. The above question description information includes stem information and / or option information, and the above question answer information includes answer information and / or analysis information. Input the above target question, the above question description information, and the above question answer information into the question representation generation model;
[0031] A semantic feature generation module, configured to obtain the semantic feature of the above target question through the masked language model layer in the above question representation generation model when the above target question is input into the above question representation generation model;
[0032] A question classification feature generation module, configured to obtain a first vector representation corresponding to the question description information through a first text encoding layer and a first pooling layer in the question representation generation model when the question description information is input into the question representation generation model, and obtain a question classification feature of the target question based on the first vector representation through a question classification layer in the question representation generation model.
[0033] A question structure composition feature generation module, configured to obtain a second vector representation corresponding to the question answer information through a second text encoding layer and a second pooling layer in the question representation generation model when the question answer information is input into the question representation generation model, and obtain a question structure composition feature of the target question based on the first vector representation and the second vector representation;
[0034] A question representation generation module, configured to generate a fusion feature of the target question as the question representation of the target question based on the semantic feature, the question classification feature, and the question structure composition feature through a feature merging layer in the question representation generation model.
[0035] In a third aspect, an embodiment of the present application provides a computer device, where the computer device includes: a processor, a memory, and a network interface;
[0036] The processor is connected to the memory and the network interface. Among them, the network interface is used to provide a data communication function, the memory is used to store program codes, and the processor is used to call the program codes to execute the method in the first aspect of the embodiment of the present application.
[0037] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the processor executes the program instructions, it executes the method in the first aspect of the embodiment of the present application.
[0038] In a fifth aspect, an embodiment of the present application provides a computer program product, where the computer program product includes a computer program, the computer program is stored in a computer-readable storage medium, and the computer program is adapted to be read and executed by a processor so that a computer device having the processor executes the method in the first aspect of the embodiment of the present application.
[0039] In this application, when a target question is received, the question description information and question answer information included in the target question are obtained. The question description information includes stem information and / or option information, and the question answer information includes answer information and / or analysis information. Then, the target question, the question description information, and the question answer information are input into a question representation generation model. The above target question can be input into the masked language model layer in the question representation generation model, and the semantic features of the target question can be obtained through the masked language model layer. Next, the above question description information can be sequentially input into the first text encoding layer and the first pooling layer in the question representation generation model, and the question classification features of the above target question can be obtained through the question classification layer based on the first vector representation output by the first pooling layer. Finally, the above question description information is sequentially input into the second text encoding layer and the second pooling layer in the question representation generation model, and the question classification features of the target question are obtained based on the second vector representation output by the second pooling layer and the above first vector representation. Through the feature merging layer in the question representation generation model, the above semantic features, the above question classification features, and the above question structure composition features are fused to generate the fusion features of the above target question as the question representation of the target question, so that the question representation includes question semantic information, question structure composition information, and question category information to more fully represent the corresponding question, and the generation efficiency of the question representation is high, enhancing the applicability of the question representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 is a schematic diagram of the system architecture provided by the embodiments of the present application;
[0042] Figure 2 is a schematic flowchart of the method for obtaining the question representation provided by the embodiments of the present application;
[0043] Figure 3 is a schematic diagram of the composition of the target question provided by the embodiments of the present application;
[0044] Figure 4 is a schematic diagram of the structure of the masked language model layer provided by the embodiments of the present application;
[0045] Figure 5 is a schematic diagram of the structure of the question representation generation model provided by the embodiments of the present application;
[0046] Figure 6 is a schematic flowchart of the generation process of the question classification features provided by the embodiments of the present application;
[0047] Figure 7 It is a schematic diagram of the process for generating the composition features of the question structure provided by the embodiments of the present application;
[0048] Figure 8 It is a schematic diagram of the process for generating the question representation provided by the embodiments of the present application;
[0049] Figure 9 It is a schematic diagram of the process for generating similar questions provided by the embodiments of the present application;
[0050] Figure 10 It is a schematic diagram of the structure of the device for obtaining the question representation provided by the embodiments of the present application;
[0051] Figure 11 It is a schematic diagram of the structure of the computer device provided by the embodiments of the present application. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0053] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0054] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0055] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0056] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0057] The solution provided in the embodiments of this application involves natural language processing and machine learning technologies in the field of artificial intelligence, and is specifically described through the following embodiments:
[0058] The method for obtaining problem representations provided by the embodiments of the present application (or simply referred to as the method provided by the embodiments of the present application) is applicable to generating corresponding problem representations based on problems in an application program (such as a learning application). The problem representations corresponding to each problem may include various information about the problem (such as problem semantics, problem category, etc.). Thus, problem data processing (which may be problem clustering, question bank construction, similar problem recommendation, etc.) can be performed based on the problem representations corresponding to each problem to enhance the learning effect of an object (which may be a student, etc.) using the learning application. For example, the learning application may push multiple recommended problems to an object (or the target push object). The above-mentioned multiple recommended problems may be multiple similar problems pushed based on one or more problems. The target push object may practice the multiple similar problems pushed by the above-mentioned learning application to further consolidate knowledge points. During the process of similar problem recommendation, the problem representations corresponding to the multiple problems to be recommended can be obtained, and some similar problems can be selected from the multiple problems to be recommended for pushing based on the multiple problem representations. However, in the usual process of generating problem representations, multiple dimensional information of the problem (such as problem semantic information, problem structure composition information, and problem category information, etc.) cannot be incorporated into the problem representation simultaneously, or the representation of various information is not sufficient (not explicitly incorporated, that is, some information is not in the optimized objective function), resulting in the inability to achieve the optimal clustering effect and recommendation effect when performing problem clustering and similar problem recommendation based on the generated problem representations. Therefore, information such as problem semantic information, problem structure composition information, and problem category information (which may be question type, problem difficulty, and problem knowledge points) can be incorporated into the problem representation to more fully represent the corresponding problem, so as to better achieve problem clustering and similar problem recommendation based on the problem representation, enhance the problem practice experience of the target push object in the learning application, and have good problem representation acquisition effect and strong applicability.
[0059] In the method provided by the embodiments of the present application, during the acquisition of the question representation, a question (which can be called the target question) for generating the corresponding question representation can be received, and the question description information and the question answer information in the above target question can be obtained. Here, the question description information may include the stem information of the target question. When the target question is a multiple-choice question, the question description information may include the stem information and the option information. The question answer information may include the answer information and the analysis information of the target question. During the acquisition of the question representation, the above target question can be input into the masked language model layer in the question representation generation model, and the semantic features of the target question can be obtained through the masked language model layer. During the acquisition of the question representation, the above question description information (which may include the stem information, or the stem information and the option information) can be sequentially input into the first text encoding layer and the first pooling layer in the question representation generation model, and the question classification features of the above target question (which can be the question classification features generated based on the classification results corresponding to the question type, the question difficulty, and the question knowledge points respectively) can be obtained through the question classification layer based on the output of the first pooling layer (which can be the first vector representation). Finally, the above question description information is sequentially input into the second text encoding layer and the second pooling layer in the question representation generation model, and the question classification features of the above target question are obtained based on the output of the second pooling layer (which can be the second vector representation) and the above first vector representation. Through the feature merging layer in the question representation generation model, the above semantic features, the above question classification features, and the above question structure composition features are fused to generate the fused features of the above target question as the question representation of the target question, so that the question representation includes question semantic information, question structure composition information, and question category information to more fully represent the corresponding question, and the question representation generation effect is good. In addition, by sharing some model layers (such as the first text encoding layer and the first pooling layer) in the question representation generation model during the process of obtaining the question classification features and the question structure composition features, the efficiency of obtaining the question representation through the question representation generation model can be further improved, as well as the training effect and training efficiency of training the question representation generation model (such as multi-task pre-training) to obtain the ability to generate question representations.
[0060] See Figure 1 , Figure 1 is the schematic diagram of the system architecture provided by the embodiments of the present application. As Figure 1As shown, the system architecture may include a business server 100 and a terminal cluster. The terminal cluster may include terminal devices such as terminal device 200a, terminal device 200b, terminal device 200c, …, terminal device 200n. Among them, the above-mentioned business server 100 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal devices (including terminal device 200a, terminal device 200b, terminal device 200c, …, terminal device 200n) may be intelligent terminals such as smart phones, tablet computers, laptop computers, desktop computers, handheld computers, mobile internet devices (MIDs), wearable devices (such as smart watches, smart bracelets, etc.), intelligent computers, and intelligent vehicles. Among them, the business server 100 can establish communication connections with each terminal device in the terminal cluster, and communication connections can also be established between the terminal devices in the terminal cluster. In other words, the business server 100 can establish communication connections with each of the terminal devices such as terminal device 200a, terminal device 200b, terminal device 200c, …, terminal device 200n. For example, a communication connection can be established between terminal device 200a and the business server 100. A communication connection can be established between terminal device 200a and terminal device 200b, and a communication connection can also be established between terminal device 200a and terminal device 200c. Among them, the above-mentioned communication connection is not limited to the connection method, and can be directly or indirectly connected through wired communication methods, or can be directly or indirectly connected through wireless communication methods, etc., which can be specifically determined according to the actual application scenario, and this application does not make any restrictions here.
[0061] It should be understood that as Figure 1 shown, each terminal device in the terminal cluster may be installed with an application client. When the application client runs on each terminal device, it can respectively communicate with the above-mentioned Figure 1Data interaction is carried out between the business servers 100 shown, enabling the business servers 100 to receive business data from each terminal device, or the business servers 100 to push business data (such as similar questions) to each terminal device. Among them, the above application client can be a learning application, a social application, an instant messaging application, a live broadcast application, a news application, a short video application, a video application, a music application, a shopping application, a novel application, a payment application, etc., which have the function of displaying data information such as text, images, and videos. Specifically, it can be determined according to the actual application scenario requirements and is not limited here. Among them, the application client can be an independent client or an embedded sub-client integrated in a certain client (such as an instant messaging client, a social client, etc.). Specifically, it can be determined according to the actual application scenario and is not limited here. Taking the learning application as an example, when the target push object uses the learning application through the terminal device, the target push object can view and practice the questions in the learning application through the terminal device. The business server 100, as the server of the learning application, can be a collection of multiple servers including the background server corresponding to the application client, the data processing server, etc. The business server 100 can receive business data from the terminal device (for example, a similar question recommendation instruction based on the target question sent by the target push object through the terminal device), generate a corresponding question representation based on the target question, and thus obtain the corresponding similar questions from multiple candidate questions, and return the above similar questions to the above terminal device to recommend and display them to the target push object through the installed learning application. The method provided by the embodiments of the present application can be executed by a business server 100 as shown in Figure 1 , or can be executed by a terminal device (such as any one of the terminal devices 200a, 200b,..., 200n shown in Figure 1 ), or can also be jointly executed by the terminal device and the business server. Specifically, it can be determined according to the actual application scenario and is not limited here.
[0062] In some feasible embodiments, the service server 100 may obtain a target question and acquire the question description information and question answer information in the above-mentioned target question. Here, the question description information may include the stem information of the target question. When the target question is a multiple-choice question, the question description information may include the stem information and option information. The question answer information may include the answer information and analysis information of the target question. A question representation generation model may be deployed in the service server 100, that is, the service server 100 may input the above-mentioned target question into the masked language model layer in the question representation generation model, and obtain the semantic features of the target question through the masked language model layer. The service server 100 may also input the above-mentioned question description information (which may include the stem information, or the stem information and option information) into the first text encoding layer and the first pooling layer in the question representation generation model in sequence, and obtain the question classification features of the above-mentioned target question (which may be question classification features generated based on the classification results corresponding to the question type, question difficulty, and question knowledge points respectively) through the question classification layer based on the output of the first pooling layer (which may be the first vector representation). The service server 100 inputs the above-mentioned question description information into the second text encoding layer and the second pooling layer in the question representation generation model in sequence, and obtains the question classification features of the above-mentioned target question based on the output of the second pooling layer (which may be the second vector representation) and the above-mentioned first vector representation. The service server 100 fuses the above-mentioned semantic features, the above-mentioned question classification features, and the above-mentioned question structure composition features through the feature merging layer in the question representation generation model to generate the fusion features of the above-mentioned target question as the question representation of the target question, so that the question representation includes question semantic information, question structure composition information, and question category information to more fully represent the corresponding question, and the question representation generation effect is good. In addition, the service server 100 shares some model layers (such as the first text encoding layer and the first pooling layer) in the above-mentioned question representation generation model during the process of obtaining the question classification features and the question structure composition features, which can further improve the efficiency of obtaining the question representation through the question representation generation model, as well as the training effect and training efficiency of training the question representation generation model (such as multi-task pre-training) to obtain the ability to generate question representations. The service server 100 may obtain multiple candidate recommendation features corresponding to multiple candidate recommended questions, and acquire target recommendation features from the above-mentioned multiple candidate recommendation features based on the above-mentioned fusion features, and obtain a part of candidate questions (which may be the first candidate questions) based on the above-mentioned target recommendation features. The service server 100 may also obtain second candidate questions through text similarity matching from the above-mentioned multiple candidate recommended questions based on the target question, and select some candidate questions from the above-mentioned first candidate questions and the above-mentioned second candidate questions as the similar questions of the target question, and send the similar questions to each terminal device to display to the target push object, and the similar question recommendation effect is good, enhancing the recommendation effect of similar question recommendation.
[0063] In some feasible embodiments, it may be that the terminal device 200a obtains the target question through the application client (such as a learning application) installed thereon, and obtains the question description information and the question answer information in the above-mentioned target question. Here, the question description information may include the stem information of the target question. When the target question is a multiple-choice question, the question description information may include the stem information and the option information. The question answer information may include the answer information and the analysis information of the target question. A question representation generation model may be deployed in the terminal device 200a, that is, the terminal device 200a may input the above-mentioned target question into the masked language model layer in the question representation generation model, and obtain the semantic features of the target question through the masked language model layer. The terminal device 200a may also input the above-mentioned question description information (which may include the stem information, or the stem information and the option information) into the first text encoding layer and the first pooling layer in the question representation generation model in sequence, and obtain the question classification features of the above-mentioned target question (which may be question classification features generated based on the classification results corresponding to the question type, question difficulty, and question knowledge points respectively) through the question classification layer based on the output of the first pooling layer (which may be the first vector representation). The terminal device 200a inputs the above-mentioned question description information into the second text encoding layer and the second pooling layer in the question representation generation model in sequence, and obtains the question classification features of the above-mentioned target question based on the output of the second pooling layer (which may be the second vector representation) and the above-mentioned first vector representation. The terminal device 200a fuses the above-mentioned semantic features, the above-mentioned question classification features, and the above-mentioned question structure composition features through the feature merging layer in the question representation generation model to generate the fusion features of the above-mentioned target question as the question representation of the target question, so that the question representation includes question semantic information, question structure composition information, and question category information to more fully represent the corresponding question, and the question representation generation effect is good. In addition, the terminal device 200a shares some model layers (such as the first text encoding layer and the first pooling layer) in the above-mentioned question representation generation model during the process of obtaining the question classification features and the question structure composition features, which can further improve the efficiency of obtaining the question representation through the question representation generation model, as well as the training effect and training efficiency of training the question representation generation model (such as multi-task pre-training) to obtain the ability to generate question representations. The terminal device 200a may obtain multiple candidate recommendation features corresponding to multiple candidate recommended questions, and obtain the target recommendation features from the above-mentioned multiple candidate recommendation features based on the above-mentioned fusion features, and obtain a part of the candidate questions (which may be the first candidate questions) based on the above-mentioned target recommendation features. The terminal device 200a may also obtain the second candidate questions from the above-mentioned multiple candidate recommended questions through text similarity matching based on the target question, and select some candidate questions from the above-mentioned first candidate questions and the above-mentioned second candidate questions as the similar questions of the target question to be displayed to the target push object, and the similar question recommendation effect is good, enhancing the recommendation effect of similar question recommendation.
[0064] For ease of description, the terminal device will be used as the execution subject of the method provided in the embodiments of the present application, and the method for obtaining the topic representation through the terminal device will be specifically described through an embodiment.
[0065] See Figure 2 , Figure 2 which is a schematic flowchart of the method for obtaining the topic representation provided in the embodiments of the present application. As Figure 2 shown, the method includes the following steps:
[0066] S101, when receiving a target topic, obtain the topic description information and topic solution information included in the target topic, and input the target topic, the topic description information, and the topic solution information into the topic representation generation model.
[0067] In some feasible implementation manners, the terminal device (such as the terminal device 200a) can receive the target topic, and the target topic can be obtained based on a similar topic recommendation instruction sent by an application client (such as a learning application) loaded in the terminal device for a target push object. When receiving the target topic, the terminal device can obtain the topic description information and topic solution information included in the target topic. Specifically, the above topic description information may include the stem information of the target topic. When the target topic is a multiple-choice question, the above topic description information may include the stem information and option information, and the above topic solution information may include the answer information and analysis information of the target topic. For example, please see Figure 3 , Figure 3 which is a schematic diagram of the composition of the target topic provided in the embodiments of the present application. As Figure 3 shown, the topic description information of the target topic is: "Among the following four propositions, which one is wrong? (); A), A parallelogram with a set of adjacent angles equal is a rectangle; B), A quadrilateral with three angles all equal is a rectangle; C), A quadrilateral with a set of opposite sides equal and a set of opposite angles being right angles is a rectangle; D), A quadrilateral with diagonals bisecting each other and being equal is a rectangle." Among them, the target topic is a multiple-choice question, and the topic description information includes the stem information: "Among the following four propositions, which one is wrong? ()" and the option information: "A), A parallelogram with a set of adjacent angles equal is a rectangle; B), A quadrilateral with three angles all equal is a rectangle; C), A quadrilateral with a set of opposite sides equal and a set of opposite angles being right angles is a rectangle; D), A quadrilateral with diagonals bisecting each other and being equal is a rectangle." See again Figure 3, in the solution information of the target question, the answer information includes: "Answer: B", and the analysis information includes: "A. Since quadrilateral ABCD is a parallelogram, then AD∥BC, so ∠A + ∠B = 180°. Since ∠A = ∠B, then ∠A = ∠B = 90°, so parallelogram ABCD is a rectangle. This option is incorrect. B. Since ∠A = ∠B = ∠C, it cannot be proved that ∠D is equal to ∠A, ∠B, and ∠C. This option is correct. C. Connect AC. Since ∠B = ∠D = 90°, AC = AC, and AB = CD, then Rt△ABC≌Rt△CDA (HL), so AD = BC, and parallelogram ABCD is a rectangle. This option is incorrect. D. Since OA = OC and OB = OD, then parallelogram ABCD is a parallelogram. Since AC = BD, then parallelogram ABCD is a rectangle. This option is incorrect. Therefore, the answer is B." The terminal device can input the received target question into the question representation generation model, and at the same time input the above question description information and question solution information into the question representation generation model, and obtain the question representation of the target question through the question representation generation model. Optionally, when inputting the above question description information and question solution information into the question representation generation model, it can be input into the question representation generation model in the format of "stem information + option information + answer information + analysis information". If some information is missing (such as the target question does not include option information), the corresponding information is set to empty to improve the data input efficiency.
[0068] S102, obtain the semantic features of the target question based on the input target question through the masked language model layer in the question representation generation model.
[0069] In some feasible embodiments, the terminal device can input the target question into the masked language model layer in the question representation generation model to obtain the semantic features of the target question through the above masked language model layer. Specifically, the above masked language model layer can include a text encoding layer (which can be called the third text encoding layer) and a masked classification layer. Please refer to Figure 4 , Figure 4 is the schematic diagram of the masked language model layer structure provided by the embodiment of the present application. As Figure 4 shown, the masked language model layer in the question representation generation model includes a third text encoding layer and a masked classification layer. The terminal device can obtain semantic features based on the input target question through the third text encoding layer and the masked classification layer in the masked language model layer. Optionally, the third text encoding layer in the above masked language model layer can be an independent text encoding layer in the above question representation generation model, or can be composed of other text encoding layers in the question representation generation model. For the convenience of description, the embodiment of the present application takes the third text encoding layer in the masked language model layer being composed of other text encoding layers in the question representation generation model as an example for illustration. Please refer to Figure 5 , Figure 5It is a schematic structural diagram of a topic representation generation model provided by an embodiment of the present application. As Figure 5 shown, the masked language model layer in the topic representation generation model includes a third text encoding layer and a masked classification layer. The third text encoding layer is jointly composed of a first text encoding layer and a second text encoding layer. That is to say, in the process of obtaining the semantic features of the target topic through the masked language model layer in the topic representation generation model, some text encoding layers (the first text encoding layer and the second text encoding layer) are shared with the process of obtaining other features of the target topic (such as topic classification features, topic structure composition features) through the topic representation generation model. Thus, the efficiency of obtaining the topic representation through the topic representation generation model can be further improved, as well as the training effect and training efficiency of training the topic representation generation model (such as multi-task pre-training) to obtain the ability to generate topic representations.
[0070] In some feasible embodiments, before the terminal device inputs the above target device into the masked language model layer in the topic representation generation model, the terminal device can randomly select some words (which can be called target words) in the above target topic and replace them with corresponding masked labels (or called mask labels), and input the target topic with the above masked labels into the topic representation generation model. For example, please refer to Figure 5 again. The terminal device can replace the target words in the topic description information (which can include stem information, or stem information and option information) in the target topic with masked labels, and replace the target words in the topic solution information (which can include answer information and analysis information) in the target topic with masked labels, and input the target topic with masked labels into the third text encoding model in the topic representation generation model (the third text encoding model is shared with the first text encoding model and the second text encoding model, that is, the topic description information with masked labels can be input into the first text encoding model, and the topic solution information with masked labels can be input into the second text encoding model). The terminal device obtains the word vectors corresponding to the above masked labels through the third text encoding model, and inputs the word vectors generated by the third text encoding model into the masked classification layer to predict (softmax activation can be used) the target words corresponding to the above masked labels (that is, the replaced target words) through the masked classification layer. The predicted target words corresponding to the above masked labels output by the masked classification layer contain the semantic information of the above target topic. Thus, the predicted target words corresponding to the above masked labels are used as the semantic features of the target topic. The semantic features obtained through the above masked language model layer can extract the semantic information of the above target topic, so that a more sufficient topic representation can be obtained based on the semantic features, and the semantic feature extraction effect is good.
[0071] In some feasible embodiments, the above-mentioned first text encoding layer and second text encoding layer may be text encoding layers based on the albert-tiny text encoding model, text encoding layers based on the albert text encoding model, and text encoding layers based on the bert text encoding model, which can be specifically determined according to the actual application scenario, and the present application does not limit this here. Optionally, the above-mentioned first text encoding model and second text encoding model may use the same model encoding (i.e., share parameters).
[0072] S103. Obtain a first vector representation corresponding to the input question description information through the first text encoding layer and the first pooling layer in the question representation generation model, and obtain the question classification feature of the target question based on the first vector representation through the question classification layer in the question representation generation model.
[0073] In some feasible embodiments, the terminal device may pass through the above-mentioned question representation generation model (the structure of the question representation generation model can be seen in Figure 5 ) The first text encoding layer and the first pooling layer (which may be a sum-pooling layer) in obtain a corresponding vector representation (which may be a first vector representation) based on the question description information of the above-mentioned target question, and obtain the question classification feature of the target question based on the first vector representation through the question classification layer in the question representation generation model. Specifically, the above-mentioned question representation generation model may include one or more question classification layers, and each question classification layer may correspond to different question classification tasks, so as to obtain different classifications based on the first vector representation through each question classification layer, and one question classification layer generates one classification of the target question. Please refer to Figure 6 , Figure 6 is a schematic diagram of the question classification feature generation process provided by the embodiment of the present application. As Figure 6 shown, Figure 6It includes topic classification layer 1, topic classification layer 2, and topic classification layer 3. Among them, topic classification layer 1 can correspond to the classification of question difficulty (a target question can correspond to a difficulty level, and the difficulty levels can include easy, relatively easy, medium, relatively difficult, difficult, etc.). Topic classification layer 2 can correspond to the classification of question types (a target question can correspond to a question type, which can include multiple-choice questions, fill-in-the-blank questions, true-or-false questions, etc.). Topic classification layer 3 can correspond to the classification of question knowledge points (a target question can include multiple knowledge point labels, such as mathematics, quadratic equations of one variable, systems of equations, etc.). The classification tasks corresponding to the above topic classification layer 1 and topic classification layer 2 can be called single-classification tasks, and the classification task corresponding to topic classification layer 3 can be called a multi-classification task. For single-classification tasks, the classification result can be in the form of one-hot encoding. For multi-classification tasks, the classification result can be in the form of multi-hot encoding and then normalized so that the sum is 1. For example, in the classification of question knowledge points, [0.5, 0, 0, 0.5, 0, …] represents the knowledge point classification including knowledge point 1 and knowledge point 4. The terminal device can Figure 6 obtain the word vectors corresponding to each word in the question description information of the target question through the first text encoding model in it, and perform summation on the sequence dimension of the word vectors corresponding to each word through the first pooling layer to obtain the first vector representation. Based on the above first vector representation, the above multiple topic classification layers (topic classification layer 1, topic classification layer 2, and topic classification layer 3) respectively obtain multiple classifications corresponding to the target question (which can include the classification of question difficulty, the classification of question types, and the classification of question knowledge points), and use the above multiple classifications as the topic classification features of the target question. Through the multiple topic classification layers, topic classification features that more fully reflect the characteristics of the target question can be obtained. The classification information contained in the topic classification features is rich, so that a more sufficient question representation can be obtained based on the topic classification features.
[0074] S104, obtain the second vector representation corresponding to the input question answer information through the second text encoding layer and the second pooling layer in the question representation generation model, and obtain the question structure composition feature of the target question based on the first vector representation and the second vector representation.
[0075] In some feasible implementation manners, the terminal device can, through the second text encoding layer and the second pooling layer (which can be a sum-pooling layer) in the above question representation generation model (for the structure of the question representation generation model, see Figure 5 ), obtain the corresponding vector representation (which can be the second vector representation) based on the question description information of the target question, and obtain the question structure composition feature of the target question based on the first vector representation obtained by the above first text encoding layer and the first pooling layer based on the question description information of the target question and the second vector representation. Please see Figure 7 ,Figure 7 is a schematic diagram of the generation process of the topic structure composition feature provided by an embodiment of the present application. As Figure 7 shown, the terminal device can obtain the word vectors corresponding to the words in the topic description information of the target topic through the Figure 7 first text encoding model in, and perform summation on the word vectors corresponding to the words in the sequence dimension through the above-mentioned second pooling layer to obtain a second vector representation. The terminal device can also obtain the word vectors corresponding to the words in the topic answer information of the target topic through the Figure 7 second text encoding model in, and perform summation on the word vectors corresponding to the words in the sequence dimension through the above-mentioned second pooling layer to obtain a second vector representation. Finally, based on the above first vector representation and the second vector representation, the topic structure composition feature of the target topic is obtained. By obtaining the first vector representation from the above topic description information and the second vector representation from the topic answer information to obtain the topic structure composition feature, the topic structure composition feature includes the matching relationship between the topic description information and the topic answer information in the target topic, so as to obtain a topic structure composition feature that more fully reflects the target topic structure composition, and thus a more accurate topic representation can be obtained based on this topic structure composition feature. In addition, the Figure 7 first text encoding model and the first pooling layer in can be shared with the Figure 6 first text encoding model and the first pooling layer in above (for example, see the first text encoding model and the first pooling layer in the topic representation generation model structure shown in Figure 5 ), that is, the first vector representation in the above topic classification feature generation process can be directly used in the topic structure composition feature generation process, thereby improving the feature generation efficiency and having a good topic structure composition feature generation effect.
[0076] S105. Generate a fusion feature of the target topic as the topic representation of the target topic based on the semantic feature, the topic classification feature, and the topic structure composition feature through the feature merging layer in the topic representation generation model.
[0077] In some feasible implementation manners, the terminal device can obtain the fusion feature of the target topic based on the above semantic feature, the topic classification feature, and the topic structure composition feature through the feature merging layer in the above topic representation generation model, and use this fusion feature as the topic representation of the target topic. Specifically, please refer to Figure 5 , Figure 5The title shown indicates that the feature merging layer in the generation model can receive semantic features output from the mask classification layer, title classification features output from multiple title classification layers (title classification layer 1, title classification layer 2, and title classification layer 3), and title structure composition features output from the first pooling layer and the second pooling layer. The feature merging layer can merge the features to obtain a fused feature as the title representation of the target title. The above-mentioned feature merging layer can obtain the fused feature by summing the semantic features, title classification features, and title structure composition features in the sequence dimension. As the title representation of the target title, the fused feature contains information in multiple dimensions of the target title, including semantic information, classification information (which can also be called metadata information and can include title difficulty, title type, and title knowledge points), and title structure composition information, enabling the title representation to more fully represent the target title and achieving good title representation generation effects.
[0078] In some feasible embodiments, before the terminal device obtains the question description information and question answer information included in the above target question, it may also obtain a plurality of sample questions, and obtain a first loss function, a second loss function, and a third loss function based on the above plurality of sample questions through a question representation generation model. Specifically, the terminal device may obtain the first loss function based on the plurality of sample questions and the above masked language model layer, obtain the second loss function based on the plurality of sample questions and the above first text encoding layer, the above first pooling layer, and the above question classification layer, and obtain the third loss function based on the plurality of sample questions and the above first text encoding layer, the above first pooling layer, the above second text encoding layer, and the above second pooling layer. The terminal device may perform weighted summation on the above first loss function, the above second loss function, and the above third loss function to obtain a target loss function, and train the above question representation generation model based on the above target loss function and the above plurality of sample questions, so that the masked language model layer of the question representation generation model obtains the ability to obtain the semantic features of any target question for any input target question, so that the first text encoding layer, the first pooling layer, and the question classification layer obtain the ability to obtain the question classification features of any target question for the question description information of any input target question, and enable the first text encoding layer, the first pooling layer, the second text encoding layer, and the second pooling layer to obtain the ability to obtain the question structure composition features of any target question for the question description information and question answer information of any input target question. By performing weighted summation on the loss functions corresponding to the generation of each feature (semantic feature, question classification feature, and question structure composition feature) to obtain a target loss function, the question representation generation model is trained based on the target loss function to simultaneously optimize the effects of the model generating semantic features, question classification features, and question structure composition features based on the target statement. At the same time, by sharing some model layers (such as the first text encoding layer and the first pooling layer) in the question representation generation model during the process of obtaining the question classification features and question structure composition features, the training efficiency of the model can be improved.
[0079] In some feasible embodiments, each of the multiple sample questions may include sample question description information and sample question answer information. When the terminal device trains the first text encoding layer, the first pooling layer, the second text encoding layer, and the second pooling layer based on the multiple sample questions to obtain the third loss function, the terminal device may set the sample question description information and the sample question answer information in any sample question as the positive sample (which may be the first training sample) of the corresponding sample question. The terminal device pairs the sample question description information in any sample question with the sample question answer information of other sample questions except the sample question answer information of the above-mentioned any sample question in the multiple sample questions pairwise to form the negative sample (which may be the second training sample) of the above-mentioned any sample question. Optionally, the above-mentioned second training samples may also be shared in the same batch to save the computing power of the terminal device and improve the training efficiency of the model. At the same time, since different samples are independent during the training process, it can be conveniently extended to multi-card training. The terminal device may train to obtain the third loss function loss through the first text encoding layer, the first pooling layer, the second text encoding layer, and the second pooling layer based on the first training sample and the second training sample, which can be expressed as:
[0080]
[0081] Among them, after the terminal device obtains the corresponding vector representation of the sample question description information in the first training sample through the first text encoding layer and the first pooling layer, and obtains the corresponding vector representation of the sample question answer information in the first training sample through the second text encoding layer and the second pooling layer, the cosine similarity between the two obtained vectors (the vector representation corresponding to the sample question description information and the vector representation corresponding to the sample question answer information) is obtained to get c+. Similarly, c1-, c2-, and c3- etc. can be obtained based on the cosine similarity between the vector representations generated by the terminal device based on the second training sample through the first text encoding layer, the first pooling layer, the second text encoding layer, and the second pooling layer. The terminal device may train the first text encoding layer, the first pooling layer, the second text encoding layer, and the second pooling layer based on the weighted target loss function of the above-mentioned third loss function, so that each text encoding layer and each pooling layer can obtain a question structure composition feature closer to the relationship between the question description information and the question answer information for the input target question (including question description information and question answer information), and enhance the generalization ability of the question representation generation model.
[0082] In some feasible embodiments, please refer to Figure 8 , Figure 8 which is a schematic flowchart of question representation generation provided by an embodiment of the present application. As Figure 8As shown, the terminal device can obtain sample questions and input the sample questions into a pre-trained text encoding layer (which can include a first text encoding layer and a second text encoding layer. The above first text encoding layer and second text encoding layer can be text encoding layers based on the albert-tiny text encoding model, text encoding layers based on the albert text encoding model, and text encoding layers based on the bert text encoding model, etc.) and each classification layer (which can include a masked classification layer and a question classification layer) to perform multi-task pre-training based on the sample questions through the above pre-trained text encoding layer and each classification layer. Among them, each of the above sample questions can include sample question description information and sample question answer information. The above multi-task pre-training can include training the pre-trained text encoding layer and each classification layer based on the above objective loss function (which can be obtained by weighted summation of the above first loss function, the above second loss function, and the above third loss function) to obtain a question representation generation model, so that the above question representation generation model has the ability to obtain corresponding semantic features, question classification features, and question structure composition features for any input target question. The terminal device can input the obtained target question into the trained question representation generation model, and obtain a question representation based on the target question through the above question representation generation model. The question representation of the above target question contains information in multiple dimensions such as semantic information, classification information (which can also be called metadata information, and can include question difficulty, question type, and question knowledge points) and question structure composition information of the target question, so that the question representation can more fully represent the target question, and the question representation generation effect is good.
[0083] In some feasible implementation manners, after the terminal device obtains the question representation of the target question, it can perform question clustering through the generated question representation, that is, obtain the cosine similarity corresponding to each question representation after obtaining the question representations of different questions to divide each question into different categories. The terminal device can also apply the question representation to various downstream tasks at the question level, or continue to train the above question representation generation model on the data of the downstream tasks, so as to implement other question processing tasks (such as similar question recall tasks) through the question representation generation model.
[0084] In some feasible implementation manners, after the terminal device obtains the question representation of the target question, it can recommend similar questions based on the question representation. Specifically, please refer to Figure 9 , Figure 9 which is a schematic flowchart of generating similar questions provided by an embodiment of the present application. As shown in Figure 9As shown, the terminal device can obtain the corresponding topic representation (or called fusion feature) based on the target topic through the topic representation generation model, and obtain multiple candidate recommended topics and the corresponding candidate recommended features of each candidate recommended topic. The terminal device can obtain the cosine similarity between the above topic representation and each candidate recommended feature through vector retrieval based on the above topic representation, and measure the similarity degree between topics based on the cosine similarity to obtain the target recommended feature from multiple candidate recommended features, so as to obtain the candidate topic corresponding to the target recommended feature (which can be called the first candidate topic). Further, the terminal device can also obtain the second candidate topic with a text similarity to the target topic not less than a set threshold (for example, the text similarity is not less than 90%) from multiple candidate recommended topics through text matching (retrieving ES) based on the above target topic, and perform similar topic selection based on the above first candidate topic and the second candidate topic. The similar topics can be selected by retaining the repeated partial topics in the first candidate topic and the second candidate topic as the similar topics. In addition to the repeated part, some topics can also be selected from the non-repeated topics in the first candidate topic and the second candidate topic according to the cosine similarity from high to low as the similar topics. The finally determined similar topics can be pushed through learning and application after sorting and business rule filtering and displayed to the target push object. By adding the use of the topic representation generated by the topic representation generation model for similar topic recommendation, the accuracy and richness of the obtained similar topics can be improved. Therefore, the accuracy and coverage rate indicators of the similar topics requested by the target push object can be improved. (After testing, through the topic representation generated by the topic representation generation model proposed in this application for similar topic recommendation, the accuracy and coverage rate of similar topic recommendation have been significantly improved (the accuracy of junior high school mathematics has increased by 3.9%, and the coverage rate has increased by 2.3%), and the training model takes less than 10 ms), and the similar topic recommendation effect is good.
[0085] In the method provided in the embodiments of the present application, the terminal device may receive a target question, and the target question may be obtained based on a similar question recommendation instruction sent by a target push object through a learning application installed in the terminal device. When receiving the target question, the terminal device may obtain the question description information and question answer information included in the target question. Specifically, the above-mentioned question description information may include the stem information of the target question. When the target question is a multiple-choice question, the above-mentioned question description information may include the stem information and option information, and the above-mentioned question answer information may include the answer information and analysis information of the target question. The terminal device may randomly select some words (which may be called target words) in the above-mentioned target question and replace them with corresponding mask labels (or called mask tags), and input the target question with the above-mentioned mask tags into the question representation generation model. The terminal device obtains the word vectors corresponding to the above-mentioned mask tags through the third text encoding model in the masked language model layer of the above-mentioned question representation generation model, and inputs the word vectors generated by the third text encoding model into the masked classification layer to predict (softmax activation can be used) the target words corresponding to the above-mentioned mask tags. The predicted target words corresponding to the above-mentioned mask tags output by the masked classification layer contain the semantic information of the above-mentioned target question, so as to use the predicted target words corresponding to the above-mentioned mask tags as the semantic features of the target question. The terminal device may also obtain the word vectors corresponding to each word in the question description information of the target question through the first text encoding model in the above-mentioned question representation generation model, and perform a sum on the word vectors corresponding to each word in the sequence dimension through the first pooling layer to obtain a first vector representation. One or more question classification layers in the question representation generation model are respectively based on the above-mentioned first vector representation to obtain multiple masked language model layers corresponding to the target question to obtain the question classification features of the above-mentioned target question. The terminal device may also obtain a corresponding second vector representation based on the question description information of the above-mentioned target question through the second text encoding layer and the second pooling layer in the above-mentioned question representation generation model, and obtain the question structure composition features of the target question based on the first vector representation and the second vector representation obtained by the first text encoding layer and the first pooling layer based on the question description information of the above-mentioned target question. The terminal device may obtain the fusion features of the target question through the feature merging layer in the above-mentioned question representation generation model based on the above-mentioned semantic features, question classification features, and question structure composition features, so as to use the fusion features as the question representation of the target question. The above-mentioned question representation contains information in multiple dimensions such as semantic information, classification information (which may also be called metadata information, and may include question difficulty, question type, and question knowledge points), and question structure composition information of the target question, so that the question representation can more fully represent the target question, and the question representation generation effect is good.
[0086] Based on the description of the method embodiments for obtaining the topic representation above, embodiments of the present application also disclose an apparatus for obtaining the topic representation. This apparatus for obtaining the topic representation can be applied to Figures 1 to 9 the method for obtaining the topic representation in the embodiment shown, to be used for executing the steps in the method for obtaining the topic representation. Here, the apparatus for obtaining the topic representation can be the business server or the terminal device in the above Figures 1 to 9 shown embodiment, that is, this apparatus for obtaining the topic representation can be the execution subject of the method for obtaining the topic representation in the above Figures 1 to 9 shown embodiment. Please refer to Figure 10 , Figure 10 which is the structural schematic diagram of the apparatus for obtaining the topic representation provided by embodiments of the present application. In embodiments of the present application, the apparatus can operate the following modules:
[0087] An obtaining module 31, configured to obtain the topic description information and the topic solution information included in the target topic when receiving the target topic. The above topic description information includes the stem information and / or option information, and the above topic solution information includes the answer information and / or analysis information. Input the above target topic, the above topic description information, and the above topic solution information into the topic representation generation model;
[0088] A semantic feature generation module 32, configured to obtain the semantic feature of the target topic through the masked language model layer in the topic representation generation model when the above target topic is input into the topic representation generation model;
[0089] A topic classification feature generation module 33, configured to obtain a first vector representation corresponding to the above topic description information through the first text encoding layer and the first pooling layer in the topic representation generation model when the above topic description information is input into the topic representation generation model, and obtain the topic classification feature of the above target topic based on the above first vector representation through the topic classification layer in the topic representation generation model.
[0090] A topic structure composition feature generation module 34, configured to obtain a second vector representation corresponding to the above topic solution information through the second text encoding layer and the second pooling layer in the topic representation generation model when the above topic solution information is input into the topic representation generation model, and obtain the topic structure composition feature of the above target topic based on the above first vector representation and the above second vector representation;
[0091] A topic representation generation module 35, configured to generate a fusion feature of the above target topic as the topic representation of the above target topic based on the above semantic feature, the above topic classification feature, and the above topic structure composition feature through the feature merging layer in the topic representation generation model.
[0092] In some feasible embodiments, the above-mentioned masked language model layer includes a third text encoding layer and a masked classification layer, and the above-mentioned semantic feature generation module 32 is further configured to:
[0093] Replace one or more target words in the above-mentioned target title with one or more masked labels, and carry the above-mentioned one or more masked labels in the above-mentioned target title and input the above-mentioned target title into the title representation generation model;
[0094] The above-mentioned obtaining the semantic feature of the above-mentioned target title by the masked language model layer in the above-mentioned title representation generation model based on the input above-mentioned target title includes:
[0095] Obtain the word vectors corresponding to the above-mentioned one or more masked labels through the above-mentioned third text encoding model to obtain the word vectors of the above-mentioned one or more target words;
[0096] Use the predicted target words corresponding to the above-mentioned one or more masked labels obtained by the above-mentioned masked classification layer based on the above-mentioned word vectors as the semantic feature of the above-mentioned target title.
[0097] In some feasible embodiments, the above-mentioned title representation generation model includes at least one of the above-mentioned title classification layers, and the above-mentioned title classification feature generation module 33 is further configured to:
[0098] Obtain the word vectors corresponding to the words in the above-mentioned title description information through the first text encoding model in the above-mentioned title representation generation model, and perform summation on the sequence dimension of the word vectors corresponding to the above-mentioned words through the first pooling layer in the above-mentioned title representation generation model to obtain the first vector representation corresponding to the above-mentioned title description information;
[0099] Obtain any classification corresponding to the above-mentioned target title through any of the above-mentioned title classification layers in the above-mentioned title representation generation model based on the above-mentioned first vector representation, obtain the respective classifications obtained through each of the above-mentioned title classification layers, and obtain the title classification feature corresponding to the above-mentioned title description information based on the above-mentioned respective classifications.
[0100] In some feasible embodiments, the above-mentioned title structure composition feature generation module 34 is further configured to:
[0101] Obtain the word vectors corresponding to the words in the above-mentioned title answer information through the second text encoding model in the above-mentioned title representation generation model, and perform summation on the sequence dimension of the word vectors corresponding to the above-mentioned words through the second pooling layer in the above-mentioned title representation generation model to obtain the second vector representation.
[0102] In some feasible embodiments, before obtaining the question description information and question answer information included in the above-mentioned target question, the semantic feature generation module 32, the question classification feature generation module 33, and the question structure composition feature generation module 34 are further configured to:
[0103] Based on multiple sample questions and the above-mentioned masked language model layer, obtain a first loss function corresponding to the semantic features generated by the question representation generation model. Based on the multiple sample questions and the above-mentioned first text encoding layer, the first pooling layer, and the question classification layer, obtain a second loss function corresponding to the question classification features generated by the question representation generation model. And based on the multiple sample questions and the above-mentioned first text encoding layer, the first pooling layer, the second text encoding layer, and the second pooling layer, obtain a third loss function corresponding to the question structure composition features generated by the question representation generation model;
[0104] Weighted sum the above-mentioned first loss function, the second loss function, and the third loss function to obtain a target loss function. Based on the target loss function and the multiple sample questions, train the question representation generation model, so that the masked language model layer of the question representation generation model obtains the ability to obtain the semantic features of any input target question, enabling the first text encoding layer, the first pooling layer, and the question classification layer to obtain the question classification features of the question description information of any input target question, and enabling the first text encoding layer, the first pooling layer, the second text encoding layer, and the second pooling layer to obtain the question structure composition features of the question description information and the question answer information of any input target question.
[0105] In some feasible embodiments, each of the multiple sample questions includes at least sample question description information and sample question answer information. The question structure composition feature generation module 34 is further configured to:
[0106] Set the sample question description information and the sample question answer information in any sample question as the first training sample of the any sample question. Pair the sample question description information in the any sample question with the remaining sample information in the multiple sample questions pairwise to form the second training sample of the any sample question. The remaining sample information is other sample question answer information included in the multiple sample questions except the sample question answer information of the any sample question;
[0107] The first text encoding layer, the first pooling layer, the second text encoding layer and the second pooling layer are trained based on the first training samples and the second training samples of the above-mentioned respective sample questions to obtain the third loss function.
[0108] In some feasible implementations, after the feature merging layer in the title representation generation model generates the fused feature of the target title based on the semantic feature, the title classification feature and the title structure composition feature as the title representation of the target title, the acquisition module 31 is further used to: obtain the cosine similarity between the fused feature and the candidate recommendation feature corresponding to each candidate recommendation title in multiple candidate recommendation titles, obtain the target recommendation feature from the multiple candidate recommendation features based on the cosine similarity between the fused feature and each candidate recommendation feature, and use the candidate recommendation title associated with the target recommendation feature as the first candidate title;
[0109] A second candidate title whose text similarity with the target title is not less than a set threshold is obtained from the multiple candidate recommended titles through text similarity matching, and a similar title of the target title obtained based on the first candidate title and the second candidate title is sent to the target push object.
[0110] According to the above Figure 2 The corresponding embodiment, Figure 2 The implementation method described in steps S101 to S105 in the method for obtaining the title representation shown in the figure can be implemented by Figure 10 Each module of the device shown in the figure is executed. For example, the above Figure 2 The implementation method described in step S101 of the method for obtaining the title representation shown in FIG. Figure 10 The implementation method described in step S102 can be performed by the semantic feature generation module 32, the implementation method described in step S103 can be performed by the topic classification feature generation module 33, the implementation method described in step S104 can be performed by the topic structure composition feature generation module 34, and the implementation method described in step S105 can be performed by the topic representation generation module 35. The implementation methods performed by the acquisition module 31, the semantic feature generation module 32, the topic classification feature generation module 33, the topic structure composition feature generation module 34, and the topic representation generation module 35 can be referred to in the above Figure 2 The implementation methods provided in each step of the corresponding embodiment will not be repeated here.
[0111] In an embodiment of the present application, the acquisition device for the topic representation can receive a target topic, and the target topic can be obtained based on a similar topic recommendation instruction sent by a learning application loaded in the acquisition device for the topic representation through a target push object. When the target topic is received, the acquisition device for the topic representation can obtain the topic description information and the topic solution information included in the target topic. Specifically, the above-mentioned topic description information can include the stem information of the target topic. When the target topic is a multiple-choice question, the above-mentioned topic description information can include the stem information and the option information, and the above-mentioned topic solution information can include the answer information and the analysis information of the target topic. The acquisition device for the topic representation can randomly select some words (which can be called target words) in the above-mentioned target topic and replace them with corresponding mask labels (or called mask tags), and input the target topic with the above-mentioned mask tags into the topic representation generation model. The acquisition device for the topic representation obtains the word vectors corresponding to the above-mentioned mask tags through the third text encoding model in the masked language model layer of the above-mentioned topic representation generation model, and inputs the word vectors generated by the third text encoding model into the mask classification layer to predict (softmax activation can be used) the target words corresponding to the above-mentioned mask tags through the mask classification layer. The predicted target words corresponding to the above-mentioned mask tags output by the mask classification layer contain the semantic information of the above-mentioned target topic, so as to use the predicted target words corresponding to the above-mentioned mask tags as the semantic features of the target topic. The acquisition device for the topic representation can also obtain the word vectors corresponding to the words in the topic description information of the target topic through the first text encoding model in the above-mentioned topic representation generation model, and perform a sum of the word vectors corresponding to the words in the sequence dimension through the first pooling layer to obtain a first vector representation, and obtain multiple mask language model layers corresponding to the target topic respectively based on the above-mentioned first vector representation through one or more topic classification layers in the topic representation generation model to obtain the topic classification features of the target topic. The acquisition device for the topic representation can also obtain a corresponding second vector representation based on the topic description information of the above-mentioned target topic through the second text encoding layer and the second pooling layer in the above-mentioned topic representation generation model, and obtain the topic structure composition features of the target topic based on the first vector representation and the second vector representation obtained by the above-mentioned first text encoding layer and the first pooling layer based on the topic description information of the above-mentioned target topic. The acquisition device for the topic representation can obtain the fusion features of the target topic through the feature merging layer in the above-mentioned topic representation generation model based on the above-mentioned semantic features, topic classification features and topic structure composition features, so as to use the fusion features as the topic representation of the target topic. The above-mentioned topic representation contains information in multiple dimensions such as semantic information, classification information (which can also be called metadata information, and can include topic difficulty, topic type and topic knowledge points) and topic structure composition information of the target topic, so that the topic representation can more fully represent the target topic, and the topic representation generation effect is good.
[0112] In the embodiments of the present application, each module in the device shown in the above figure may be separately or entirely combined into one or several other modules to form, or a certain (some) module may also be further split into multiple smaller modules in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In practical applications, the function of one module may also be realized by multiple modules, or the functions of multiple modules may be realized by one module. In other feasible implementation manners of the present application, the above device may also include other modules. In practical applications, these functions may also be assisted by other modules and may be realized by the cooperation of multiple modules, which are not limited herein.
[0113] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 11 shown, the computer device 1000 may be the terminal device in the corresponding embodiment of the above Figures 2 - 9 . The computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 may further include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As Figure 11 shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0114] Among them, the network interface 1004 in the computer device 1000 may also be network-connected to the terminal 200a in the corresponding embodiment of the above Figure 1 , and optionally, the user interface 1003 may further include a display screen (Display) and a keyboard (Keyboard). In Figure 11In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface for developers; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement the acquisition method represented by the title in the corresponding embodiment described above Figure 2 of the acquisition method represented by the title in the corresponding embodiment.
[0115] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the acquisition method represented by the title in the corresponding embodiment described above Figure 2 and will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.
[0116] In addition, it should be noted here that the embodiments of the present application also provide a computer-readable storage medium, and the computer program executed by the acquisition device represented by the title mentioned above is stored in the above computer-readable storage medium. The above computer program includes program instructions. When the above processor executes the above program instructions, it can execute the description of the acquisition method represented by the title in the corresponding embodiment described above. Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application. Figure 2
[0117] In addition, it should be noted that: the embodiments of the present application also provide a computer program product, which may include a computer program, and the computer program can be stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor can execute the computer program, so that the computer device executes the description of the acquisition method represented by the title in the corresponding embodiment described above Figures 2 to 9 and will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer program product involved in the present application, please refer to the description of the method embodiments of the present application.
[0118] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above methods. Among them, the above storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0119] The above disclosure is only for the preferred embodiments of the present application. Of course, it cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A method for obtaining problem representation, characterized in that, The method includes: Obtain the question description information and question answer information included in the target question. The question description information includes stem information and / or option information, and the question answer information includes answer information and / or analysis information. Input the target question, the question description information, and the question answer information into a question representation generation model in a preset format; Through the masked language model layer in the question representation generation model, based on the input target question, obtain the semantic features of the target question; Obtain a first vector representation corresponding to the input question description information through the first text encoding layer and the first pooling layer in the question representation generation model, and through the question classification layer in the question representation generation model, obtain the question classification features of the target question based on the first vector representation; Obtain a second vector representation corresponding to the input question answer information through the second text encoding layer and the second pooling layer in the question representation generation model, and obtain the question structure composition features of the target question based on the first vector representation and the second vector representation; Through the feature merging layer in the question representation generation model, generate the fusion features of the target question based on the semantic features, the question classification features, and the question structure composition features as the question representation of the target question, and the question representation is used for question clustering and / or similar question recommendation of the target application.
2. The method according to claim 1, wherein The masked language model layer includes a third text encoding layer and a masked classification layer; before obtaining the semantic features of the target question through the masked language model layer in the question representation generation model based on the input target question, the method further includes: Replace one or more target words in the target question with one or more masked labels, and input the target question carrying the one or more masked labels into the question representation generation model; The obtaining the semantic features of the target question through the masked language model layer in the question representation generation model based on the input target question includes: Obtain the word vectors corresponding to the one or more masked labels through the third text encoding layer to obtain the word vectors of the one or more target words; Use the predicted target words corresponding to the one or more masked labels obtained by the masked classification layer based on the word vectors as the semantic features of the target question.
3. The method according to claim 2, wherein The question representation generation model includes at least one of the question classification layers. The obtaining a first vector representation corresponding to the input question description information through the first text encoding layer and the first pooling layer in the question representation generation model, and obtaining the question classification features of the target question through the question classification layer in the question representation generation model based on the first vector representation includes: Obtain the word vectors corresponding to the words in the question description information through the first text encoding model in the question representation generation model, and perform a sum of the word vectors corresponding to the words in the sequence dimension through the first pooling layer in the question representation generation model to obtain the first vector representation corresponding to the question description information; Any one of the topic classification layers in the topic representation generation model obtains any classification corresponding to the target topic based on the first vector representation, obtains each classification obtained through each of the topic classification layers, and obtains a topic classification feature corresponding to the topic description information based on each of the classifications.
4. The method according to claim 3, characterized in that, The obtaining of the second vector representation corresponding to the input topic answer information through the second text encoding layer and the second pooling layer in the topic representation generation model includes: The second text encoding model in the topic representation generation model obtains word vectors corresponding to each word in the topic answer information, and the second pooling layer in the topic representation generation model sums the word vectors corresponding to each word in the sequence dimension to obtain the second vector representation.
5. The method according to any one of claims 1 to 4, characterized in that, Before obtaining the topic description information and the topic answer information included in the target topic, the method further includes: Based on multiple sample topics and the masked language model layer, obtaining a first loss function for the topic representation generation model to generate the semantic feature, based on the multiple sample topics and the first text encoding layer, the first pooling layer, and the topic classification layer, obtaining a second loss function for the topic representation generation model to generate the topic classification feature, and based on the multiple sample topics and the first text encoding layer, the first pooling layer, the second text encoding layer, and the second pooling layer, obtaining a third loss function for the topic representation generation model to generate the topic structure composition feature; Weightedly summing the first loss function, the second loss function, and the third loss function to obtain a target loss function, and training the topic representation generation model based on the target loss function and the multiple sample topics.
6. The method according to claim 5, wherein Each of the multiple sample topics includes at least sample topic description information and sample topic answer information. The obtaining of the third loss function based on the multiple sample topics and the first text encoding layer, the first pooling layer, the second text encoding layer, and the second pooling layer includes: Setting the sample topic description information and the sample topic answer information in any one of the sample topics as the first training sample of the any one of the sample topics, and pairwise pairing the sample topic description information in the any one of the sample topics with the remaining sample information in the multiple sample topics to form the second training sample of the any one of the sample topics, where the remaining sample information is other sample topic answer information included in the multiple sample topics except the sample topic answer information of the any one of the sample topics; Training the first text encoding layer, the first pooling layer, the second text encoding layer, and the second pooling layer based on the first training sample and the second training sample of each of the sample topics to obtain the third loss function.
7. The method according to any one of claims 1, 2, 3, 4 and 6, characterized in that, After generating a fusion feature of the target topic as the topic representation of the target topic through the feature merging layer in the topic representation generation model based on the semantic feature, the topic classification feature, and the topic structure composition feature, the method further includes: Obtain the cosine similarity between the fusion feature and the candidate recommendation features corresponding to each candidate recommended question among multiple candidate recommended questions, obtain the target recommendation feature from the multiple candidate recommendation features based on the cosine similarity between the fusion feature and each candidate recommendation feature, and use the candidate recommended question associated with the target recommendation feature as the first candidate question; Obtain a second candidate question with a text similarity to the target question not less than a set threshold from the multiple candidate recommended questions through text similarity matching, and send the similar questions of the target question obtained based on the first candidate question and the second candidate question to the target push object.
8. An apparatus for obtaining a problem representation, characterized in that, Comprising: An acquisition module, configured to, when receiving a target question, obtain the question description information and question answer information included in the target question, where the question description information includes stem information and / or option information, and the question answer information includes answer information and / or analysis information, and input the target question, the question description information, and the question answer information into a question representation generation model in a preset format; A semantic feature generation module, configured to, when the target question is input into the question representation generation model, obtain the semantic feature of the target question through the masked language model layer in the question representation generation model; A question classification feature generation module, configured to, when the question description information is input into the question representation generation model, obtain a first vector representation corresponding to the question description information through the first text encoding layer and the first pooling layer in the question representation generation model, and obtain the question classification feature of the target question based on the first vector representation through the question classification layer in the question representation generation model; A question structure composition feature generation module, configured to, when the question answer information is input into the question representation generation model, obtain a second vector representation corresponding to the question answer information through the second text encoding layer and the second pooling layer in the question representation generation model, and obtain the question structure composition feature of the target question based on the first vector representation and the second vector representation; A question representation generation module, configured to generate a fusion feature of the target question as the question representation of the target question based on the semantic feature, the question classification feature, and the question structure composition feature through the feature merging layer in the question representation generation model.
9. A computer device, characterized in that, Comprising: A processor, a memory, and a network interface; The processor is connected to the memory and the network interface, where the network interface is used to provide a data communication function, the memory is used to store program codes, and the processor is used to call the program codes to execute the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is suitable for being loaded and executed by the processor to execute the method according to any one of claims 1-7.
11. A computer program product, characterized in that, The computer program product includes a computer program which is stored in a computer-readable storage medium and is adapted to be read and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Topic detection method and device, electronic equipment and storage medium
CN114282531A