Question generation method and device, electronic equipment and storage medium
By obtaining the correlation feature information of the target object, and using the topic generation and quality analysis model to optimize the conversation topic, the problem of low correlation between topic content and users in the existing dialogue system is solved, and the user experience and retention rate are improved.
Patent Information
- Application Number
- CN202410002627.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-08
AI Technical Summary
In the existing dialogue system, the topic content has low correlation with users, resulting in a decrease in user viscosity and poor user experience.
By obtaining the associated feature information of the target object, generating target topic generation instructions, and using the target topic generation model and quality analysis model, predicting and optimizing the conversation topic, and generating question information that is closer to user preferences.
Improves the prediction accuracy and user satisfaction of conversation topics, and increases user stickiness and retention.
Smart Images

Figure CN120277174A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a question generation method, apparatus, electronic device, and storage medium. Background Art
[0002] Artificial intelligence is a technology that simulates human intelligence and aims to create intelligent machines that can learn and process information independently. The artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. Specifically, natural language processing is mainly applied to aspects such as machine translation, automatic summarization, text classification, dialogue systems, text semantic comparison, and speech recognition. Among them, a dialogue system is a computer system that simulates humans and aims to form a coherent and smooth dialogue with humans. For example, for a dialogue system with a certain role setting, the duration and number of rounds of daily conversations are very limited, and it can only passively answer questions raised by users, and it is easy to fall into an awkward chat state.
[0003] Currently, it is possible to obtain dialogue topics from multiple platforms and aggregate relevant information including the specified dialogue topic to obtain a dialogue topic cluster for the current dialogue topic, so as to calculate the popularity of the dialogue topic; then, based on the dialogue topic cluster and the popularity of the dialogue topic, a title or descriptive text can be generated, and the dialogue system can implement an active dialogue with the user based on the above title or descriptive text. However, the topic content in the above method has a low relevance to the user, resulting in a decrease in user viscosity and a poor user experience. Summary of the Invention
[0004] In view of the above existing technical problems, the present disclosure proposes a question generation method, apparatus, electronic device, and storage medium.
[0005] According to one aspect of the embodiments of the present disclosure, a question generation method is provided, including:
[0006] Obtaining target associated feature information corresponding to a target object, where the target associated feature information includes target historical dialogue information corresponding to the target object and target object attribute information corresponding to the target object;
[0007] Generating a target theme generation instruction based on the target associated feature information;
[0008] Inputting the target theme generation instruction into a target theme generation model for dialogue theme prediction processing to obtain a plurality of target theme word segments corresponding to the target object; the target theme generation model is obtained by training a preset machine learning model based on target loss information;
[0009] Input the multiple target topic word segmentations into a target quality analysis model for quality analysis processing to obtain target topic quality information; the target quality analysis model is obtained by training a preset reinforcement learning model based on the target loss information, and the target loss information represents the degree of difference between a first topic quality information and a second topic quality information in the sample topic quality information. The first topic quality information is the quality information corresponding to a preset sample topic word segmentation in the sample topic quality information, and the second topic quality information is the maximum quality information in the sample topic quality information; the preset sample topic word segmentation is determined based on the topic word segmentations selected by the sample object.
[0010] Based on the multiple target topic word segmentations, the target topic quality information, and a preset dialogue generation model, generate target question information for the target object.
[0011] According to another aspect of the embodiments of the present disclosure, there is provided a question generation device, including:
[0012] A target feature acquisition module, configured to acquire target associated feature information corresponding to a target object, where the target associated feature information includes target historical dialogue information corresponding to the target object and target object attribute information corresponding to the target object;
[0013] A target instruction generation module, configured to generate a target topic generation instruction based on the target associated feature information;
[0014] A target topic prediction module, configured to input the target topic generation instruction into a target topic generation model for dialogue topic prediction processing to obtain multiple target topic word segmentations corresponding to the target object; the target topic generation model is obtained by training a preset machine learning model based on target loss information;
[0015] A first quality analysis module, configured to input the multiple target topic word segmentations into a target quality analysis model for quality analysis processing to obtain target topic quality information; the target quality analysis model is obtained by training a preset reinforcement learning model based on the target loss information, and the target loss information represents the degree of difference between a first topic quality information and a second topic quality information in the sample topic quality information. The first topic quality information is the quality information corresponding to a preset sample topic word segmentation in the sample topic quality information, and the second topic quality information is the maximum quality information in the sample topic quality information; the preset sample topic word segmentation is determined based on the topic word segmentations selected by the sample object.
[0016] A target question generation module, configured to generate target question information for the target object based on the multiple target topic word segmentations, the target topic quality information, and a preset dialogue generation model.
[0017] Optionally, the device further includes:
[0018] A sample feature acquisition module, configured to acquire sample associated feature information, where the sample associated feature information includes sample historical dialogue information corresponding to a sample object and sample object attribute information corresponding to the sample object;
[0019] A sample instruction generation module, configured to generate a sample topic generation instruction based on the sample associated feature information;
[0020] A sample topic prediction module, configured to input the sample topic generation instruction into the preset machine learning model for dialogue topic prediction processing to obtain multiple sample topic word segments corresponding to the sample object;
[0021] A selected topic acquisition module, configured to acquire a preset sample topic word segment corresponding to the sample object;
[0022] A second quality analysis module, configured to input the multiple sample topic word segments into the preset reinforcement learning model for quality analysis processing to obtain the sample topic quality information;
[0023] A loss determination module, configured to determine the target loss information based on the sample topic quality information and the preset sample topic word segment;
[0024] A model training module, configured to train the preset machine learning model and the preset reinforcement learning model based on the target loss information to obtain the target topic generation model and the target quality analysis model.
[0025] Optionally, the preset machine learning model includes a sample encoding module, a sample fine-tuning module, a sample feature processing module, and a sample prediction module; the sample topic prediction module includes:
[0026] A first encoding processing module, configured to input the sample topic generation instruction into the sample encoding module for encoding processing to obtain sample instruction encoding information;
[0027] A first feature processing module, configured to input the sample instruction encoding information into the sample fine-tuning module for feature processing to obtain first sample feature information;
[0028] A second feature processing module, configured to input the sample instruction encoding information into the sample feature processing module for feature processing to obtain second sample feature information;
[0029] A first prediction processing module, configured to input the first sample feature information and the second sample feature information into the sample prediction module for topic prediction processing to obtain the multiple sample topic word segments.
[0030] Optionally, the model training module includes:
[0031] A first training module, configured to train the sample fine-tuning module and the sample prediction module in the preset machine learning model based on the target loss information to obtain the target topic generation model;
[0032] A second training module, configured to train the preset reinforcement learning model based on the target loss information to obtain the target quality analysis model.
[0033] Optionally, the selected topic acquisition module includes:
[0034] A historical selection acquisition module, configured to acquire the historical selected topic word segmentation corresponding to the sample object;
[0035] A similarity analysis module, configured to perform similarity analysis on the historical selected topic word segmentation and each sample topic word segmentation in the multiple sample topic word segmentations to obtain the corresponding similarity index data for each sample topic word segmentation;
[0036] A selected topic determination module, configured to use the sample topic word segmentation with the largest corresponding similarity index data in the multiple sample topic word segmentations as the preset sample topic word segmentation.
[0037] Optionally, the device further includes:
[0038] A sample information acquisition module, configured to acquire at least one first preset high-frequency word segmentation and sample model attribute information corresponding to a first historical time period; the first historical time period is a preset time period before a first current moment;
[0039] Correspondingly, the sample instruction generation module includes:
[0040] A first instruction generation module, configured to generate the sample topic generation instruction based on the sample association feature information, the at least one first preset high-frequency word segmentation, and the sample model attribute information.
[0041] Optionally, the device further includes:
[0042] A target information acquisition module, configured to acquire at least one second preset high-frequency word segmentation and target model attribute information corresponding to the preset dialogue generation model corresponding to a second historical time period; the second historical time period is a preset time period before a second current moment;
[0043] Correspondingly, the target instruction generation module includes:
[0044] A second instruction generation module, configured to generate the target topic generation instruction based on the target association feature information, the at least one second preset high-frequency word segmentation, and the target model attribute information.
[0045] Optionally, the target information acquisition module includes:
[0046] A short text acquisition module, configured to acquire a plurality of preset short text information corresponding to the second historical time period from at least one preset information platform;
[0047] A word segmentation processing module, configured to perform word segmentation processing on the plurality of preset short text information to obtain a plurality of text word segments;
[0048] A text word segment determination module, configured to determine a variety of text word segments based on the plurality of text word segments;
[0049] A frequency analysis module, configured to perform frequency analysis on the plurality of text word segments to obtain current word frequency index data corresponding to each of the variety of text word segments;
[0050] A historical word frequency acquisition module, configured to acquire historical word frequency index data corresponding to each of the variety of text word segments;
[0051] A high-frequency word segmentation determination module, configured to determine the at least one second preset high-frequency word segmentation from the variety of text word segments based on the current word frequency index data and the historical word frequency index data.
[0052] Optionally, the high-frequency word segmentation determination module includes:
[0053] A word frequency ratio analysis module, configured to perform word frequency ratio analysis on each text word segment based on the current word frequency index data and the historical word frequency index data to obtain word frequency ratio index data corresponding to each text word segment; the word frequency ratio index data corresponding to each text word segment represents the ratio between the current word frequency index data corresponding to each text word segment and the historical word frequency index data corresponding to each text word segment;
[0054] An initial word segment acquisition module, configured to select text word segments from the variety of text word segments whose corresponding word frequency ratio index data is greater than a preset ratio index data to obtain at least one initial text word segment;
[0055] A word segment screening module, configured to delete initial text word segments in the at least one initial text word segment whose corresponding current word frequency index data and corresponding historical word frequency index data are both less than a preset word frequency index data to obtain the at least one second preset high-frequency word segmentation.
[0056] Optionally, the target topic generation model includes a target encoding module, a target fine-tuning module, a target feature processing module, and a target prediction module; the target topic prediction module includes:
[0057] A second encoding processing module, configured to input the target topic generation instruction into the target encoding module for encoding processing to obtain target instruction encoding information;
[0058] A third feature processing module, configured to input the target instruction encoding information into the target fine-tuning module for feature processing to obtain first target feature information;
[0059] A fourth feature processing module, configured to input the target instruction encoding information into the target feature processing module for feature processing to obtain second target feature information;
[0060] A second prediction processing module, configured to input the first target feature information and the second target feature information into the target prediction module for topic prediction processing to obtain the multiple target topic segmentations.
[0061] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the above-mentioned question generation method.
[0062] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the above-mentioned question generation method.
[0063] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product including instructions, when it runs on a computer, enabling the computer to execute the above-mentioned question generation method.
[0064] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0065] Obtain the target associated feature information corresponding to the target object. The target associated feature information includes the target historical conversation information corresponding to the target object and the target object attribute information corresponding to the target object, which can realize the acquisition of associated features of the target object in multiple dimensions. Then, combined with the target associated feature information, generate a target topic generation instruction, which can facilitate the accurate prediction of the target topic word segmentation associated with the target object. Secondly, input the target topic generation instruction into the target topic generation model for conversation topic prediction processing to obtain multiple target topic word segmentations corresponding to the target object. The target topic generation model is obtained by training a preset machine learning model based on the target loss information, which can realize the efficient prediction of the conversation topic of the active conversation for the target object. Then, input the multiple target topic word segmentations into the target quality analysis model for quality analysis processing to obtain the target topic quality information. The target quality analysis model is obtained by training a preset reinforcement learning model based on the target loss information. The target loss information represents the difference degree between the first topic quality information and the second topic quality information in the sample topic quality information. The first topic quality information is the quality information corresponding to the preset sample topic word segmentation in the sample topic quality information, and the second topic quality information is the largest quality information in the sample topic quality information. The preset sample topic word segmentation is determined based on the topic word segmentation selected by the sample object. The above target loss information can be used to train the model to improve the relevance between the topic result predicted by the model and the topic result preferred by the target object, thereby improving the prediction accuracy of the target quality analysis model and the target topic generation model. Then, combined with the multiple target topic word segmentations, the target topic quality information and the preset conversation generation model, generate the target question information for the target object, which can make the target question information closer to the preference of the target object, thereby increasing user viscosity, improving user satisfaction and user retention rate, and bringing a better chat conversation experience.
[0066] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. Brief Description of the Drawings
[0067] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0068] Figure 1 is a schematic diagram of an application system shown according to an exemplary embodiment;
[0069] Figure 2 is a flowchart of a question generation method shown according to an exemplary embodiment;
[0070] Figure 3It is a schematic flowchart of a question generation method shown according to an exemplary embodiment;
[0071] Figure 4 It is a block diagram of a question generation device shown according to an exemplary embodiment;
[0072] Figure 5 It is a block diagram of an electronic device for generating target question information shown according to an exemplary embodiment;
[0073] Figure 6 It is a block diagram of another electronic device for generating target question information shown according to an exemplary embodiment. Detailed implementation manners
[0074] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0075] The special word "exemplary" here means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" here does not have to be construed as superior to or better than other embodiments.
[0076] In addition, for a better description of the present application, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present application can be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present application.
[0077] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0078] In recent years, with the research and progress of artificial intelligence technology, artificial intelligence technology has been widely applied in multiple fields. The solution provided in the embodiments of the present application relates to technologies such as natural language processing, and will be specifically described through the following embodiments:
[0079] Please refer to Figure 1 , Figure 1 It is a schematic diagram of an application system shown according to an exemplary embodiment. The application system can be used for the question generation method of the present application. As Figure 1As shown, the application system can at least include server 01 and terminal 02.
[0080] In the embodiment of the present application, server 01 can be used to generate target question information. Specifically, the above-mentioned server 01 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0081] In the embodiment of the present application, terminal 02 can be used to generate target object attribute information corresponding to the target object. The above-mentioned terminal 02 can include entity devices such as smart phones, desktop computers, tablet computers, laptop computers, smart speakers, vehicle-mounted terminals, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices, or can also include software running on the entity devices, such as application programs. In the embodiment of the present application, the operating system running on the above-mentioned terminal 02 can include, but is not limited to, Android system, IOS system, linux, windows, etc.
[0082] In addition, it should be noted that Figure 1 The application environment shown is only one provided by the present disclosure. In actual applications, other application environments may also be included. For example, for the generation process of the target question information, it can also be implemented on terminal 02.
[0083] In the embodiment of this specification, the above-mentioned terminal 02 and server 01 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any limitations in this regard.
[0084] It should be noted that the step sequence shown in the following figure is a possible one, and in fact, it is not necessary to strictly follow this order. Some steps can be executed in parallel without relying on each other.
[0085] Specifically, Figure 2 is a flowchart of a question generation method shown according to an exemplary embodiment. As Figure 2 shown, the question generation method can be used in electronic devices such as terminals or servers, and specifically can include the following steps:
[0086] S201: Obtain target associated feature information corresponding to the target object.
[0087] In a specific embodiment, the target object can refer to an object that needs to provide an active conversation. Specifically, the target object can include a user account.
[0088] In a specific embodiment, the target associated feature information may characterize the associated features of the target object in multiple dimensions. The target associated feature information may include the target historical dialogue information corresponding to the target object and the target object attribute information corresponding to the target object. Among them, the target historical dialogue information may refer to the dialogue information between the target object and the target dialogue system corresponding to the preset dialogue generation model within a preset historical time range. The target historical dialogue information may include multiple question information and the answer information corresponding to each question information. The target object attribute information may be used to characterize the attributes of the target object from multiple perspectives. The target object attribute information may include multiple object attribute information of the target object. Exemplarily, the target object attribute information may include one or more of the attribute information such as the age information corresponding to the target object, the location information where the target object is located, and the preference resource information corresponding to the target object.
[0089] In a specific embodiment, the target historical dialogue information and the target object attribute information may be obtained from the preset memory of the server.
[0090] S203: Generate a target topic generation instruction based on the target associated feature information.
[0091] In a specific embodiment, the target topic generation instruction may be used to instruct the target topic generation model to perform dialogue topic prediction based on the target associated feature information. Specifically, the specific manifestation form of the target topic generation instruction may include text, etc. Exemplarily, the target topic generation instruction may be "Please help me generate a topic related to [XXX (person's name)] [music], which can meet the attribute characteristics of [young users], and is required to be brief, attractive, and not exceed 20 characters".
[0092] In a specific embodiment, the target topic generation instruction may be generated based on the target associated feature information and the preset instruction prompt information. Specifically, multiple feature information in the target associated feature information may be respectively added to multiple areas to be supplemented in the preset instruction prompt information to obtain the target topic generation instruction. Exemplarily, the preset instruction prompt information may include "Please help me generate a topic related to [XX] [XX] [XX], which can meet the attribute characteristics of [XX] [XX], and is required to be concise, attractive, and not exceed N characters", "Please generate an attractive topic in combination with the [XX] attribute characteristics for the recent [XX] social hot topic, not exceeding N characters, and be concise and powerful", or "Please generate an attractive topic in combination with the interest points of the recent [XX] movie and TV works and [XX] stars, not exceeding N characters, and be concise and powerful", where the areas to be supplemented in the above preset instruction prompt information may refer to the areas corresponding to "[XX]".
[0093] In the above embodiments, by combining the association features of the target object in multiple dimensions in the target association feature information, a target theme generation instruction is generated, which can facilitate the accurate prediction of the target theme word segmentation associated with the target object.
[0094] In a specific embodiment, the above method may further include:
[0095] Obtain at least one second preset high-frequency word segmentation corresponding to the second historical time period and the target model attribute information corresponding to the preset dialogue generation model;
[0096] Correspondingly, generating the target theme generation instruction based on the target association feature information may include:
[0097] Generate a target theme generation instruction based on the target association feature information, at least one second preset high-frequency word segmentation, and the target model attribute information.
[0098] In a specific embodiment, the second historical time period is a preset time period before the second current moment. The second current moment may refer to the current moment during the model application process. Exemplarily, the second historical time period may refer to the current day during the model application process.
[0099] In a specific embodiment, at least one second preset high-frequency word segmentation may refer to a word segmentation with a frequency higher than the preset frequency in the second historical time period. At least one second preset high-frequency word segmentation may include popular words, etc.
[0100] In a specific embodiment, the above at least one second preset high-frequency word segmentation may be obtained in the following manner:
[0101] Obtain multiple preset short text information corresponding to the second historical time period from at least one preset information platform;
[0102] Perform word segmentation processing on the multiple preset short text information to obtain multiple text word segmentations;
[0103] Based on the multiple text word segmentations, determine multiple text word segmentations;
[0104] Perform frequency analysis on the multiple text word segmentations to obtain the current word frequency index data corresponding to each of the multiple text word segmentations;
[0105] Obtain the historical word frequency index data corresponding to each of the multiple text word segmentations;
[0106] Based on the current word frequency index data and the historical word frequency index data, determine at least one second preset high-frequency word segmentation from the multiple text word segmentations.
[0107] In a specific embodiment, the preset information platform can be used to provide recent hot information. The preset information platform can include a preset information platform.
[0108] In a specific embodiment, the multiple preset short text information corresponding to the second historical time period may refer to the short text information in the resource description information publicly disclosed by the at least one preset information platform within the second historical time period. Exemplarily, the preset short text information may be "XX issues a blue cold wave warning".
[0109] In a specific embodiment, each preset short text information among the multiple preset short text information can be subjected to word segmentation processing to obtain multiple text segments. Exemplarily, when the preset short text information is "XX issues a blue cold wave warning", performing word segmentation processing on the above preset short text information can obtain multiple text segments including "cold wave", "XX", "issues", and "warning", etc.
[0110] In a specific embodiment, multiple types of text segments can be obtained by performing deduplication processing on the multiple text segments.
[0111] In a specific embodiment, the current word frequency index data corresponding to any text segment can represent the frequency of occurrence of the above any text segment within the second historical time period.
[0112] In a specific embodiment, the number of times each type of text segment appears in the multiple text segments can be used as the current word frequency index data corresponding to each type of text segment.
[0113] In a specific embodiment, the historical word frequency index data corresponding to any text segment can represent the frequency of occurrence of the above any text segment within the third historical time period. Among them, the third historical time period can be a preset time period before the second historical time period. Exemplarily, when the second historical time period is the current day, the third historical time period can be the day before yesterday.
[0114] In a specific embodiment, each time the current word frequency index data corresponding to multiple types of text segments is obtained through frequency analysis, based on the correspondence between the second historical time period and the current word frequency index data corresponding to each of the multiple types of text segments, the above word frequency index data can be stored in a preset memory. Correspondingly, based on the second historical time period, the third historical time period can be determined. Then, based on the text segment to be obtained and the third historical time period, the corresponding word frequency index data can be read from the above preset memory and used as the historical word frequency index data.
[0115] In a specific embodiment, a third historical time period may be determined based on a second historical time period; multiple historical short text information corresponding to the third historical time period may be obtained from at least one preset information platform; the multiple historical short text information may be subjected to word segmentation processing to obtain multiple historical word segments; based on the multiple historical word segments, frequency word segmentation may be performed on the multiple text word segments to obtain historical word frequency index data corresponding to each of the multiple text word segments.
[0116] In a specific embodiment, determining at least one second preset high-frequency word segment from multiple text word segments based on the current word frequency index data and the historical word frequency index data may include:
[0117] Performing word frequency ratio analysis on each text word segment based on the current word frequency index data and the historical word frequency index data to obtain word frequency ratio index data corresponding to each text word segment;
[0118] Selecting text word segments with corresponding word frequency ratio index data greater than the preset ratio index data from the multiple text word segments to obtain at least one initial text word segment;
[0119] Deleting initial text word segments in which both the corresponding current word frequency index data and the corresponding historical word frequency index data are less than the preset word frequency index data from the at least one initial text word segment to obtain at least one second preset high-frequency word segment.
[0120] In a specific embodiment, the word frequency ratio index data corresponding to each text word segment may represent the ratio between the current word frequency index data corresponding to each text word segment and the historical word frequency index data corresponding to each text word segment.
[0121] In a specific embodiment, the current word frequency index data corresponding to each text word segment may be divided by the historical word frequency index data corresponding to each text word segment to obtain the word frequency ratio index data corresponding to each text word segment.
[0122] In a specific embodiment, the preset ratio index data may be set according to actual application needs, and the present disclosure does not make any limitation.
[0123] In a specific embodiment, at least one initial text word segment may refer to text word segments with corresponding word frequency ratio index data greater than the preset ratio index data among the multiple text word segments.
[0124] In a specific embodiment, the preset word frequency index data may be set according to actual application needs, and the present disclosure does not make any limitation.
[0125] In a specific embodiment, at least one initial text segment in which the corresponding current word frequency index data or the corresponding historical word frequency index data is greater than or equal to a preset word frequency index data can be used as at least one second preset high-frequency text segment.
[0126] In the above embodiment, based on the current word frequency index data and the historical word frequency index data, word frequency ratio analysis is performed on each text segment to obtain the word frequency ratio index data corresponding to each text segment. From multiple text segments, the text segments with the corresponding word frequency ratio index data greater than the preset ratio index data are selected to obtain at least one initial text segment. The initial text segments in which both the corresponding current word frequency index data and the corresponding historical word frequency index data are less than the preset word frequency index data are deleted from the at least one initial text segment to obtain at least one second preset high-frequency text segment, which can avoid the frequent use of common words from being misselected as the second preset high-frequency text segment, improve the screening accuracy of the second preset high-frequency text segment, and further improve the accuracy of the predicted target topic text segment. It can be understood that the frequency of occurrence of some common words is very high, and the absolute value of the fluctuation will also be relatively high, and it is often easy to exceed the words with a small frequency base; in addition, for example, text segment W1 and text segment W2 appear 2000 times and 36 times respectively on the statistical day, and 100 times and 6 times respectively in the past two days. It can be determined that the word frequency ratio index data of text segment W1 is 2, and the word frequency ratio index data of text segment W2 is 6. Through the above method, text segment W1 can be determined as the second preset high-frequency text segment, that is, the accurate judgment of the second preset high-frequency text segment can be realized.
[0127] In a specific embodiment, the target model attribute information can characterize the model attribute features of a preset dialogue generation model. The target model attribute information can include "emotional communication robot" or "virtual anchor robot", etc. Among them, the preset dialogue generation model can be used to generate dialogue information for the target object based on the dialogue topic. Exemplarily, taking the application scenario where the attribute of the dialogue robot is an emotional communication robot as an example, the dialogue robot can generate dialogue information for the target object based on the preset dialogue generation model to achieve dialogue communication with the target object.
[0128] In a specific embodiment, the application developer can preset the model attribute information corresponding to the preset dialogue generation model in the dialogue robot based on the role development setting of the dialogue robot and store it in the preset memory of the server. Correspondingly, the target model attribute information corresponding to the preset dialogue generation model can be obtained from the above preset memory.
[0129] In a specific embodiment, a target topic generation instruction may be generated based on target associated feature information, at least one second preset high-frequency word segmentation, target model attribute information, and preset instruction prompt information. Specifically, two or more of the target associated feature information, at least one second preset high-frequency word segmentation, and target model attribute information may be added to multiple areas to be supplemented in the preset instruction prompt information to obtain the target topic generation instruction.
[0130] In the above embodiment, generating the target topic generation instruction based on the target associated feature information, at least one second preset high-frequency word segmentation, and target model attribute information can facilitate improving the diversity of the predicted target topic word segmentation on the basis of enhancing the prediction accuracy.
[0131] S205: Input the target topic generation instruction into a target topic generation model for dialogue topic prediction processing to obtain multiple target topic word segmentations corresponding to the target object.
[0132] In a specific embodiment, the target topic generation model may be used to generate multiple target topic word segmentations for the target object. The target topic generation model may be obtained by training a preset machine learning model based on target loss information. Specifically, the target topic generation model may include a natural language processing model. Optionally, the target topic generation model may include a ChatGPT (Chat Generative Pre-trained Transformer) or BERT (Bidirectional Encoder Representation from Transformers) model, etc.
[0133] In a specific embodiment, the target loss information may represent the degree of difference between the first topic quality information and the second topic quality information in the sample topic quality information. The sample topic quality information may be obtained by performing quality analysis processing based on multiple sample topic word segmentations and a preset reinforcement learning model. The sample topic quality information may include the quality information corresponding to each of the multiple sample topic word segmentations. The quality information corresponding to any sample topic word segmentation may represent the pros and cons of the above-mentioned any sample topic word segmentation among the multiple sample topic word segmentations. The first topic quality information may refer to the quality information corresponding to a preset sample topic word segmentation in the sample topic quality information. The second topic quality information may refer to the maximum quality information in the sample topic quality information. Among them, the preset sample topic word segmentation may be used to represent the expected result of the sample object. The preset sample topic word segmentation may be determined based on the topic word segmentation selected by the sample object. Specifically, the manifestation form of the sample topic quality information may include a vector, a matrix, etc., which is not limited in the present disclosure.
[0134] In a specific embodiment, multiple target topic word segments can characterize the conversation topics for an active conversation with a target object. Exemplarily, the multiple target topic word segments can include "light music" or "weather", etc.
[0135] In a specific embodiment, the target topic generation model and the target quality analysis model can be obtained in the following manner:
[0136] Obtain sample associated feature information;
[0137] Based on the sample associated feature information, generate a sample topic generation instruction;
[0138] Input the sample topic generation instruction into a preset machine learning model for conversation topic prediction processing to obtain multiple sample topic word segments corresponding to the sample object;
[0139] Obtain the preset sample topic word segments corresponding to the sample object;
[0140] Input the multiple sample topic word segments into a preset reinforcement learning model for quality analysis processing to obtain sample topic quality information;
[0141] Based on the sample topic quality information and the preset sample topic word segments, determine the target loss information;
[0142] Based on the target loss information, train the preset machine learning model and the preset reinforcement learning model to obtain the target topic generation model and the target quality analysis model.
[0143] In a specific embodiment, the target quality analysis model can be used to perform quality analysis on multiple target topic word segments. The target quality analysis model can be obtained by training the preset reinforcement learning model based on the target loss information.
[0144] In a specific embodiment, the sample associated feature information can characterize the associated features of the sample object in multiple dimensions. The sample associated feature information can include the sample historical conversation information corresponding to the sample object and the sample object attribute information corresponding to the sample object. Among them, the sample object can be an object used to train the target topic generation model and the target quality analysis model. The sample historical conversation information corresponding to the sample object can refer to the conversation information between the sample object and a preset conversation system within a historical time range. The above-mentioned preset conversation system and the target conversation system can be the same conversation system or different conversation systems. The sample object attribute information can characterize the attributes of the sample object from multiple perspectives.
[0145] In a specific embodiment, the sample topic generation instruction can be used to instruct the preset machine learning model to perform conversation topic prediction based on the sample associated feature information.
[0146] In a specific embodiment, a sample topic generation instruction may be generated based on sample association feature information and preset instruction prompt information. Specifically, the specific generation process of the sample topic generation instruction may refer to the generation process of the above-mentioned target topic generation instruction, which will not be elaborated in this disclosure.
[0147] In a specific embodiment, the above method may further include:
[0148] Obtain at least one first preset high-frequency word segment corresponding to the first historical time period and sample model attribute information;
[0149] Correspondingly, the generation of the sample topic generation instruction based on the sample association feature information may include:
[0150] Generate a sample topic generation instruction based on the sample association feature information, at least one first preset high-frequency word segment, and sample model attribute information.
[0151] In a specific embodiment, the first historical time period may refer to a preset time period before the first current moment. Among them, the first current moment may refer to the current moment in the model training process. Exemplarily, the first historical time period may refer to the current day in the model training process.
[0152] In a specific embodiment, at least one first preset high-frequency word segment may refer to a word segment whose frequency of occurrence in the first historical time period is higher than the preset frequency. At least one first preset high-frequency word segment may include popular words, etc.
[0153] In a specific embodiment, the acquisition method of at least one first preset high-frequency word segment may refer to the acquisition method of at least one second preset high-frequency word segment above, which will not be elaborated in this disclosure.
[0154] In a specific embodiment, the sample model attribute information may refer to the model attribute information of the dialogue generation model in the preset dialogue system corresponding to the sample historical dialogue information. The sample model attribute information may include "emotional communication robot" or "virtual anchor robot", etc.
[0155] In a specific embodiment, a sample topic generation instruction may be generated based on the sample association feature information, at least one first preset high-frequency word segment, sample model attribute information, and preset instruction prompt information. Specifically, two or more of the sample association feature information, at least one first preset high-frequency word segment, and sample model attribute information may be added to multiple areas to be supplemented in the preset instruction prompt information to obtain the sample topic generation instruction.
[0156] In a specific embodiment, multiple sample topic word segments may represent the dialogue topics for the active dialogue of the sample object.
[0157] In a specific embodiment, the preset machine learning model may include a sample encoding module, a sample fine-tuning module, a sample feature processing module, and a sample prediction module.
[0158] In a specific embodiment, the above-mentioned inputting the sample topic generation instruction into the preset machine learning model for dialogue topic prediction processing to obtain multiple sample topic segmentations corresponding to the sample object may include:
[0159] Input the sample topic generation instruction into the sample encoding module for encoding processing to obtain sample instruction encoding information;
[0160] Input the sample instruction encoding information into the sample fine-tuning module for feature processing to obtain the first sample feature information;
[0161] Input the sample instruction encoding information into the sample feature processing module for feature processing to obtain the second sample feature information;
[0162] Input the first sample feature information and the second sample feature information into the sample prediction module for topic prediction processing to obtain multiple sample topic segmentations.
[0163] In a specific embodiment, the sample instruction encoding information can be used to represent the semantic features of the sample topic generation instruction. Specifically, the manifestation form of the sample instruction encoding information may include vectors or matrices, etc. It can be understood that the sample instruction encoding information can represent the underlying semantic features of the sample topic generation instruction.
[0164] In a specific embodiment, the first sample feature information can represent the high-level semantic features of the sample topic generation instruction. Specifically, the manifestation form of the first sample feature information may include vectors or matrices, etc.
[0165] In a specific embodiment, the sample fine-tuning module may adopt a low-rank matrix structure, that is, the original d*d-dimensional matrix W is represented by the product of a d*r-dimensional matrix A and an r*d-dimensional matrix B through matrix decomposition, and r << d.
[0166] In a specific embodiment, the second sample feature information can represent the high-level semantic features of the sample topic generation instruction. Specifically, the manifestation form of the second sample feature information may include vectors or matrices, etc.
[0167] In a specific embodiment, the sample feature processing module may include multiple feature extraction layers; the high-level semantic features in the sample instruction encoding information can be extracted through the sample feature processing module.
[0168] In a specific embodiment, the sample prediction module may include a sample feature superposition layer, a sample prediction layer, and a sample output layer. Specifically, inputting the first sample feature information and the second sample feature information into the sample feature superposition layer for superposition processing can obtain the third sample feature information; inputting the third sample feature information into the sample prediction layer for prediction processing can obtain a sample prediction sequence; inputting the sample prediction sequence into the sample output layer for decoding processing can obtain multiple sample topic segmentations. Further, the third sample feature information and the preset start identification information may be first input into the sample prediction layer for text prediction to obtain the identification information at the next position in the sample prediction sequence; then, based on the identification information at the next position in the sample prediction sequence, the preset start identification information is updated, and based on the updated preset start identification information, the step of inputting the third sample feature information and the preset start identification information into the sample prediction layer for text prediction to obtain the identification information at the next position in the sample prediction sequence is repeated until the result output by the sample prediction layer is the preset end identification information. The sample prediction sequence can be generated in the order from the first to the last of the identification information output by the sample prediction layer, in combination with the preset start identification information.
[0169] In a specific embodiment, the preset sample topic segmentation may represent the expected result of the sample object. The preset sample topic segmentation may be determined based on the historical selected topic segmentations corresponding to the sample object.
[0170] In a specific embodiment, the obtaining of the preset sample topic segmentation corresponding to the sample object may include:
[0171] Obtain the historical selected topic segmentations corresponding to the sample object;
[0172] Perform similarity analysis on each sample topic segmentation in the historical selected topic segmentations and the multiple sample topic segmentations to obtain the similarity index data corresponding to each sample topic segmentation;
[0173] Take the sample topic segmentation with the largest corresponding similarity index data among the multiple sample topic segmentations as the preset sample topic segmentation.
[0174] In a specific embodiment, the historical selected topic word segments corresponding to the sample object may refer to the topic word segments selected by the sample object through performing a selection operation when the sample object is a historical object. Among them, the historical object may refer to an object that actively conducts a conversation through a dialogue system corresponding to a preset machine learning model during the application process of the model. Specifically, during the application process of the preset machine learning model by the above-mentioned dialogue system, multiple historical topic word segments for the above-mentioned sample object may be generated based on the preset machine learning model; in response to a selection instruction corresponding to the sample object, the historical selected topic word segments may be obtained, and the historical selected topic word segments may be stored based on the corresponding relationship between the sample object and the historical selected topic word segments; among them, the historical selected topic word segments may be any one of the above-mentioned multiple historical topic word segments. Correspondingly, the historical selected topic word segments corresponding to the sample object may be obtained from a preset memory.
[0175] In a specific embodiment, the similarity index data corresponding to any sample topic word segment may characterize the similarity degree between the historical selected topic word segments and the above-mentioned any sample topic word segment.
[0176] In a specific embodiment, by performing semantic distance analysis on the historical selected topic word segments and any sample topic word segment, the similarity index data corresponding to the above-mentioned any sample topic word segment may be obtained.
[0177] In a specific embodiment, based on the similarity index data corresponding to multiple sample topic word segments, a preset sample topic word segment may be determined from the multiple sample topic word segments. Specifically, the preset sample topic word segment may be the sample topic word segment with the largest corresponding similarity index data among the multiple sample topic word segments.
[0178] In a specific embodiment, in the case where there is a sample topic word segment identical to the historical selected topic word segment among the multiple sample topic word segments, the above-mentioned historical selected topic word segment may be used as the preset sample topic word segment.
[0179] In the above embodiment, by obtaining the historical selected topic word segments corresponding to the sample object, performing similarity analysis on the historical selected topic word segments and each sample topic word segment among the multiple sample topic word segments, obtaining the similarity index data corresponding to each sample topic word segment, and using the sample topic word segment with the largest corresponding similarity index data among the multiple sample topic word segments as the preset sample topic word segment, the feedback data return of the dialogue topic selected by the online user can be realized. Furthermore, by combining the target loss information to train the model, the relevance between the topic result predicted by the model and the topic result preferred by the target object can be improved, and the prediction accuracy of the target quality analysis model and the target topic generation model can be improved.
[0180] In a specific embodiment, in the case of obtaining multiple sample topic word segments through dialogue topic prediction processing, it may be that the model developer selects one of the multiple sample topic word segments by performing a sample selection operation; correspondingly, in response to the sample selection operation instruction, the preset sample topic word segment corresponding to the sample object can be obtained.
[0181] In a specific embodiment, the sample topic quality information can characterize the quality level of each sample topic word segment among the multiple sample topic word segments. The sample topic quality information can include the quality information corresponding to each of the multiple sample topic word segments.
[0182] In a specific embodiment, the quality information corresponding to the preset sample topic word segment can be searched from the sample topic quality information, and the quality information corresponding to the preset sample topic word segment can be used as the first topic quality information. The maximum quality information can be filtered out from the sample topic quality information and used as the second topic quality information. Based on the above first topic quality information and the second topic quality information, the target loss information can be determined. Specifically, the target loss information can be determined based on the above first topic quality information, the second topic quality information, and the cross-entropy loss function.
[0183] In a specific embodiment, the training of the preset machine learning model and the preset reinforcement learning model based on the target loss information to obtain the target topic generation model and the target quality analysis model may include:
[0184] Training the sample fine-tuning module and the sample prediction module in the preset machine learning model based on the target loss information to obtain the target topic generation model;
[0185] Training the preset reinforcement learning model based on the target loss information to obtain the target quality analysis model.
[0186] In a specific embodiment, the parameter update gradient can be determined based on the target loss information; based on the above parameter update gradient, the sample fine-tuning module and the sample prediction module in the preset machine learning model can be trained to obtain the target topic generation model.
[0187] In a specific embodiment, the parameter update gradient can be determined based on the target loss information; based on the above parameter update gradient, the preset reinforcement learning model can be trained to obtain the target quality analysis model.
[0188] In the above embodiments, based on the target loss information, the sample fine-tuning module and the sample prediction module in the preset machine learning model are trained to obtain the target topic generation model. On the basis of greatly reducing the training cost by training the sample fine-tuning module, a small number of trainable parameters can be added by training the sample prediction module, so that the target topic generation model has better fitting ability, thereby improving the prediction accuracy of the target topic generation model.
[0189] In a specific embodiment, based on the target loss information, some parameters in the sample feature processing module, the sample fine-tuning module, and the sample prediction module in the preset machine learning model can be trained to obtain the target topic generation model.
[0190] In the above embodiments, based on the target loss information, some parameters in the sample feature processing module, the sample fine-tuning module, and the sample prediction module in the preset machine learning model are trained to obtain the target topic generation model. On the basis of greatly reducing the training cost, a small number of trainable parameters can be added by training some parameters in the sample feature processing module, the sample fine-tuning module, and the sample prediction module, so that the target topic generation model has better fitting ability, thereby improving the prediction accuracy of the target topic generation model.
[0191] In the above embodiments, by obtaining the sample association feature information, based on the sample association feature information, generating a sample topic generation instruction, inputting the sample topic generation instruction into the preset machine learning model for dialogue topic prediction processing, obtaining multiple sample topic word segments corresponding to the sample object, obtaining the preset sample topic word segments corresponding to the sample object, inputting the multiple sample topic word segments into the preset reinforcement learning model for quality analysis processing, obtaining the sample topic quality information, based on the sample topic quality information and the preset sample topic word segments, determining the target loss information, and based on the target loss information, training the preset machine learning model and the preset reinforcement learning model to obtain the target topic generation model and the target quality analysis model, the relevance between the topic result predicted by the model and the topic result preferred by the target object can be improved, thereby improving the prediction accuracy of the target quality analysis model and the target topic generation model.
[0192] In a specific embodiment, the target topic generation model may include a target encoding module, a target fine-tuning module, a target feature processing module, and a target prediction module.
[0193] In a specific embodiment, the above step S205 may include:
[0194] Input the target topic generation instruction into the target encoding module for encoding processing to obtain the target instruction encoding information;
[0195] Input the target instruction encoding information into the target fine-tuning module for feature processing to obtain the first target feature information;
[0196] Input the target instruction encoding information into the target feature processing module for feature processing to obtain the second target feature information;
[0197] Input the first target feature information and the second target feature information into the target prediction module for topic prediction processing to obtain multiple target topic word segments.
[0198] In a specific embodiment, the target instruction encoding information can represent the semantic features of the target topic generation instruction. Specifically, the manifestation form of the target instruction encoding information can include vectors or matrices, etc.
[0199] In a specific embodiment, the first target feature information can represent the high-level semantic features of the target topic generation instruction.
[0200] In a specific embodiment, the second target feature information can represent the high-level semantic features of the target topic generation instruction. Specifically, the manifestation form of the second target feature information can include vectors or matrices, etc.
[0201] In a specific embodiment, the above-mentioned target encoding module, target fine-tuning module, target feature processing module, and target prediction module can respectively refer to the specific processing procedures of the above-mentioned sample encoding module, above-mentioned sample fine-tuning module, above-mentioned sample feature processing module, and above-mentioned sample prediction module, which will not be elaborated in this disclosure.
[0202] In the above embodiment, input the target topic generation instruction into the target encoding module for encoding processing to obtain the target instruction encoding information, input the target instruction encoding information into the target fine-tuning module for feature processing to obtain the first target feature information, input the target instruction encoding information into the target feature processing module for feature processing to obtain the second target feature information, input the first target feature information and the second target feature information into the target prediction module for topic prediction processing to obtain multiple target topic word segments. Through the target fine-tuning module and the target feature processing module, a multi-layer learnable method can be implemented to support the extraction of high-level semantic features, thereby improving the prediction effect and semantic expression ability of the target topic generation model, and improving the semantic accuracy of the generated target topic word segments.
[0203] S207: Input multiple target topic word segments into the target quality analysis model for quality analysis processing to obtain the target topic quality information.
[0204] In a specific embodiment, the target quality analysis model can be obtained by training a preset reinforcement learning model based on target loss information. Among them, the target loss information can characterize the degree of difference between the first topic quality information and the second topic quality information in the sample topic quality information. The first topic quality information can be the quality information corresponding to the preset sample topic word segmentation in the sample topic quality information. The second topic quality information can be the maximum quality information in the sample topic quality information.
[0205] In a specific embodiment, multiple target topic word segmentations are input into the target quality analysis model, and the target quality analysis model can perform quality analysis on each target topic word segmentation in the multiple target topic word segmentations to obtain the quality information corresponding to each target topic word segmentation; correspondingly, the target topic quality information can be obtained.
[0206] S209: Generate target question information for the target object based on multiple target topic word segmentations, target topic quality information, and a preset dialogue generation model.
[0207] In a specific embodiment, the preset dialogue generation model can be used to generate question information for an active dialogue. Specifically, the preset dialogue generation model can include a natural language processing model. Optionally, the preset dialogue generation model can include models such as ChatGPT (Chat Generative Pre-trained Transformer) or BERT (Bidirectional Encoder Representation from Transformers).
[0208] In a specific embodiment, the target question information can refer to the question information actively proposed by the dialogue system to the target object. Exemplarily, the target question information can be "What do you think of XX's music?".
[0209] In a specific embodiment, at least one target topic word segmentation can be selected from the multiple target topic word segmentations based on the target topic quality information, and the at least one target topic word segmentation is input into the preset dialogue generation model for text generation processing, and target question information for the target object can be obtained.
[0210] In a specific embodiment, based on the target topic quality information, the multiple target topic word segmentations are sorted in descending order of quality information to obtain a target topic word segmentation sequence; correspondingly, the first preset number of target topic word segmentations selected from the target topic word segmentation sequence can be used as the at least one target topic word segmentation and input into the preset dialogue generation model for text generation processing, and target question information for the target object can be obtained.
[0211] In the above embodiments, by obtaining the target associated feature information corresponding to the target object, where the target associated feature information includes the target historical conversation information corresponding to the target object and the target object attribute information corresponding to the target object, the acquisition of associated features in multiple dimensions of the target object can be realized. Then, by combining the target associated feature information, a target topic generation instruction can be generated, which can facilitate the accurate prediction of the target topic word segmentation associated with the target object. Secondly, inputting the target topic generation instruction into the target topic generation model for dialogue topic prediction processing, multiple target topic word segmentations corresponding to the target object can be obtained. The target topic generation model is obtained by training a preset machine learning model based on the target loss information, which can realize the efficient prediction of the dialogue topic for the active dialogue of the target object. Then, inputting the multiple target topic word segmentations into the target quality analysis model for quality analysis processing, target topic quality information can be obtained. The target quality analysis model is obtained by training a preset reinforcement learning model based on the target loss information. The target loss information represents the degree of difference between the first topic quality information and the second topic quality information in the sample topic quality information. The first topic quality information is the quality information corresponding to the preset sample topic word segmentation in the sample topic quality information, and the second topic quality information is the largest quality information in the sample topic quality information. The preset sample topic word segmentation is determined based on the topic word segmentation selected by the sample object. The model can be trained through the above target loss information to improve the relevance between the topic result predicted by the model and the topic result preferred by the target object, thereby improving the prediction accuracy of the target quality analysis model and the target topic generation model. Then, by combining the multiple target topic word segmentations, the target topic quality information and the preset dialogue generation model, target question information for the target object can be generated, which can make the target question information closer to the preferences of the target object, thereby increasing the conversion rate of the active dialogue, improving user satisfaction and user retention rate.
[0212] Figure 3 It is a flowchart of a question generation method shown according to an exemplary embodiment. Specifically, as Figure 3As shown, the target historical dialogue information corresponding to the target object, the target object attribute information corresponding to the target object, at least one second preset high-frequency word segment corresponding to the second historical time period, and the target model attribute information corresponding to the preset dialogue generation model can be obtained; based on the target association feature information, at least one second preset high-frequency word segment, and the target model attribute information, a target topic generation instruction can be generated; inputting the target topic generation instruction into the target topic generation model for dialogue topic prediction processing, multiple target topic word segments corresponding to the target object can be obtained; inputting the multiple target topic word segments into the target quality analysis model for quality analysis processing, target topic quality information can be obtained; based on the multiple target topic word segments, the target topic quality information, and the preset dialogue generation model, target question information for the target object can be generated.
[0213] Figure 4 It is a block diagram of a question generation device shown according to an exemplary embodiment. As Figure 4 shown, the device may include:
[0214] A target feature acquisition module 410, which can be used to acquire target association feature information corresponding to a target object, and the target association feature information includes target historical dialogue information corresponding to the target object and target object attribute information corresponding to the target object;
[0215] A target instruction generation module 420, which can be used to generate a target topic generation instruction based on the target association feature information;
[0216] A target topic prediction module 430, which can be used to input the target topic generation instruction into the target topic generation model for dialogue topic prediction processing to obtain multiple target topic word segments corresponding to the target object; the target topic generation model is obtained by training a preset machine learning model based on target loss information;
[0217] A first quality analysis module 440, which can be used to input the multiple target topic word segments into the target quality analysis model for quality analysis processing to obtain target topic quality information; the target quality analysis model is obtained by training a preset reinforcement learning model based on target loss information, and the target loss information represents the difference degree between the first topic quality information and the second topic quality information in the sample topic quality information, the first topic quality information is the quality information corresponding to the preset sample topic word segment in the sample topic quality information, and the second topic quality information is the maximum quality information in the sample topic quality information; the preset sample topic word segment is determined based on the topic word segment selected by the sample object;
[0218] A target question generation module 450, which can be used to generate target question information for the target object based on the multiple target topic word segments, the target topic quality information, and the preset dialogue generation model.
[0219] In a specific embodiment, the above device may further include:
[0220] A sample feature acquisition module, which can be used to acquire sample associated feature information, where the sample associated feature information includes sample historical dialogue information corresponding to the sample object and sample object attribute information corresponding to the sample object;
[0221] A sample instruction generation module, which can be used to generate a sample topic generation instruction based on the sample associated feature information;
[0222] A sample topic prediction module, which can be used to input the sample topic generation instruction into a preset machine learning model for dialogue topic prediction processing to obtain multiple sample topic word segments corresponding to the sample object;
[0223] A selected topic acquisition module, which can be used to acquire a preset sample topic word segment corresponding to the sample object;
[0224] A second quality analysis module, which can be used to input the multiple sample topic word segments into a preset reinforcement learning model for quality analysis processing to obtain sample topic quality information;
[0225] A loss determination module, which can be used to determine target loss information based on the sample topic quality information and the preset sample topic word segment;
[0226] A model training module, which can be used to train the preset machine learning model and the preset reinforcement learning model based on the target loss information to obtain a target topic generation model and a target quality analysis model.
[0227] In a specific embodiment, the above sample topic prediction module may include:
[0228] A first encoding processing module, which can be used to input the sample topic generation instruction into a sample encoding module for encoding processing to obtain sample instruction encoding information;
[0229] A first feature processing module, which can be used to input the sample instruction encoding information into a sample fine-tuning module for feature processing to obtain first sample feature information;
[0230] A second feature processing module, which can be used to input the sample instruction encoding information into a sample feature processing module for feature processing to obtain second sample feature information;
[0231] A first prediction processing module, which can be used to input the first sample feature information and the second sample feature information into a sample prediction module for topic prediction processing to obtain multiple sample topic word segments.
[0232] In a specific embodiment, the above model training module may include:
[0233] The first training module can be used to train the sample fine-tuning module and the sample prediction module in a preset machine learning model based on target loss information to obtain a target topic generation model;
[0234] The second training module can be used to train a preset reinforcement learning model based on target loss information to obtain a target quality analysis model.
[0235] In a specific embodiment, the above-mentioned selected topic acquisition module may include:
[0236] The historical selection acquisition module can be used to acquire the historical selected topic word segmentation corresponding to the sample object;
[0237] The similarity analysis module can be used to perform similarity analysis on the historical selected topic word segmentation and each sample topic word segmentation in multiple sample topic word segmentations to obtain the corresponding similarity index data for each sample topic word segmentation;
[0238] The selected topic determination module can be used to use the sample topic word segmentation with the largest corresponding similarity index data among multiple sample topic word segmentations as the preset sample topic word segmentation.
[0239] In a specific embodiment, the above-mentioned device may further include:
[0240] The sample information acquisition module can be used to acquire at least one first preset high-frequency word segmentation and sample model attribute information corresponding to the first historical time period; the first historical time period is a preset time period before the first current moment;
[0241] Correspondingly, the above-mentioned sample instruction generation module may include:
[0242] The first instruction generation module can be used to generate a sample topic generation instruction based on the sample association feature information, at least one first preset high-frequency word segmentation, and the sample model attribute information.
[0243] In a specific embodiment, the above-mentioned device may further include:
[0244] The target information acquisition module can be used to acquire at least one second preset high-frequency word segmentation and target model attribute information corresponding to the second historical time period; the second historical time period is a preset time period before the second current moment;
[0245] Correspondingly, the above-mentioned target instruction generation module 420 may include:
[0246] The second instruction generation module can be used to generate a target topic generation instruction based on the target association feature information, at least one second preset high-frequency word segmentation, and the target model attribute information.
[0247] In a specific embodiment, the above-mentioned target information acquisition module may include:
[0248] A short text acquisition module, which can be used to acquire a plurality of preset short text information corresponding to a second historical time period from at least one preset information platform;
[0249] A word segmentation processing module, which can be used to perform word segmentation processing on the plurality of preset short text information to obtain a plurality of text word segments;
[0250] A text word segment determination module, which can be used to determine a variety of text word segments based on the plurality of text word segments;
[0251] A frequency analysis module, which can be used to perform frequency analysis on the plurality of text word segments to obtain current word frequency index data corresponding to each of the variety of text word segments;
[0252] A historical word frequency acquisition module, which can be used to acquire historical word frequency index data corresponding to each of the variety of text word segments;
[0253] A high-frequency word segment determination module, which can be used to determine at least one second preset high-frequency word segment from the variety of text word segments based on the current word frequency index data and the historical word frequency index data.
[0254] In a specific embodiment, the above-mentioned high-frequency word segment determination module may include:
[0255] A word frequency ratio analysis module, which can be used to perform word frequency ratio analysis on each text word segment based on the current word frequency index data and the historical word frequency index data to obtain word frequency ratio index data corresponding to each text word segment; the word frequency ratio index data corresponding to each text word segment represents the ratio between the current word frequency index data corresponding to each text word segment and the historical word frequency index data corresponding to each text word segment;
[0256] An initial word segment acquisition module, which can be used to select text word segments from the variety of text word segments whose corresponding word frequency ratio index data is greater than the preset ratio index data to obtain at least one initial text word segment;
[0257] A word segment screening module, which can be used to delete the initial text word segments in the at least one initial text word segment whose corresponding current word frequency index data and corresponding historical word frequency index data are both less than the preset word frequency index data to obtain at least one second preset high-frequency word segment.
[0258] In a specific embodiment, the above-mentioned target topic prediction module 430 may include:
[0259] A second encoding processing module, which can be used to input the target topic generation instruction into the target encoding module for encoding processing to obtain target instruction encoding information;
[0260] The third feature processing module can be used to input the target instruction encoding information into the target fine-tuning module for feature processing to obtain the first target feature information;
[0261] The fourth feature processing module can be used to input the target instruction encoding information into the target feature processing module for feature processing to obtain the second target feature information;
[0262] The second prediction processing module can be used to input the first target feature information and the second target feature information into the target prediction module for topic prediction processing to obtain multiple target topic word segments.
[0263] Regarding the device in the above embodiments, the specific manners in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0264] Figure 5 is a block diagram of an electronic device for generating target question information shown according to an exemplary embodiment. The electronic device can be a server, and its internal structure diagram can be as Figure 5 shown. The electronic device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a question generation method.
[0265] Figure 6 is a block diagram of another electronic device for generating target question information shown according to an exemplary embodiment. The electronic device can be a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a question generation method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0266] Those skilled in the art can understand that Figure 5 or Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0267] In an exemplary embodiment, an electronic device is further provided, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the question generation method as in the embodiment of the present disclosure.
[0268] In an exemplary embodiment, a computer-readable storage medium is further provided. When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the question generation method in the embodiment of the present disclosure.
[0269] In an exemplary embodiment, a computer program product including instructions is further provided. When it runs on a computer, the computer executes the question generation method in the embodiment of the present disclosure.
[0270] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When this computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0271] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.
[0272] It can be understood that in the specific implementation manners of this application, when it comes to relevant data such as user information, when the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0273] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0274] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for generating questions, characterized in that, The method includes: Obtaining target associated feature information corresponding to a target object, where the target associated feature information includes target historical conversation information corresponding to the target object and target object attribute information corresponding to the target object; Generating a target topic generation instruction based on the target associated feature information; Inputting the target topic generation instruction into a target topic generation model for dialogue topic prediction processing to obtain multiple target topic word segments corresponding to the target object; the target topic generation model is obtained by training a preset machine learning model based on target loss information; Inputting the multiple target topic word segments into a target quality analysis model for quality analysis processing to obtain target topic quality information; the target quality analysis model is obtained by training a preset reinforcement learning model based on the target loss information, and the target loss information represents the degree of difference between a first topic quality information and a second topic quality information in sample topic quality information, the first topic quality information is the quality information corresponding to a preset sample topic word segment in the sample topic quality information, the second topic quality information is the maximum quality information in the sample topic quality information; the preset sample topic word segment is determined based on the topic word segment selected by the sample object; Generating target question information for the target object based on the multiple target topic word segments, the target topic quality information, and a preset dialogue generation model.
2. The method according to claim 1, wherein The target topic generation model and the target quality analysis model are obtained in the following manner: Obtaining sample associated feature information, where the sample associated feature information includes sample historical conversation information corresponding to a sample object and sample object attribute information corresponding to the sample object; Generating a sample topic generation instruction based on the sample associated feature information; Inputting the sample topic generation instruction into the preset machine learning model for dialogue topic prediction processing to obtain multiple sample topic word segments corresponding to the sample object; Obtaining a preset sample topic word segment corresponding to the sample object; Inputting the multiple sample topic word segments into the preset reinforcement learning model for quality analysis processing to obtain the sample topic quality information; Determining the target loss information based on the sample topic quality information and the preset sample topic word segment; Training the preset machine learning model and the preset reinforcement learning model based on the target loss information to obtain the target topic generation model and the target quality analysis model.
3. The method according to claim 2, wherein The preset machine learning model includes a sample encoding module, a sample fine-tuning module, a sample feature processing module, and a sample prediction module; the step of inputting the sample topic generation instruction into the preset machine learning model for dialogue topic prediction processing to obtain multiple sample topic word segments corresponding to the sample object includes: Inputting the sample topic generation instruction into the sample encoding module for encoding processing to obtain sample instruction encoding information; Inputting the sample instruction encoding information into the sample fine-tuning module for feature processing to obtain first sample feature information; Input the sample instruction encoding information into the sample feature processing module for feature processing to obtain second sample feature information; Input the first sample feature information and the second sample feature information into the sample prediction module for topic prediction processing to obtain the multiple sample topic word segments.
4. The method according to claim 3, wherein Training the preset machine learning model and the preset reinforcement learning model based on the target loss information to obtain the target topic generation model and the target quality analysis model includes: Training the sample fine-tuning module and the sample prediction module in the preset machine learning model based on the target loss information to obtain the target topic generation model; Training the preset reinforcement learning model based on the target loss information to obtain the target quality analysis model.
5. The method according to claim 2, wherein Obtaining the preset sample topic word segment corresponding to the sample object includes: Obtaining the historical selected topic word segment corresponding to the sample object; Performing similarity analysis on the historical selected topic word segment and each sample topic word segment in the multiple sample topic word segments to obtain the corresponding similarity index data for each sample topic word segment; Taking the sample topic word segment with the largest corresponding similarity index data in the multiple sample topic word segments as the preset sample topic word segment.
6. The method according to claim 2, wherein The method further includes: Obtaining at least one first preset high-frequency word segment and sample model attribute information corresponding to a first historical time period; the first historical time period is a preset time period before a first current moment; Generating a sample topic generation instruction based on the sample association feature information includes: Generating the sample topic generation instruction based on the sample association feature information, the at least one first preset high-frequency word segment, and the sample model attribute information.
7. The method according to claim 1, characterized in that, The method further includes: Obtaining at least one second preset high-frequency word segment and target model attribute information corresponding to a second historical time period; the second historical time period is a preset time period before a second current moment; Generating a target topic generation instruction based on the target association feature information includes: Generating the target topic generation instruction based on the target association feature information, the at least one second preset high-frequency word segment, and the target model attribute information.
8. The method according to claim 7, wherein The at least one second preset high-frequency word segment is obtained in the following manner: Obtaining multiple preset short text information corresponding to the second historical time period from at least one preset information platform; Performing word segmentation processing on the multiple preset short text information to obtain multiple text word segments; Determining multiple text word segments based on the multiple text word segments; Performing frequency analysis on the multiple text word segments to obtain the corresponding current word frequency index data for each of the multiple text word segments; Obtaining the historical word frequency index data corresponding to each of the multiple text word segments; Determining the at least one second preset high-frequency word segment from the multiple text word segments based on the current word frequency index data and the historical word frequency index data.
9. The method according to claim 8, characterized in that Determining the at least one second preset high-frequency word segmentation from the multiple text word segmentations based on the current word frequency index data and the historical word frequency index data includes: Performing word frequency ratio analysis on each text word segmentation based on the current word frequency index data and the historical word frequency index data to obtain word frequency ratio index data corresponding to each text word segmentation; the word frequency ratio index data corresponding to each text word segmentation represents the ratio between the current word frequency index data corresponding to each text word segmentation and the historical word frequency index data corresponding to each text word segmentation; Selecting, from the multiple text word segmentations, the text word segmentations whose corresponding word frequency ratio index data is greater than the preset ratio index data to obtain at least one initial text word segmentation; Deleting the initial text word segmentations in the at least one initial text word segmentations whose corresponding current word frequency index data and corresponding historical word frequency index data are both less than the preset word frequency index data to obtain the at least one second preset high-frequency word segmentation.
10. The method according to any one of claims 1-9, characterized in that The target topic generation model includes a target encoding module, a target fine-tuning module, a target feature processing module, and a target prediction module; inputting the target topic generation instruction into the target topic generation model for dialogue topic prediction processing to obtain multiple target topic word segmentations corresponding to the target object includes: Inputting the target topic generation instruction into the target encoding module for encoding processing to obtain target instruction encoding information; Inputting the target instruction encoding information into the target fine-tuning module for feature processing to obtain first target feature information; Inputting the target instruction encoding information into the target feature processing module for feature processing to obtain second target feature information; Inputting the first target feature information and the second target feature information into the target prediction module for topic prediction processing to obtain the multiple target topic word segmentations.
11. A question generation device, characterized in that The device includes: A target feature acquisition module, configured to acquire target associated feature information corresponding to a target object, where the target associated feature information includes target historical dialogue information corresponding to the target object and target object attribute information corresponding to the target object; A target instruction generation module, configured to generate a target topic generation instruction based on the target associated feature information; A target topic prediction module, configured to input the target topic generation instruction into a target topic generation model for dialogue topic prediction processing to obtain multiple target topic word segmentations corresponding to the target object; the target topic generation model is obtained by training a preset machine learning model based on target loss information; The first quality analysis module is configured to input the multiple target topic word segmentations into a target quality analysis model for quality analysis processing to obtain target topic quality information; the target quality analysis model is obtained by training a preset reinforcement learning model based on the target loss information, the target loss information represents the degree of difference between the first topic quality information and the second topic quality information in the sample topic quality information, the first topic quality information is the quality information corresponding to a preset sample topic word segmentation in the sample topic quality information, and the second topic quality information is the maximum quality information in the sample topic quality information; the preset sample topic word segmentation is determined based on the topic word segmentation selected by the sample object. The target question generation module is configured to generate target question information for the target object based on the multiple target topic word segmentations, the target topic quality information, and a preset dialogue generation model.
12. An electronic device, characterized in that, Comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the executable instructions to implement the question generation method according to any one of claims 1 to 10.
13. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the question generation method according to any one of claims 1 to 10 is implemented.