Information extraction method, dialogue method, electronic equipment, storage medium and product
By generating and using the first prompt text and the second prompt text, combined with the pre-trained language model, directly extracting and indirectly inferring user information, the problem of low accuracy of information extraction in the prior art is solved, and the intelligent and personalized service capabilities of the dialogue assistant are improved.
Patent Information
- Application Number
- CN202510119432.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, the interaction content between users and dialogue assistants is complex and diverse, resulting in low accuracy of information extraction.
By obtaining the user's multiple rounds of conversation content, the first prompt text and the second prompt text are generated, and the user information is directly extracted and indirectly inferred from the multiple rounds of conversation content using the pre-trained language model.
It improves the depth and breadth of understanding of user intentions, improves the accuracy of information extraction, and allows conversation assistants to better understand user needs and preferences.
Smart Images

Figure CN119940372A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this application relate to the field of computer software technology, and in particular to an information extraction method, a dialogue method, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the widespread use of conversational assistants (such as intelligent voice assistants and chatbots), user interactions with these assistants are frequent and diverse. To provide more accurate services, one approach in related technologies is to extract user-related information based on the conversation between the assistant and the user. This helps the assistant better understand the user's needs, preferences, and personalized information in subsequent conversations. However, the content of user interactions with conversational assistants is often complex and diverse, resulting in relatively low accuracy for information extraction methods in these technologies. Summary of the Invention
[0003] In view of the above, one or more embodiments of this application provide an information extraction method, a dialogue method, an electronic device, a computer-readable storage medium, and a computer program product.
[0004] To achieve the above objectives, one or more embodiments of this application provide the following technical solutions:
[0005] According to a first aspect of one or more embodiments of this application, an information extraction method is proposed, comprising:
[0006] Obtain the content of multiple rounds of user conversations;
[0007] Based on the content of the multi-turn dialogue, generate a first prompt text and a second prompt text;
[0008] The first prompt text is input into a pre-trained language model, so that the pre-trained language model can directly extract target user information from the multi-turn dialogue content based on the first prompt text; and...
[0009] The second prompt text is input into the pre-trained language model so that the pre-trained language model can indirectly infer the target user information from the multi-turn dialogue content based on the second prompt text.
[0010] In one implementation, the first prompt text includes a first processing requirement for the multi-turn dialogue content and the multi-turn dialogue content; wherein, the first processing requirement is at least used to instruct the pre-trained language model to extract directly stated target user information from the multi-turn dialogue content;
[0011] And / or,
[0012] The second prompt text includes a second processing requirement for the multi-turn dialogue content and the multi-turn dialogue content; wherein, the second processing requirement is at least used to instruct the pre-trained language model to determine the target topic that appears more frequently than a set threshold from the multi-turn dialogue content, infer target user information based on the target topic and provide the reason for the inference.
[0013] In one implementation, the first prompt text further includes at least one of the following: role setting information, the type of user information to be extracted, and at least one processing example that meets the first processing requirement;
[0014] And / or,
[0015] The second prompt text also includes at least one of the following: role setting information, the type of user information to be extracted, and at least one processing example that meets the second processing requirements.
[0016] In one implementation, obtaining the multi-turn dialogue content between the user and the dialogue assistant includes:
[0017] In response to the fact that the dialogue content between the user and the dialogue assistant meets the preset conditions, the system obtains the multi-turn dialogue content between the user and the dialogue assistant.
[0018] The preset conditions include any one of the following: the number of rounds in the dialogue is greater than the set number of rounds, the number of topics involved in the dialogue exceeds the set number, the number of questions asked about the same topic in the dialogue exceeds the set number, or the frequency of use of a specific keyword in the dialogue exceeds the set frequency.
[0019] In one implementation, generating the first prompt text and the second prompt text based on the multi-turn dialogue content includes:
[0020] The multi-turn dialogue content is embedded into a preset first prompt template to generate the first prompt text; and
[0021] The multi-turn dialogue content is embedded into a preset second prompt template to generate the second prompt text.
[0022] In one implementation, the first prompt template and the second prompt template are obtained in the following way:
[0023] The content of the multi-turn dialogue is analyzed to determine the target expression style;
[0024] Based on the target expression style, a first prompt template that matches the target expression style is obtained from a first prompt template library, and a second prompt template that matches the target expression style is obtained from a second prompt template library;
[0025] The first prompt template library pre-stores first prompt templates corresponding to different expression styles; the second prompt template library pre-stores second prompt templates corresponding to different expression styles.
[0026] One implementation also includes:
[0027] After obtaining the target user information output by the pre-trained model, it is detected whether the user database stores historical user information of the same type as the target user information;
[0028] If the user database does not store historical user information of the same type as the target user information, the target user information is stored in the user database.
[0029] One implementation also includes:
[0030] If the user database stores historical user information of the same type as the target user information, retrieve the historical user information from the user database.
[0031] Determine the semantic similarity between the target user information and the historical user information;
[0032] If the semantic similarity is greater than the first similarity threshold, the target user information output by the pre-trained model is discarded;
[0033] If the semantic similarity is less than the second similarity threshold, the historical user information is deleted, and the target user information is stored in the user database;
[0034] If the semantic similarity is between the first similarity threshold and the second similarity threshold, the target user information and the historical user information are merged, and the merged user information is stored in the user database.
[0035] Wherein, the first similarity threshold is greater than the second similarity threshold.
[0036] According to a second aspect of the embodiments of this application, a dialogue method is provided, including:
[0037] The system retrieves questions entered by the user during a conversation with a dialogue assistant; wherein the dialogue assistant is deployed in a search platform that provides several search terms.
[0038] Retrieve target user information that matches the problem from a user database containing pre-stored user information; the user information stored in the user database is obtained based on the information extraction method described in any one of the first aspects;
[0039] Based on the target user information, the target search term is retrieved from several search terms provided by the search platform, and an answer to the question is output based on the target search term.
[0040] In one implementation, the plurality of search terms include notes posted on the search platform by user accounts of the search platform; the notes include at least one of text, images, videos, or audio.
[0041] According to a third aspect of the embodiments of this application, a dialogue method is provided, including:
[0042] Obtain the questions entered by the user during their conversation with the chat assistant;
[0043] Retrieve target user information that matches the problem from a user database containing pre-stored user information; the user information stored in the user database is obtained based on the information extraction method described in any one of the first aspects;
[0044] Based on the question and the target user information, generate a prompt text;
[0045] The prompt text is input into a pre-trained dialogue model, which then generates an answer to the question based on the prompt text and the target user information; wherein the dialogue model is obtained by fine-tuning a pre-trained language model.
[0046] Output the answer.
[0047] According to a fourth aspect of the embodiments of this application, an electronic device is provided, comprising:
[0048] processor;
[0049] Memory used to store processor-executable instructions;
[0050] Wherein, when the processor executes the executable instructions, it is used to implement the method described in the first aspect, the second aspect, or the third aspect.
[0051] According to a fifth aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the methods described above.
[0052] According to a sixth aspect of the embodiments of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of any of the methods described above.
[0053] The technical solutions provided by the embodiments of this application may include the following beneficial effects:
[0054] In this embodiment, by acquiring the user's multi-turn dialogue content and generating a first prompt text and a second prompt text, the combined use of the two prompt texts enables the pre-trained model to not only extract the user's explicit information based on the first prompt text, but also to mine potential information based on the second prompt text. This dual processing mechanism improves the depth and breadth of understanding the user's intent and enhances the accuracy of information extraction.
[0055] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0056] Figure 1 This is a schematic diagram of a dialogue system provided in an exemplary embodiment.
[0057] Figure 2 This is a flowchart illustrating an information extraction method provided in an exemplary embodiment.
[0058] Figure 3 This is a schematic diagram of another schematic diagram of an extraction method provided in an exemplary embodiment.
[0059] Figure 4 This is a flowchart illustrating a dialogue method provided in an exemplary embodiment.
[0060] Figure 5 This is an exemplary embodiment of an interactive timing diagram of a dialogue process.
[0061] Figure 6 This is a flowchart illustrating another dialogue method provided in an exemplary embodiment.
[0062] Figure 7 This is an interactive timing diagram of another dialogue process provided in an exemplary embodiment.
[0063] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment. Detailed Implementation
[0064] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of this application as detailed in the appended claims.
[0065] It should be noted that the steps of the corresponding methods in other embodiments are not necessarily performed in the order shown and described in this application. In some other embodiments, the methods may include more or fewer steps than those described in this application. Furthermore, a single step described in this application may be broken down into multiple steps in other embodiments; and multiple steps described in this application may be combined into a single step in other embodiments.
[0066] With the widespread use of conversational assistants (such as intelligent voice assistants and chatbots), user interactions with these assistants are frequent and diverse. To provide more accurate services, one approach in related technologies is to extract user-related information based on the conversation between the assistant and the user. This helps the assistant better understand the user's needs, preferences, and personalized information in subsequent conversations. However, the content of user interactions with conversational assistants is often complex and diverse, resulting in relatively low accuracy for information extraction methods in these technologies.
[0067] Based on this, the embodiments of this application provide an information extraction method that combines direct extraction and indirect inference to obtain user information more comprehensively, thereby improving the intelligence and personalized service capabilities of the dialogue assistant.
[0068] In some embodiments, please refer to Figure 1 A dialogue system is provided, which includes a server 100 and several terminals 200.
[0069] Server 100 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. During operation, server 100 can run server-side programs for a specific application to implement its related functions. For example, when server 100 runs a dialog service program, it can act as the corresponding dialog server.
[0070] Terminal 200 includes, but is not limited to, smartphones, personal digital assistants, tablets, personal computers, laptops, wearable devices, virtual reality terminal devices, and augmented reality terminal devices. During operation, terminal 200 can run client-side programs for a specific application to implement its functions. For example, when terminal 200 runs a dialogue service program, it can act as a client for that dialogue service. The client application for this dialogue service can be launched and run on the electronic device. This client-side program can be an application installed on the electronic device, a webpage, a mini-program, a plugin, a component, or other similar forms.
[0071] The server 100 and several terminals 200 can interact via a network. The specific choice between wired or wireless networks for communication is not limited in this specification.
[0072] In some embodiments, please refer to Figure 2 and Figure 3 , Figure 2 A flowchart illustrating an information extraction method is shown. This information extraction method can be executed by an electronic device, which can be... Figure 1 The server 100 shown, or related to Figure 1 Other devices that the server 100 is connected to are not limited in this embodiment. The information extraction method includes:
[0073] In S201, the user's multi-turn conversation content is obtained.
[0074] In this process, the question-and-answer session between the user and the dialogue assistant constitutes one round of dialogue.
[0075] In this step, the electronic device acquires multi-turn dialogue content between the user and the dialogue assistant. This dialogue content is a detailed record of the interaction between the user and the dialogue assistant. Each round of dialogue includes a question and answer, recording the user's question and the dialogue assistant's response. This data forms the basis for subsequent information extraction and inference. By acquiring multi-turn dialogue content, the electronic device can capture the user's expressed needs, interests, and potential preferences during the dialogue.
[0076] In S202, based on the multi-turn dialogue content, the first prompt text and the second prompt text are generated.
[0077] In this step, after acquiring the content of multiple rounds of dialogue, the electronic device generates two different prompts based on this content. The first prompt aims to directly extract information explicitly expressed by the user from the multi-turn dialogue, such as containing explicit questions or instructions, to guide the pre-trained language model in direct information extraction. The second prompt, however, tends to infer information not explicitly expressed by the user, using suggestive or inferential questions to help the pre-trained language model indirectly infer the user's needs or preferences from the dialogue content.
[0078] By generating targeted prompt text, electronic devices can flexibly apply pre-trained language models to process different types of information. The first prompt text helps to accurately extract the user's directly expressed needs, while the second prompt text can uncover the user's implicit preferences and needs, thus providing more multi-dimensional information support for subsequent user interactions and personalized recommendations.
[0079] In S203, the first prompt text is input into the pre-trained language model so that the pre-trained language model can directly extract the target user information from the multi-turn dialogue content based on the first prompt text.
[0080] In this step, please refer to Figure 3 The electronic device inputs the initial prompt text into a pre-trained language model. Based on this prompt text, the pre-trained language model directly extracts the user's explicit information from the multi-turn dialogue. This explicit information typically includes the user's clearly expressed needs, preferences, or opinions. For example, a user might directly mention their liking for a particular product or their preference for a certain behavior; the pre-trained language model will directly capture this information through the initial prompt text. Through this direct extraction process, the electronic device can effectively accumulate the user's explicitly expressed interests and needs, and in subsequent services, assist the dialogue assistant in making personalized recommendations or responses based on these needs.
[0081] In S204, the second prompt text is input into the pre-trained language model so that the pre-trained language model can indirectly infer the target user information from the multi-turn dialogue content based on the second prompt text.
[0082] In this step, please refer to Figure 3 The electronic device inputs a second prompt text into a pre-trained language model. Based on this prompt text, the pre-trained language model indirectly infers user information from multiple rounds of dialogue. Unlike direct extraction, indirect inference relies more on the model's reasoning ability, aiming to extract potential information from content not explicitly expressed by the user. For example, a user may repeatedly mention terms or topics in a certain field without explicitly expressing interest; the system can identify the user's potential interests in that field through indirect inference.
[0083] Through this indirect extraction process, electronic devices can capture information that users do not explicitly express but is implied in the dialogue, expanding the depth of understanding of user needs and preferences. This allows the dialogue assistant to more comprehensively perceive the user's true intentions, thereby providing more personalized and intelligent responses and recommendations in subsequent interactions.
[0084] The information extraction method provided in this embodiment obtains the content of multiple rounds of dialogue between the user and the dialogue assistant, and generates a first prompt text and a second prompt text. This enables the pre-trained model to extract explicit information of the user based on the first prompt text and mine potential information based on the second prompt text, thereby improving the accuracy of information extraction.
[0085] The information extraction method provided in the embodiments of this application will be further described below:
[0086] In some embodiments, for S201, the electronic device can, in response to the dialogue content between the user and the dialogue assistant meeting preset conditions, acquire the multi-turn dialogue content between the user and the dialogue assistant, thereby automatically triggering the information extraction process. The preset conditions include, but are not limited to, any of the following: the number of turns in the dialogue content exceeds a set number of turns; the number of topics involved in the dialogue content exceeds a set number; the number of questions asked about the same topic in the dialogue content exceeds a set number; or the frequency of use of a specific keyword in the dialogue content exceeds a set frequency.
[0087] In one possible implementation, when a user engages in multiple rounds of interaction with the conversational assistant, it typically indicates that the user has a sustained interest in a particular topic or series of questions. This sustained interaction may reflect the user's genuine needs, questions, or new ideas formed during the conversation. Setting a threshold for the number of conversation rounds can help electronic devices identify such in-depth conversations and thus determine whether information extraction is necessary. This threshold can be specifically set according to the actual application scenario; this implementation does not impose any restrictions on it.
[0088] By triggering information extraction when the number of dialogue rounds between the user and the dialogue assistant exceeds a set threshold, electronic devices can capture user information based on sufficient content. This not only helps to understand the user's long-term needs, but also obtains more contextual information when the user expresses complex questions or preferences, thereby providing more accurate and personalized services. This approach can effectively avoid the problem of incomplete information extraction due to insufficient dialogue, while also improving the dialogue assistant's ability to understand the user's multi-turn dialogue.
[0089] In the second possible implementation, the increase in the number of topics in a user's multi-turn conversation usually means that the user is involved in multiple points of interest or needs in a short period of time. This diversified conversation may reflect the user's current complex needs or interest in multiple fields. Setting a threshold for the number of topics allows the system to determine whether information extraction is needed when the user discusses multiple topics. The threshold for the number of topics can be set according to the actual application scenario, and this implementation does not impose any restrictions on it.
[0090] When the number of topics discussed in a conversation between a user and a conversational assistant exceeds a set threshold, information extraction is triggered. This helps electronic devices capture the user's interests and needs in different areas. This not only enables the conversational assistant to understand the user's broad needs but also allows for more comprehensive suggestions and answers in subsequent interactions. When the user's interests are scattered, electronic devices can integrate information from multiple sources to provide the user with more relevant services. This helps identify the user's cross-domain interests and avoids overlooking potential needs or concerns.
[0091] In the third possible implementation, when a user repeatedly asks questions about the same topic, it usually indicates that the topic is very important to the user, and there may be unresolved issues or a strong interest in the topic. In this pattern of repeated questioning, the user may be hoping for more in-depth answers or there may be hidden information that has not been captured. Therefore, setting a threshold for the number of questions asked about the same topic can help electronic devices identify topics that users are highly interested in. The threshold for the number of questions asked about the same topic can be set according to the actual application scenario, and this implementation does not impose any restrictions on it.
[0092] When a user asks the same question more times on the same topic than a set threshold, information extraction is triggered. This allows the electronic device to focus more on the user's core needs or questions, ensuring that the user receives sufficient support and feedback on topics of interest. The electronic device can perform in-depth analysis of the user's frequently asked questions, extracting more targeted information. This extracted information can assist the dialogue assistant, preventing the assistant from providing incomplete or superficial answers when the user continues to focus on a particular topic.
[0093] In the fourth possible implementation, the repeated use of specific keywords by the user in the conversation usually indicates that these keywords are closely related to the user's current needs, interests, or emotional state. High-frequency use of keywords may reflect the user's focus or emotional inclination. Setting a frequency threshold for specific keywords allows electronic devices to trigger information extraction when users frequently use certain keywords. The frequency threshold for specific keywords can be set according to the actual application scenario; this implementation does not impose any restrictions on it.
[0094] When the frequency of certain keywords in a conversation between a user and the dialogue assistant exceeds a set threshold, information extraction is triggered. This allows the electronic device to capture the user's core interests or emotional expressions within a specific time period, improving the accuracy of information extraction. The extracted information better reflects the user's key needs and emotional state, which not only helps the dialogue assistant provide more accurate answers but also helps track and reinforce the user's focus in subsequent conversations. By identifying high-frequency keywords, dialogue strategies can be quickly adjusted to provide more personalized services when user preferences or needs change.
[0095] In some embodiments, the first prompt text generated by the electronic device based on the multi-turn dialogue content may include a first processing requirement for the multi-turn dialogue content and the multi-turn dialogue content itself; wherein, the first processing requirement is at least used to instruct the pre-trained language model to extract directly stated target user information from the multi-turn dialogue content. With an explicit first processing requirement, the pre-trained language model can focus on identifying and extracting information directly expressed by the user in the dialogue, including preferences, needs, or important information explicitly mentioned by the user, such as "I like to travel" or "I want to buy a new phone." This direct extraction ensures that explicitly expressed user interests and needs are not overlooked.
[0096] The second prompt text generated by the electronic device based on multi-turn dialogue content may include a second processing requirement for the multi-turn dialogue content and the multi-turn dialogue content itself. The second processing requirement at least instructs the pre-trained language model to identify target topics that appear more frequently than a set threshold in the multi-turn dialogue content, infer target user information based on the target topics, and provide reasons for the inference. The design of the second prompt text aims to uncover the user's implicit needs or interests from the potential patterns in the dialogue content. For example, if a user frequently mentions a topic in multiple dialogues but does not explicitly express interest, the pre-trained language model can infer that the topic is a possible area of interest for the user. By analyzing these frequently occurring topics, it can infer the user's possible hobbies, concerns, or future needs, and these inferences will be accompanied by reasons, thereby increasing the credibility and interpretability of the inferences.
[0097] The combined use of these two prompts allows electronic devices, with the help of a pre-trained language model, to not only accurately extract information explicitly expressed by the user but also infer the user's implicit needs or interests. This dual processing mechanism significantly improves the depth and breadth of understanding user intent, contributing to more personalized and precise services in subsequent interactions. Simultaneously, by recording the reasons for the inferences, the electronic device can verify and optimize the results, further enhancing the reliability of information processing.
[0098] For example, the first prompt text may also include at least one of the following: role setting information, the type of user information to be extracted, and at least one processing example that meets the first processing requirements.
[0099] The second prompt text also includes at least one of the following: role setting information, the type of user information to be extracted, and at least one processing example that meets the second processing requirements.
[0100] Role-setting information specifies the role or perspective that the pre-trained language model should play when processing dialogue content. This helps the pre-trained language model better understand the user's needs and context, thus making it more accurate in extracting and inferring user information. For example, role-setting might allow the model to simulate a consultant, customer service representative, or personal assistant, so that the pre-trained language model will process information according to the characteristics of different roles, making it closer to real-world application scenarios.
[0101] For example, in the initial prompt text, when the pre-trained language model knows the role it needs to play, it can more accurately extract explicit user information related to that role. For instance, if the role is set as "health consultant," the pre-trained language model may focus more on the user's directly expressed health preferences or needs.
[0102] For example, in the second prompt text, when inferring implicit information, role-playing can help the model understand the dialogue from a specific perspective. For instance, as a "travel planner," the model might pay more attention to travel destinations mentioned repeatedly by the user, inferring the user's implicit interest in travel.
[0103] Specifically, the type of user information to be extracted clearly informs the pre-trained language model of the information type it should extract or infer. This ensures that the pre-trained language model can focus on specific information categories, such as interests, behavioral habits, and purchasing intentions, when processing dialogue content. This helps improve the relevance and effectiveness of the pre-trained language model's processing and avoids generalization or bias in information extraction.
[0104] For example, explicitly specifying the type of user information to be extracted in the initial prompt text allows the pre-trained language model to focus on extracting explicit information, such as "extracting the user's interests and hobbies" or "extracting the user's occupation information." In this way, the pre-trained language model can accurately capture information that matches these types when parsing the dialogue.
[0105] For example, in the second prompt text, the explicit information type is just as important as the implicit information inference. For instance, when specifying the extraction of the user's "potential consumption intention," the pre-trained language model tends to look for purchase-related content in frequently used topics for inference.
[0106] The processing examples provide concrete operational samples demonstrating how a pre-trained language model should handle similar tasks. These examples help the pre-trained language model understand task requirements, reduce comprehension bias, and improve execution consistency. Particularly for complex extraction or inference tasks, the processing examples effectively guide the behavior of the pre-trained language model.
[0107] For example, by providing examples in the initial prompt text, the pre-trained language model can understand what information meets the criteria for "direct extraction." For instance, an example might be "the user explicitly stated that they like ebooks," and the model would understand that similar direct expressions need to be extracted.
[0108] For example, in the second prompt text, when inferring implicit information, examples can help the pre-trained language model understand what kind of inference is reasonable and how to make reasonable inferences based on dialogue frequency or content features. For instance, an example might demonstrate the process of "the user mentions environmental topics multiple times, therefore inferring that the user is interested in sustainable products," guiding the pre-trained language model in handling similar situations.
[0109] For example, the first and second prompt texts can contain all of the above. By including role setting information, the type of user information to be extracted, and processing examples in the prompt texts, the pre-trained language model can be guided to perform complex tasks more effectively. This content enables the pre-trained language model to more accurately understand the task objective, make reasonable inferences in specific contexts, and ultimately improve the accuracy and practicality of information extraction and inference. This not only optimizes the processing capabilities of the pre-trained language model but also ensures that the parsing of dialogue content is closer to user needs and real-world application scenarios.
[0110] For example, the initial prompt text can be represented in the following form:
[0111] Imagine you are a seasoned summarizing expert. Here are some questions users have asked a chatbot. Extract the following information from these questions:
[0112] Basic information: Name, Gender, Age, Occupation, Birthday, Zodiac Sign, Personality, Relationship Status (Single / Married / With Children), Location
[0113] Lifestyle: hobbies, lifestyle habits, eating habits, travel habits, spending habits, daily routine (whether you stay up late, wake up early, or take a nap), work situation (occupation, work location, job content), entertainment (favorite types of books, movies, games, TV shows, whether you like going to exhibitions, whether you like outdoor activities, etc.).
[0114] Major life events: life experiences (marriage, buying a house, renovation, studying abroad, having a baby, various exams, birth, aging, illness, and death, etc.), and emotional life (arguments, dating, breakups, etc.).
[0115] Note that the summary result must be something the user directly states. For example, the user should directly state in the question what they like, what content they enjoy, or what their dietary preferences are, such as "I like a certain singer," "I like a certain game," "I love spicy food," or "I work at an internet company." It *must* be something the user directly expresses, and *not* a possible result inferred from the user's question. Ignore irrelevant information. The result must be given with certainty. If the above information cannot be summarized, please output "None."
[0116]
Example
[0117] Users are asking questions such as: "Are there any campsites near location A?", "Do you have any pictures? Or recommendations from platform B?", "What are some good songs?", "I like singer B, do they have any song recommendations?", "Is this model a platform B's own model? Does it use any other interfaces?", "What features makes this model different from other models?", "What are the most popular tags right now?", "I'm a blogger on platform B, what kind of content is currently popular with friends from other regions?", "What are the most popular tourist attractions in location A?", "What World Heritage sites are there in location A?", "I want to recommend a certain food in location A, it's a unique dish, can you help me write a recommendation?", "The above text has too few emojis, can you add some?", "This restaurant is located in location A, the average cost per person is 100, please add the above information and rewrite the text.", "C University travel guide." ','C University Tourism Power,' 'How much does a certain food cost?', 'A certain brand's new food,' 'How to allocate skill points for a D game character,' 'How to allocate skill points for a D game character,' 'How to match a certain artifact,' 'The matching of a certain artifact,' 'Give the matching of a certain artifact,' 'Are there any good places to eat meat in A, with large portions and good taste?', 'How much does the live-fire shooting range in A cost?', 'How much does the live-fire shooting range in A cost?'
[0118] Summary of results:
[0119]
Basic Information
[0120] [Lifestyle] Regarding entertainment, the user explicitly stated that they like singer B. Regarding work, the user explicitly stated that they are a blogger on platform B.
[0121] [Major Events] None
[0122] Below is a multi-turn dialogue between users in a real-world scenario. Please summarize it following the content and format of the example above: ******.
[0123] For example, the second prompt text can be represented in the following form:
[0124] Imagine you are a seasoned expert at summarizing information. Here are some questions users ask a chatbot. Infer some information about the users from these questions, as follows:
[0125] Basic information: Name, Gender, Age, Occupation, Birthday, Zodiac Sign, Personality, Relationship Status (Single / Married / With Children), Location
[0126] Lifestyle: hobbies, lifestyle habits, eating habits, travel habits, spending habits, daily routine (whether you stay up late, wake up early, or take a nap), work situation (occupation, work location, job content), entertainment (favorite types of books, movies, games, TV shows, whether you like going to exhibitions, whether you like outdoor activities, etc.).
[0127] Major life events: life experiences (marriage, buying a house, renovation, studying abroad, having a baby, various exams, birth, aging, illness, and death, etc.), and emotional life (arguments, dating, breakups, etc.).
[0128] Note that you need to provide the inferred result and reasoning. This must be based on the user asking the same question multiple times; questions asked only once are ignored. Below are two examples:
[0129]
Example
[0130] The questions users asked were: ['Are there any campsites near location A?', 'Are there any restaurants suitable for couples?', 'What are some dining recommendations for location A?', 'Which attractions in location A are good for taking photos?', 'What fun indoor activities can you recommend?', 'I want to do some weight loss exercises, which ones should I start with?', 'I want to go on a beach vacation, what beach resorts in location A would you recommend?', 'I feel like I'm not in a good state lately, how should I adjust?', 'How can I overcome procrastination?', 'Are there any suggestions on how to improve emotional intelligence?', 'I always feel stressed at work, how can I relieve stress?', 'What are some tourist attractions in location A?', 'What types of movies do I like to watch?', 'My favorite books are history books, and I recently read a book about science and technology.']
[0131] The conclusion is:
[0132]
Basic Information
[0133] Living conditions
[0134] Interests and hobbies: (1) The user may enjoy outdoor activities, as they repeatedly asked about campsites and photo spots in location A. (2) The user may enjoy traveling, as they asked for suggestions on restaurants, tourism, and seaside resorts in location A. (3) The user may be interested in sports, especially weight loss exercises, as they asked about suitable sports. (4) The user may enjoy movies and history books, as they mentioned that they enjoy reading history books and recently read a book about science and technology.
[0135] Daily routine: The user may have procrastination issues, as the user repeatedly asked questions about overcoming procrastination.
[0136] Work situation: Users may be facing significant work pressure, as they asked for advice on how to reduce stress and improve emotional intelligence, indicating a focus on personal psychological adjustment.
[0137] [Important Event] The user may have experienced some kind of emotional distress or poor condition, as the user mentioned that they have not been feeling well recently and asked how to adjust.
[0138] Below is the content of the user's multi-turn dialogue. Please provide your summary: ******.
[0139] In some embodiments, in order to improve the efficiency of generating prompt text, a first prompt template and a second prompt template can be preset, so that after the electronic device acquires the multi-turn dialogue content, it can embed the multi-turn dialogue content into the preset first prompt template to generate a first prompt text; and embed the multi-turn dialogue content into the preset second prompt template to generate a second prompt text.
[0140] For example, the first prompt template includes at least one of the following: a first processing requirement, role setting information, the type of user information to be extracted, and at least one processing example that satisfies the first processing requirement. The second prompt text includes at least one of the following: a second processing requirement, role setting information, the type of user information to be extracted, and at least one processing example that satisfies the second processing requirement.
[0141] One possible implementation involves setting a first prompt template and a second prompt template. This is suitable for scenarios where the nature of the dialogue content and user needs are relatively fixed, such as customer service dialogues in specific fields or dialogue assistants with fixed functions. In these cases, the structure of the dialogue content and the required prompt text format are relatively consistent. Since there are only two fixed templates, the design and maintenance are relatively simple and efficient. Furthermore, using fixed templates ensures that the generated prompt text has a consistent format and style, which helps maintain output stability.
[0142] In another possible implementation, considering that the expression style of dialogue content may vary in certain scenarios, a single template may not meet all needs. Therefore, a first prompt template library and a second prompt template library can be set up. The first prompt template library pre-stores first prompt templates corresponding to different expression styles; the second prompt template library pre-stores second prompt templates corresponding to different expression styles. This makes it suitable for scenarios that need to handle multiple different dialogue styles or different user needs, such as personalized dialogue assistants, complex customer relationship management systems, or applications that need to support multiple languages and styles.
[0143] After the electronic device acquires the aforementioned multi-turn dialogue content, it can analyze the content to determine the target expression style. Then, based on this style, it retrieves a first prompt template from a first prompt template library that matches the target style, and a second prompt template from a second prompt template library that also matches the target style. In this embodiment, the template library provides multiple expression styles, allowing the electronic device to select the most suitable template based on actual needs. This method generates prompt text that better matches user requirements and the style of the dialogue content. By selecting a suitable template, the electronic device can generate more personalized and targeted prompt text, improving the naturalness of the dialogue and the user experience.
[0144] For example, the first prompt template library and the second prompt template library pre-store templates with various expression styles. For example, these templates can be classified according to different expression styles (such as formal, informal, humorous, serious, etc.). Each template contains a fixed structure with space reserved for embedding multi-turn dialogue content. After acquiring the multi-turn dialogue content, the electronic device preprocesses the multi-turn dialogue content. The preprocessing includes, but is not limited to: (1) text segmentation, which divides the dialogue content into words or phrases; (2) grammatical parsing, which analyzes the grammatical structure of sentences and identifies components such as subject, predicate, and object; (3) sentiment analysis, which identifies the sentiment tendency in the dialogue (such as positive, neutral, negative); (4) keyword extraction, which extracts keywords or themes that appear frequently in the dialogue. Then, based on the results of the preprocessing, the style features of the dialogue content are extracted. The style features include, but are not limited to: (1) sentiment tendency, whether the dialogue content expresses strong emotions, such as anger, joy, anxiety, etc.; (2) tone, whether the tone of the dialogue content is formal, casual, or humorous; (3) vocabulary, whether professional terms, slang, or expressions with specific cultural backgrounds are used in the dialogue. Next, the electronic device matches the extracted style features with template style tags in a template library to determine the target expression style. For example, if the dialogue content displays a strong emotional tendency and uses an informal tone, the electronic device might select an "informal" or "humorous" style prompt template. Based on the determined target expression style, the electronic device selects a template matching that style from a first prompt template library to generate the first prompt text; similarly, it selects a matching template from a second prompt template library to generate the second prompt text; finally, the electronic device embeds the multi-turn dialogue content into the selected template to generate the complete prompt text. Through this implementation, the electronic device can generate more personalized prompt text that matches the user's expression style, thereby improving the accuracy of information extraction and user experience satisfaction.
[0145] In some embodiments, after acquiring the first prompt text and the second prompt text, the electronic device can input the first prompt text into a pre-trained language model so that the pre-trained language model can directly extract target user information from the multi-turn dialogue content based on the first prompt text; and input the second prompt text into the pre-trained language model so that the pre-trained language model can indirectly infer target user information from the multi-turn dialogue content based on the second prompt text.
[0146] Pre-trained language models refer to deep learning models that, after being trained on a large amount of text data, can understand and generate natural language. These models, once pre-trained, can be applied to various downstream tasks, such as text classification, translation, information extraction, and dialogue generation. By employing pre-trained language models from related technologies and designing different prompt texts, electronic devices can leverage the understanding and generation capabilities of pre-trained language models to extract explicitly expressed information and infer implicit interests or information from the user, thereby constructing a more comprehensive user profile.
[0147] Of course, the pre-trained language model can also be fine-tuned to better suit the information extraction task of this application. For example, after pre-training, the pre-trained language model can be fine-tuned using labeled sample data to make the fine-tuned pre-trained language model more suitable for the information extraction task.
[0148] In some embodiments, after obtaining the target user information output by the pre-trained model, the target user information can be stored in a user database to assist in the subsequent dialogue process.
[0149] For example, to avoid storage redundancy, after obtaining the target user information output by the pre-trained model, the electronic device can detect whether the user database stores historical user information of the same type as the target user information; if the user database does not store historical user information of the same type as the target user information, the target user information is stored in the user database.
[0150] Specifically, electronic devices first identify the type of target user information (such as user interests, preferences, habits, etc.). Then, they search the user database for historical user information that matches that target user information type. For example, if the target user information type is "interests," the electronic device will search the user database for existing interest records to see if there is any content similar to or duplicated with the newly extracted interest. Only if no similar information exists in the user database will the newly extracted target user information be stored, thus avoiding duplicate data storage.
[0151] For example, if the user database stores historical user information of the same type as the target user information, the electronic device can retrieve the historical user information from the user database, then determine the semantic similarity between the target user information and the historical user information, and perform the following steps based on the semantic similarity:
[0152] In the first possible scenario, if the semantic similarity is greater than the first similarity threshold, it means that the target user information output by the pre-trained model is highly similar to or almost identical to the historical user information already existing in the user database. Since the two are highly similar, storing new information would lead to data redundancy, wasting storage space and not increasing the effective information volume of the database. Therefore, electronic devices can discard the target user information output by the pre-trained model instead of storing it in the user database, effectively avoiding duplicate storage, optimizing the storage efficiency of the database, and reducing the trouble that redundant data may cause in subsequent data processing and analysis, thereby improving response speed and efficiency.
[0153] In the second possible scenario, if the semantic similarity is less than the second similarity threshold, it means that the target user information output by the pre-trained model differs significantly from the historical user information in the user database, and may even be completely different information. In other words, the old historical user information may be outdated or irrelevant, while the new target user information better reflects the current user's true situation and interests. Therefore, the database needs to be updated. The electronic device can delete the historical user information and store the target user information in the user database. This embodiment, by deleting outdated or irrelevant historical information and storing new information, can maintain the up-to-dateness and accuracy of the database content, ensuring more accurate services in scenarios such as user profiling and personalized recommendations. Moreover, it can also avoid the disorderly accumulation of data and keep the database clean and efficient.
[0154] In the third possible scenario, if the semantic similarity falls between the first and second similarity thresholds, meaning the target user information and historical user information have some similarity but are not completely identical, the electronic device can merge the target user information and historical user information and store the merged user information in the user database. For example, the electronic device can employ weighted averaging, information splicing, or other merging strategies to integrate the useful parts of the two pieces of information. For instance, if the historical information is "the user likes watching science fiction movies," and the target information is "the user has recently been frequently watching horror movies," the electronic device can merge these two into "the user likes watching both science fiction and horror movies." This merging method comprehensively considers the user's latest interests and historical preferences, forming a more comprehensive and accurate user profile, reducing data redundancy, and ensuring the effective integration of new and old information, thereby improving the quality and usability of the database.
[0155] The first similarity threshold is greater than the second similarity threshold. The first and second similarity thresholds can be set according to the actual application scenario, and this implementation does not impose any restrictions on them.
[0156] In some embodiments, please refer to Figure 4This application also provides a dialogue method, which can be applied to an electronic device, the electronic device being... Figure 1 The server 100 shown, or related to Figure 1 Other devices that the server 100 can communicate with are not limited in this embodiment. The dialogue method includes:
[0157] In S401, the questions entered by the user during the conversation with the dialogue assistant are obtained; the dialogue assistant is deployed in the search platform, which provides several search items.
[0158] In this step, the conversational assistant is deployed on a search platform that provides several search options.
[0159] For example, several search terms include notes posted by user accounts on the search platform; notes can include at least one of text, images, video, or audio. For instance, users can write and publish plain text content, which can be articles, comments, logs, summaries, or any other form of written expression. Alternatively, users can upload and share images, such as photos, illustrations, screenshots, etc., typically used to supplement or enhance text content. Or, users can upload and publish video content, which can include recorded videos, instructional videos, demonstrations, etc. Alternatively, users can publish audio files, such as recordings, podcasts, or lectures, allowing users to convey information or share opinions through sound. A single note can include any one or more forms of media content to provide richer information. For example, a note may simultaneously contain text descriptions, related images, and additional video or audio to comprehensively express the user's thoughts or information. Users on the search platform can use these notes to share personal insights, record important information, or engage in social interaction.
[0160] Electronic devices first acquire the query questions entered by the user when interacting with the conversational assistant. These questions typically express the user's need for specific information or content and can be natural language input, such as "I want to know the latest technology news" or "Recommend some family-friendly vacation destinations." By acquiring the user's input questions, electronic devices can identify the user's current needs and thus initiate the corresponding search and recommendation process.
[0161] In S402, target user information that matches the problem is retrieved from the user database containing pre-stored user information; the user information stored in the user database is obtained based on the information extraction method described above.
[0162] In this step, after obtaining the user's input question, the electronic device accesses the user database to retrieve target user information related to that question. The user information stored in the database is obtained through the previously described information extraction methods and includes user preferences, historical search records, and frequently used keywords. For example, if the user database records a user's preference for technology news, and the current user's question is "I want to know about the latest technology news," the electronic device will retrieve this target user information to match the user's current query.
[0163] By recalling target user information that matches the problem, electronic devices can more accurately pinpoint users' interests and preferences. This personalized information retrieval method ensures that electronic devices can provide search results that better meet user needs, greatly improving the user experience.
[0164] In S403, the target search term is retrieved from several search terms provided by the search platform based on the target user information, and an answer to the question is output based on the target search term.
[0165] In this step, after retrieving the target user information, the electronic device uses this information to retrieve the most relevant search term from multiple search options provided by the search platform. For example, if a user frequently follows technology news and their input question is about recent technology news, the system will prioritize searching for "technology news." This process not only relies on the user's current question but also comprehensively considers target user information to select the search term that best matches the user's needs. By retrieving target search terms based on user information, the electronic device can better match search results with the user's personalized needs. This approach improves search relevance, ensures the quality of search results, and ultimately increases user satisfaction and platform user stickiness.
[0166] This embodiment integrates user information and real-time queries, making search results more personalized and accurate, and improving search efficiency.
[0167] In an exemplary dialogue scenario, please refer to Figure 5 , Figure 5 An interactive sequence diagram of a dialogue process is shown.
[0168] In S501, terminal 200 responds to a question entered by the user in a dialog box by generating a dialog request carrying that question.
[0169] For example, the user input question can be in the form of text, voice, image or other forms, and this implementation does not impose any restrictions on it.
[0170] In S502, terminal 200 sends a dialogue request to server 100.
[0171] In S503, server 100 receives a dialogue request, retrieves target user information that matches the question in the dialogue request from a user database containing pre-stored user information, retrieves the target search term from several search terms provided by the search platform based on the target user information, and generates an answer to the question based on the target search term.
[0172] In S504, server 100 sends a dialogue response containing the answer to terminal 200.
[0173] In S505, terminal 200 displays the answer in the interactive interface.
[0174] In some embodiments, please refer to Figure 6 This application also provides another dialogue method, which can be applied to an electronic device, such as... Figure 1 The server 100 shown, or related to Figure 1 Other devices that the server 100 can communicate with are not limited in this embodiment. The dialogue method includes:
[0175] In S601, the questions entered by the user during the conversation with the dialogue assistant are obtained.
[0176] In this step, the electronic device first obtains the query question entered by the user when interacting with the conversational assistant. This question is usually an expression of the user's need for specific information or content, and can be natural language input, such as "I want to know the latest technology news" or "Recommend some family-friendly vacation destinations". By obtaining the user's input question, the electronic device can identify the user's current needs and thus initiate the corresponding search and recommendation process.
[0177] In S602, target user information that matches the problem is retrieved from the user database containing pre-stored user information; the user information stored in the user database is obtained based on the information extraction method described above.
[0178] In this step, after obtaining the user's input question, the electronic device accesses the user database to retrieve target user information related to that question. This information, stored in the user database, is obtained through the previously described information extraction methods and includes user preferences, historical search records, and frequently used keywords. For example, if the user database records a user's preference for technology news, and the current user's question is "I want to know about the latest technology news," the electronic device will retrieve this target user information to match the user's current query.
[0179] By recalling target user information that matches the problem, electronic devices can more accurately pinpoint users' interests and preferences. This personalized information retrieval method ensures that electronic devices can provide search results that better meet user needs, greatly improving the user experience.
[0180] In S603, a prompt text is generated based on the question and the target user information.
[0181] In this step, after obtaining the question and relevant user information, the electronic device generates a prompt text to guide the dialogue model in generating an answer. The prompt text typically includes two main parts: instructions for handling the user's question and referenced user information. For example, if the user's question is "What will the weather be like next week?", the prompt text generated by the electronic device might include, "Please provide a detailed weather forecast based on the user's historical location data and date range."
[0182] By generating precise prompt text, electronic devices can effectively guide dialogue models to produce responses that better meet user needs. This step plays a crucial role in improving the accuracy and relevance of responses, ensuring that the content generated by the dialogue model directly addresses the user's needs.
[0183] In S604, the prompt text is input into the trained dialogue model so that the dialogue model can generate an answer to the question according to the instructions of the prompt text and with reference to the target user information; wherein, the dialogue model is obtained by fine-tuning a pre-trained language model.
[0184] In this step, the generated prompt text is input into the pre-trained dialogue model. This dialogue model, fine-tuned from a pre-trained language model, possesses strong contextual understanding and generation capabilities. Following the instructions in the prompt text and incorporating target user information, the dialogue model generates a specific answer to the user's question. For example, if a user asks, "Recommend some science fiction novels?", the model will refer to the user's reading history and generate, "I recommend you read *The Three-Body Problem* and *Dune*."
[0185] Using a finely tuned dialogue model, we can better understand the prompt text and generate high-quality responses. This step ensures that the generated content is both highly relevant to the user's question and incorporates personalized information, making the responses more targeted and practical.
[0186] In S606, the answer is output.
[0187] Finally, the electronic device outputs the answer generated by the dialogue model to the user as a direct response to the user's question. The output can be presented in the form of text, voice, etc., depending on how the user interacts with the dialogue assistant.
[0188] This embodiment leverages the user's historical information and the powerful generative capabilities of a pre-trained language model to achieve efficient and personalized dialogue generation. Users can obtain responses that match their interests and needs, thereby improving the intelligence level of the dialogue assistant and the user experience.
[0189] In an exemplary dialogue scenario, please refer to Figure 7 , Figure 7 An interactive sequence diagram of a dialogue process is shown.
[0190] In S701, terminal 200 responds to a question entered by the user in a dialog box by generating a dialog request carrying that question.
[0191] For example, the user input question can be in the form of text, voice, image or other forms, and this implementation does not impose any restrictions on it.
[0192] In S702, terminal 200 sends a dialogue request to server 100.
[0193] In S703, server 100 receives a dialogue request, retrieves target user information that matches the question in the dialogue request from a user database containing pre-stored user information, generates prompt text based on the question and the target user information, and inputs the prompt text into a trained dialogue model so that the dialogue model can generate an answer to the question according to the instructions of the prompt text and with reference to the target user information.
[0194] In S704, server 100 sends a dialogue response containing the answer to terminal 200.
[0195] In S707, terminal 200 displays the answer in the interactive interface.
[0196] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.
[0197] In some embodiments, this application also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor implements the method described in any one of the above embodiments by executing the executable instructions.
[0198] For example, Figure 8 This is a schematic structural diagram of a device provided in an exemplary embodiment. Please refer to... Figure 8At the hardware level, the device includes a processor 802, an internal bus 804, a network interface 806, memory 808, and non-volatile memory 810, and may also include other hardware required for different scenarios. One or more embodiments of this application can be implemented in software, such as the processor 802 reading the corresponding computer program from the non-volatile memory 810 into memory 808 and then running it. Of course, in addition to software implementation, one or more embodiments of this application do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0199] In some embodiments, this application also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the method as described in any of the preceding claims.
[0200] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0201] In some embodiments, this application also provides a computer program product including a computer program that, when executed by a processor, implements the steps of the method as described in any of the preceding embodiments.
[0202] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0203] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0204] The terminology used in one or more embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this application. The singular forms “a,” “the,” and “the” used in one or more embodiments of this application and in the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0205] The above description is merely a preferred embodiment of one or more embodiments of this application and is not intended to limit the scope of one or more embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this application should be included within the protection scope of one or more embodiments of this application.
Claims
1. An information extraction method, characterized in that: include: Get the user's multi-round conversation content; Based on the contents of the multiple rounds of conversations, generating a first prompt text and a second prompt text; Inputting the first prompt text into a pre-trained language model, so that the pre-trained language model directly extracts target user information from the multi-round conversation content based on the first prompt text; as well as, The second prompt text is input into the pre-trained language model so that the pre-trained language model can indirectly infer the target user information from the multiple rounds of conversation content based on the second prompt text.
2. The method according to claim 1, characterized in that The first prompt text includes a first processing requirement for the multi-round dialogue content and the multi-round dialogue content; wherein the first processing requirement is at least used to instruct the pre-trained language model to extract directly stated target user information from the multi-round dialogue content; and / or, The second prompt text includes a second processing requirement for the multi-round conversation content and the multi-round conversation content; wherein the second processing requirement is at least used to instruct the pre-trained language model to determine a target topic whose occurrence frequency is greater than a set threshold from the multi-round conversation content, infer the target user information based on the target topic, and provide reasons for the inference.
3. The method according to claim 2, characterized in that The first prompt text further includes at least one of the following contents: role setting information, the type of user information to be extracted, and at least one processing example that meets the first processing requirement; and / or, The second prompt text also includes at least one of the following contents: role setting information, the type of user information to be extracted, and at least one processing example that meets the second processing requirement.
4. The method according to claim 1, characterized in that: The obtaining of multiple rounds of conversation content between the user and the conversation assistant includes: In response to the conversation content between the user and the conversation assistant meeting a preset condition, obtaining multiple rounds of conversation content between the user and the conversation assistant; Among them, the preset conditions include any one of the following: the number of rounds of the conversation content is greater than the set number of rounds, the number of topics involved in the conversation content exceeds the set number, the number of questions asked on the same topic in the conversation content exceeds the set number, or the frequency of use of specific keywords in the conversation content exceeds the set frequency.
5. The method according to claim 1, characterized in that The generating a first prompt text and a second prompt text based on the multi-round conversation content includes: Embedding the multi-round conversation content into a preset first prompt template to generate the first prompt text; and The multi-round conversation contents are embedded into a preset second prompt template to generate the second prompt text.
6. The method according to claim 5, characterized in that The first prompt template and the second prompt template are obtained in the following manner: Analyzing the contents of the multiple rounds of conversations to determine a target expression style; Based on the target expression style, acquiring a first prompt template adapted to the target expression style from a first prompt template library, and acquiring a second prompt template adapted to the target expression style from a second prompt template library; The first prompt template library pre-stores first prompt templates corresponding to different expression styles; the second prompt template library pre-stores second prompt templates corresponding to different expression styles.
7. The method according to claim 1, characterized in that Also includes: After obtaining the target user information output by the pre-training model, detecting whether historical user information of the same type as the target user information is stored in a user database; If the user database does not store historical user information of the same type as the target user information, the target user information is stored in the user database.
8. The method according to claim 7, characterized in that Also includes: If the user database stores historical user information of the same type as the target user information, obtaining the historical user information from the user database; Determining the semantic similarity between the target user information and the historical user information; If the semantic similarity is greater than a first similarity threshold, discarding the target user information output by the pre-training model; If the semantic similarity is less than a second similarity threshold, deleting the historical user information, and storing the target user information in the user database; If the semantic similarity is between the first similarity threshold and the second similarity threshold, merging the target user information and the historical user information, and storing the merged user information in the user database; The first similarity threshold is greater than the second similarity threshold.
9. A conversation method, characterized in that: include: Obtaining a question input by a user during a conversation with a conversation assistant; wherein the conversation assistant is deployed in a search platform, and the search platform provides a number of search items; Recalling target user information adapted to the question from a user database pre-stored with user information; the user information stored in the user database is obtained based on the information extraction method according to any one of claims 1 to 8; A target search term is retrieved from a number of search terms provided by a search platform based on the target user information, so as to output an answer to the question based on the target search term.
10. The method according to claim 9, characterized in that The plurality of search items include notes posted on the search platform by a user account of the search platform; The note includes at least one of text, image, video or audio.
11. A conversation method, characterized in that: include: Get the questions entered by the user during the conversation with the conversation assistant; Recalling target user information adapted to the question from a user database pre-stored with user information; The user information stored in the user database is obtained based on the information extraction method according to any one of claims 1 to 8; Generate a prompt text based on the question and the target user information; Inputting the prompt text into a trained dialogue model, so that the dialogue model generates an answer to the question according to the instruction of the prompt text and with reference to the target user information; wherein the dialogue model is obtained by fine-tuning a pre-trained language model; The answer is output.
12. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 11 by running the executable instructions.
13. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
Citation Information
Cited By
Sample generation method, model training method, data processing method and electronic equipment
CN121681783A