Voice interaction method and device, household appliance and computer readable storage medium

By pre-constructing a Q&A text library in the voice interaction system, and using interactive keywords and user identity identification numbers to query related question texts, the problem of low generation rate of recommended topics in voice interaction is solved, and the user interaction experience and device intelligence are improved.

CN119988526APending Publication Date: 2025-05-13QINDAO HAIER REFRIGERATOR CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311481462.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-05-13

Smart Images

  • Figure CN119988526A_ABST
    Figure CN119988526A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent household appliances, and discloses a voice interaction method and device, a household appliance and a computer readable storage medium. The voice interaction method is applied to the household appliance, and the voice interaction method comprises the following steps: analyzing first-round dialogue interaction information, and obtaining an interaction keyword; querying a related question text in a question and answer pair text library according to the interaction keyword and the user identity identification number; performing correlation sorting on the related problem texts to obtain a related problem text sequence; and generating a recommended topic text according to the related question text sequence and outputting the recommended topic text to the user. Compared with the prior art, the method has the advantages that selection of the subtopic from the knowledge graph and generation of the topic recommendation statement according to the subtopic are avoided, the generation rate of the recommendation topic is increased, and the user interaction experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of smart home appliances, for example, to a voice interaction method and device, home appliance equipment, and a computer-readable storage medium. Background Art

[0002] By installing artificial intelligence technology on home appliances, researchers have enabled home appliances to communicate with users through voice and achieve voice interaction. Existing voice interaction functions are mostly based on the user's historical voice interaction information, integrating the user's historical questions and determining the user's preferences. When the user asks about similar topics, the system actively asks the user's needs to achieve active question and answer. However, active question and answer based on the user's historical voice interaction information often repeatedly asks the user the same historical questions, resulting in low interest in the interaction process and low topic richness.

[0003] In order to improve the fun of the human-computer interaction process and the richness of topics, and improve the user's interactive experience, the relevant technology discloses a method and system for processing dialogue interaction. The method includes the following steps: parsing multiple rounds of dialogue interaction information to obtain entity information related to the topic; determining the main topic and subtopic to which each round of dialogue interaction belongs based on the entity information and the topic information obtained in the context dialogue; after the set round of dialogue interaction is completed, selecting other subtopics associated with the main topic to which the current round belongs, generating a topic recommendation statement based on the set topic recommendation method and outputting it to the user; parsing the multimodal feedback data of the user on the topic recommendation statement, and determining whether to continue the recommended subtopic based on the user feedback result.

[0004] In the process of implementing the embodiments of the present disclosure, it is found that there are at least the following problems in the related art:

[0005] Related technologies identify topics from the current dialogue interaction, and actively interact with sub-topics in the framework of the topic graph of main topics and sub-topics, thereby improving the richness of topics in multiple rounds of dialogue and enhancing the fun of human-computer dialogue. However, the amount of information in the knowledge graph is huge, and how to quickly generate recommended topics in real time during the interaction process is an urgent problem to be solved.

[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0007] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0008] The embodiments of the present disclosure provide a voice interaction method and device, a home appliance, and a computer-readable storage medium, which can increase the rate of generating recommended topics and improve the user interaction experience.

[0009] In some embodiments, a voice interaction method is provided, which is applied to home appliances. The voice interaction method includes: parsing the first round of dialogue interaction information to obtain interaction keywords; querying relevant question texts in a question-and-answer text library based on the interaction keywords and user identity identification numbers; sorting the relevant question texts by relevance to obtain a relevant question text sequence; generating a recommended topic text based on the relevant question text sequence and outputting it to the user.

[0010] Optionally, the step of searching for relevant question texts in the question-and-answer pair text library based on the interaction keywords and the user identity identification number includes: determining the user question-and-answer pair sub-library corresponding to the user in the question-and-answer pair text library based on the user identity identification number; extracting keywords of the question text in the user question-and-answer pair sub-library to obtain topic keywords; and performing correlation matching on the interaction keywords and the topic keywords to determine the relevant question texts.

[0011] Optionally, the first round of dialogue interaction information includes the text to be queried generated according to the user's voice; the step of sorting the relevant question texts by relevance includes: sorting the relevant question texts in descending order according to the relevance between the relevant question texts and the text to be queried.

[0012] Optionally, the step of generating a recommended topic text and outputting it to the user based on the relevant question text sequence includes: calculating the similarity of the relevant question texts in the relevant question text sequence in sequence; when the similarity is less than a similarity threshold and the correlation is greater than a correlation threshold, determining the recommended topic text and outputting it to the user.

[0013] Optionally, the steps of calculating similarity of relevant question texts in the relevant question text sequence in sequence include: retrieving the historical question and answer pair sub-library corresponding to the user in the question and answer pair text library according to the user identity identification number; retrieving the relevant question texts in the relevant question text sequence in sequence; and calculating similarity between the retrieved relevant question texts and the question texts in the historical question and answer pair sub-library.

[0014] Optionally, the voice interaction method also includes: converting the collected user voice into text to obtain a first text; generating a second text based on the first text, based on a pre-trained decision model and knowledge graph; and generating a question-and-answer text library for each user based on the first text and the second text.

[0015] Optionally, the step of generating a second text based on the first text based on a pre-trained decision model and knowledge graph includes: extracting the subject and entity information of the first text; generating the second text based on the subject and entity information of the first text based on a pre-trained decision model and knowledge graph.

[0016] Optionally, the voice interaction method also includes: extracting the topic of the first text; generating a question-and-answer pair text library based on the first text and the second text, including: matching the first text, the topic of the first text, the user identity identification number and the interaction time to generate a historical question-and-answer pair library for each user; generating a user question-and-answer pair library based on each user's historical question-and-answer pair library and the second text; obtaining a question-and-answer pair text library based on each user's historical question-and-answer pair library and the user question-and-answer pair library.

[0017] Optionally, the step of generating a user question and answer pair library based on each user's historical question and answer pair library and the second text includes: filling the second text into a pre-written prompt engineering text to obtain a filled-in text; inputting the filled-in text into a large language model to generate a question and answer pair text; based on each user's historical question and answer pair library, filtering the question and answer pair text to generate a user question and answer pair library.

[0018] Optionally, after completing the first round of dialogue interaction, the voice interaction method further includes: when the operating state of the household appliance is abnormal, feeding back abnormal information to the user.

[0019] In some embodiments, a voice interaction device is provided, comprising a processor and a memory storing program instructions, wherein the processor is configured to execute the voice interaction method as described in any of the above embodiments when running the program instructions.

[0020] In some embodiments, a household appliance is provided, including: a device body; and the voice interaction device described in the above embodiments, installed on the device body.

[0021] In some embodiments, a computer-readable storage medium is provided, storing program instructions, which, when executed, are used to cause a computer to execute the voice interaction method as described in any of the above embodiments.

[0022] The voice interaction method and device, home appliance, and computer-readable storage medium provided by the embodiments of the present disclosure can achieve the following technical effects:

[0023] The voice interaction method provided by the embodiment of the present disclosure can pre-build a question-answer text library. During the interaction process, the relevant question text is directly queried in the question-answer text library through the interaction keywords and the user identity number, and then the relevant question text is sorted by relevance, and the recommended topic text is generated and output to the user. Compared with the related technology, it avoids selecting sub-topics from the knowledge graph and then generating topic recommendation sentences based on the sub-topics, thereby improving the generation rate of recommended topics and improving the user interaction experience.

[0024] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein:

[0026] Figure 1 is a schematic diagram of a household appliance provided by an embodiment of the present disclosure;

[0027] Figure 2 is a schematic diagram of a voice interaction method provided by an embodiment of the present disclosure;

[0028] Figure 3 is a schematic diagram of another voice interaction method provided by an embodiment of the present disclosure;

[0029] Figure 4 is a schematic diagram of another voice interaction method provided by an embodiment of the present disclosure;

[0030] Figure 5 is a schematic diagram of another voice interaction method provided by an embodiment of the present disclosure;

[0031] Figure 6 is a schematic diagram of another voice interaction method provided by an embodiment of the present disclosure;

[0032] Figure 7 It is a structural diagram of the voice interaction device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0033] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.

[0034] The terms "first", "second", etc. in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.

[0035] Unless otherwise stated, the term "plurality" means two or more.

[0036] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.

[0037] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.

[0038] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.

[0039] The household appliance provided by the embodiment of the present disclosure is as follows: Figure 1 As shown, the household appliance 10 includes a device body 110 and a voice interaction device 70 .

[0040] Optionally, the voice interaction device 70 is installed on the device body 110, and the voice interaction device 70 includes a processor. The processor can parse the first round of dialogue interaction information between the user and the home appliance to obtain interaction keywords. Then, according to the interaction keywords and the user identity document (ID), the relevant question text is queried in the question-answer text library. And the relevant question text can be sorted by relevance to obtain a relevant question text sequence. Then, according to the relevant question text sequence, a recommended topic text is generated and output to the user, so as to realize multiple rounds of interaction between the home appliance and the user.

[0041] Combination Figure 1The home appliance shown in the figure, the embodiment of the present disclosure provides a voice interaction method, which is applied to the home appliance, such as Figure 2 As shown, the method includes:

[0042] S201: The processor parses the first round of dialogue interaction information to obtain interaction keywords.

[0043] S202: The processor searches for relevant question texts in a question-answer text library according to the interaction keywords and the user identity number.

[0044] S203: The processor sorts the relevant question texts by relevance to obtain a relevant question text sequence.

[0045] S204: The processor generates a recommended topic text according to the relevant question text sequence and outputs it to the user.

[0046] By adopting the voice interaction method provided by the embodiment of the present disclosure, it is possible to pre-build a question-answer text library, parse the dialogue interaction information during the interaction process, and obtain interaction keywords. Through the interaction keywords and user ID, the relevant question text is directly queried in the question-answer text library, and then the relevant question text is sorted by relevance, and the recommended topic text is generated and output to the user. Compared with the related art, it avoids selecting sub-topics from the knowledge graph and then generating topic recommendation sentences based on the sub-topics, thereby improving the generation rate of recommended topics and improving the user interaction experience. In addition, after the home appliance completes the first round of dialogue interaction with the user, by generating a recommended topic text and outputting it to the user, the home appliance actively provides a new topic entry point for the interaction process, guiding the user to participate more in the dialogue interaction, so as to improve the initiative and intelligence of the home appliance in the voice interaction process.

[0047] Optionally, the step of searching for relevant question texts in the question-and-answer pair text library based on the interaction keywords and the user identity identification number includes: determining the user question-and-answer pair sub-library corresponding to the user in the question-and-answer pair text library based on the user identity identification number; extracting keywords of the question text in the user question-and-answer pair sub-library to obtain topic keywords; and performing correlation matching on the interaction keywords and the topic keywords to determine the relevant question texts.

[0048] In this embodiment, the question-answer pair text library includes user question-answer pair libraries of multiple users, and each user's user question-answer pair library corresponds to the user's ID one by one. The user question-answer pair library corresponding to the user can be determined according to the user ID, and the user question-answer pair library stores multiple question texts that the user may ask and the answers corresponding to the questions. By extracting the keywords of the question text in the user question-answer pair library, the topic keywords are obtained, and the topic keywords here refer to the subject information and entity information contained in each question text. The interactive keywords refer to the subject information and entity information contained in the first round of dialogue interaction information between the user and the home appliance. By matching the correlation between the interactive keywords and the topic keywords, the relevant question text can be quickly determined. By improving the determination rate of the relevant question text, the efficiency of generating recommended topics is further improved.

[0049] Exemplarily, the interaction keywords obtained are "Beijing, food, roast duck, XX village (a certain food brand)". The user ID is A00001, and the user question and answer sub-library corresponding to the user is determined according to the user ID. The user question and answer sub-library of the user includes question texts such as "Where is the store of XX village?", "When was the origin of XX village?", "What are the scenic spots and historical sites in Beijing?", "How does the roast duck taste?". Keyword extraction is performed on the question text in the user question and answer sub-library, and the topic keywords obtained are "XX village, store", "XX village, origin time", "Beijing, scenic spots and historical sites", "roast duck, taste". Then the interaction keywords and topic keywords are matched for relevance, and the related question texts determined are "Where is the store of XX village?", "When was the origin of XX village?", "How does the roast duck taste?".

[0050] Optionally, the first round of dialogue interaction information includes the text to be queried generated according to the user's voice; the step of sorting the relevant question texts by relevance includes: sorting the relevant question texts in descending order according to the relevance between the relevant question texts and the text to be queried.

[0051] In this embodiment, a descending sorted sequence of related question texts is generated in advance, so that a recommended topic text is generated according to the related question text sequence, thereby further improving the efficiency of generating recommended topics and improving the user interaction experience. In addition, the text to be queried refers to the text generated according to the user's voice question. The related question texts are sorted in descending order according to the correlation between the related question texts and the text to be queried, and the related question text sequence is obtained. Then, the recommended topic text is generated according to the related question text sequence, so that the generated recommended topic text is relevant to the interactive content in the first round of the user's dialogue interaction process, so as to improve the fluency and anthropomorphism of the dialogue interaction process.

[0052] For example, the query text generated according to the user voice is "What delicacies are there in Beijing?", and the related question texts are "Where is the store of XX village?", "When did XX village originate?", "How does roast duck taste?". Then, according to the correlation between the related question text and the query text, the related question texts are sorted in descending order, and the obtained related question text sequence is "How does roast duck taste?", "Where is the store of XX village?", "When did XX village originate?".

[0053] Furthermore, the step of sorting the relevant question text in descending order according to the correlation between the relevant question text and the text to be queried includes: inputting the relevant question text and the text to be queried into a text vectorization model to obtain a first vectorized text and a second vectorized text corresponding to the relevant question text and the text to be queried, respectively; inputting the first vectorized text and the second vectorized text into a topic model to sort the first vectorized text in descending order of topic relevance.

[0054] In this embodiment, topic relevance refers to the relevance between the topic of the first vectorized text and the topic of the second vectorized text. The text vectorization model may be a Transformer model or a Word2Vector model. By inputting the first vectorized text and the second vectorized text into the topic model, the topic model calculates the topic relevance with the first vectorized text, and sorts the first vectorized text in descending order according to the calculation result. The topic model may be an LDA (Latent Dirichlet Allocation) model or an LSA (Latent Semantic Analysis) model.

[0055] It should be noted that the text vectorization model and topic model given in this application are only some embodiments. The specific text vectorization model and topic model need to be selected by technical personnel according to actual needs and are not specified here.

[0056] Optionally, the step of generating a recommended topic text and outputting it to the user based on the relevant question text sequence includes: calculating the similarity of the relevant question texts in the relevant question text sequence in sequence; when the similarity is less than a similarity threshold and the correlation is greater than a correlation threshold, determining the recommended topic text and outputting it to the user.

[0057] In this embodiment, by calculating the similarity of the related question texts in the related question text sequence in advance, the related question texts whose similarity exceeds the threshold are excluded. The related question texts whose similarity is less than the similarity threshold and whose relevance is greater than the relevance threshold are determined as recommended topic texts and output to the user. By limiting the relevance of the determined recommended topic text to be greater than the relevance threshold, so that the recommended topic text and the text to be queried belong to the same topic center, the dialogue interaction process is carried out around the content that the user is interested in, thereby further improving the user's voice interaction experience.

[0058] Optionally, the steps of calculating similarity of relevant question texts in the relevant question text sequence in sequence include: retrieving the historical question and answer pair sub-library corresponding to the user in the question and answer pair text library according to the user identity identification number; retrieving the relevant question texts in the relevant question text sequence in sequence; and calculating similarity between the retrieved relevant question texts and the question texts in the historical question and answer pair sub-library.

[0059] In this embodiment, the historical question and answer pair sub-library corresponding to the user in the question and answer pair text library can be retrieved according to the user ID. The historical question and answer pair sub-library corresponding to the user stores the dialogue interaction information that multiple users have had with home appliances, and the dialogue interaction information includes the question texts that the users have asked and the recommended topic texts that the home appliances have output. The relevant question texts in the relevant question text sequence are retrieved in turn and the similarity is calculated with the question texts in the historical question and answer pair sub-library, so as to exclude the relevant question texts whose similarity is greater than the similarity threshold. Among them, the SimHash algorithm can be used to calculate the similarity between the relevant question text and the question text in the historical question and answer pair sub-library. By excluding the relevant question texts whose similarity is greater than the similarity threshold, the same recommended topic text is avoided from being generated and output to the user.

[0060] Exemplarily, the user ID is A00001, and the historical question-answer pair sub-library corresponding to the user is determined according to the user ID. The question texts included in the historical question-answer pair sub-library of the user are "What are the delicacies in Beijing?", "What are the stories of XX village?", and "How does the roast duck taste?". The sequence of related question texts is "How does the roast duck taste?", "Where is the store of XX village?", "When did XX village originate?". And the correlation between the related question texts in the related question text sequence and the text to be queried is 0.98, 0.95, and 0.93, respectively. Then retrieve the related question texts in the related question text sequence in turn. Then the related question text retrieved for the first time is "How does the roast duck taste?". The similarity calculation is performed on the retrieved related question text and the question text in the historical question-answer pair sub-library, and the similarity is 1. The similarity threshold is 0.9. Since 1>0.9, the related question text of "How does the roast duck taste?" is excluded. According to the sequence of the relevant question texts, the relevant question text "Where is the store in XX village?" is retrieved, and the similarity is calculated with the question text in the historical question-answering sub-database, and the similarity is 0.7. The relevant threshold is 0.91. Since 0.7<0.9, and 0.95>0.91, the recommended topic text is determined to be "Where is the store in XX village?", and the recommended topic text is output to the user.

[0061] It should be noted that the specific values ​​of the similarity threshold and the correlation threshold are set by technical personnel based on actual application scenarios and experimental research and are not specified here.

[0062] Combination Figure 3 As shown, the embodiment of the present disclosure provides another voice interaction method, including:

[0063] S301: The processor converts the collected user voice into text to obtain a first text.

[0064] S302, the processor generates a second text according to the first text based on a pre-trained decision model and knowledge graph.

[0065] S303: The processor generates a question-answer text library for each user according to the first text and the second text.

[0066] S304: The processor parses the first round of dialogue interaction information to obtain interaction keywords.

[0067] S305: The processor searches for relevant question texts in the question-answer text library according to the interaction keywords and the user identity number.

[0068] S306: The processor sorts the relevant question texts by relevance to obtain a relevant question text sequence.

[0069] S307: The processor generates a recommended topic text according to the relevant question text sequence and outputs it to the user.

[0070] By adopting the voice interaction method provided by the embodiment of the present disclosure, the collected user voice can be converted into text to obtain a first text, and then based on the pre-trained decision model and knowledge graph, a second text is generated according to the first text, thereby realizing the pre-generation of a question-answer text library for each user. Compared with the related art, it avoids selecting sub-topics from the knowledge graph in real time and then generating topic recommendation sentences according to the sub-topics, further improving the generation rate of recommended topics.

[0071] Optionally, the step of generating a second text based on the first text based on a pre-trained decision model and knowledge graph includes: extracting the subject and entity information of the first text; generating the second text based on the subject and entity information of the first text based on a pre-trained decision model and knowledge graph.

[0072] In this embodiment, the first text is a text generated by converting the user's voice, where the user's voice is the historical voice in the process of the user's dialogue and interaction with the home appliance, that is, the first text refers to the user's historical interaction text information. The pre-trained decision model can determine the depth information of the query's associated keywords according to the depth information of the entity information of the first text in the process of querying keywords in the knowledge graph, and determine the association relationship between the keywords according to the theme of the first text and the pre-set matching principle. Among them, the depth of the associated keywords is greater than the depth of the theme and entity information of the first text. By limiting the depth of the associated keywords to be greater than the depth of the entity information of the first text, it is convenient to limit the second text generated according to the theme and associated keywords of the first text to be a sub-topic of the first text, that is, a topic that the user may be interested in. Using the first text, capture the topic of interest to the user, and then generate the second text based on the pre-trained decision model and knowledge graph, that is, generate the topic that the user may be interested in, so as to further make the generated question-answer text library of each user meet the user's interest needs.

[0073] Exemplarily, the content of the first text is "What delicacies are there in Beijing? Beijing's delicacies include roast duck and XX village", the theme of the extracted first text is "delicious food", and the entity information is "Beijing, roast duck, XX village". In the knowledge graph, the storage depth of "Beijing" is 1, the storage depth of "roast duck", "XX village" and "Forbidden City" is 2, and the storage depth of "crispy", "golden" and "1890s" is 3. Among them, the association between "Beijing" and "roast duck" and "XX village" is "delicious food", the association between "Beijing" and "Forbidden City" is "historical sites", the association between "roast duck" and "crispy" is "taste", the association between "roast duck" and "golden" is "color", and the association between "XX village" and "1890s" is "origin time". The decision model determines that the depth information of the entity information of the first text is 2 (when multiple entity information corresponds to different depth information, the depth information corresponding to the entity information with the largest depth is taken), then the depth information of the queried associated keyword is 3. The theme of the first text is "food". According to the pre-set matching principle, that is, "food" matches "color" and "taste", the correlation between the keywords is determined to be "color" and "taste", and the second text generated from the knowledge graph is "the roast duck has a crispy taste" and "the roast duck has a golden color".

[0074] It should be noted that the extraction of the subject and entity information of the first text can be performed through an extraction model. The extraction model can be an LSTM+CRF (Long Short Term Memory+Conditional Random Field) model or a Span model. The decision model used for training can be a decision tree model. There is no restriction on the specific extraction model type and the decision model type used for training.

[0075] Combination Figure 4 As shown, the embodiment of the present disclosure provides another voice interaction method, including:

[0076] S401: The processor converts the collected user voice into text to obtain a first text.

[0077] S402, the processor generates a second text according to the first text based on a pre-trained decision model and knowledge graph.

[0078] S403: The processor extracts a topic for the first text.

[0079] S404, the processor matches the first text, the subject of the first text, the user identification number and the interaction time to generate a historical question-answer pair library for each user.

[0080] S405: The processor generates a user question-answer pair library based on each user's historical question-answer pair library and the second text.

[0081] S406: The processor obtains a question-answer pair text library based on each user's historical question-answer pair library and the user question-answer pair library.

[0082] S407: The processor parses the first round of dialogue interaction information to obtain interaction keywords.

[0083] S408: The processor searches for relevant question texts in the question-answer text library according to the interaction keywords and the user identity number.

[0084] S409: The processor sorts the relevant question texts by relevance, and obtains a relevant question text sequence.

[0085] S410: The processor generates a recommended topic text according to the relevant question text sequence and outputs it to the user.

[0086] By using the voice interaction method provided by the embodiment of the present disclosure, it is possible to pre-generate a historical question-answer pair library for each user by matching the first text, the subject of the first text, the user ID, and the interaction time. By pre-generating a historical question-answer pair library for each user, the recommended topic generation rate is further improved.

[0087] Optionally, the step of generating a user question and answer pair library based on each user's historical question and answer pair library and the second text includes: filling the second text into a pre-written prompt engineering text to obtain a filled-in text; inputting the filled-in text into a large language model to generate a question and answer pair text; based on each user's historical question and answer pair library, filtering the question and answer pair text to generate a user question and answer pair library.

[0088] In this embodiment, the question-answer pair text can be filtered based on each user's historical question-answer pair database, and the user question-answer pair database can be pre-generated. By pre-generating the user question-answer pair database, the recommended topic generation rate is further improved. In addition, before generating the user question-answer pair database, the question-answer pair text is filtered based on each user's historical question-answer pair database to filter out the same text as in the user's historical interaction information, further avoiding repeated generation of the same recommended topic text.

[0089] Exemplarily, the pre-written prompt engineering text is "Extract key information from the above text and organize it into questions, and query scientific and rigorous answers in the text based on the questions", and the second text is "The roast duck has a crispy taste" and "The roast duck has a golden color." The second text is filled into the prompt engineering text to obtain the filled-in text "Extract key information from 'The roast duck has a crispy taste' and organize it into questions, and query scientific and rigorous answers in 'The roast duck has a crispy taste' based on the questions" and "Extract key information from 'The roast duck has a golden color' and organize it into questions, and query scientific and rigorous answers in 'The roast duck has a golden color' based on the questions." The filled-in text is input into the Large Language Model (LLM) to generate question-answer text pairs "Q: How does the roast duck taste? A: Crispy taste" and "Q: How is the color of the roast duck? A: Golden color." Then, based on the user's historical question-answer pair database, the question-answer pair text is filtered. The filtering method can use the SimHash algorithm to calculate the similarity between the question-answer pair text and the historical question-answer pair text in the user's historical question-answer pair database. If the similarity is greater than the similarity threshold, the question-answer pair text is deleted. The question-answer pair texts with similarity less than the similarity threshold are saved to generate the user question-answer pair database.

[0090] Combination Figure 5 As shown, the embodiment of the present disclosure provides another voice interaction method, including:

[0091] S501: The processor collects the user's first round of question voice.

[0092] S502: The processor pre-processes the first round of question speech.

[0093] S503, the processor converts the pre-processed first round of question speech into text to be queried.

[0094] S504: The processor searches for answers in the question-answer text library based on the text to be queried and outputs the answers to the user, completing the first round of dialogue interaction.

[0095] S505: The processor parses the first round of dialogue interaction information to obtain interaction keywords.

[0096] S506: The processor searches for relevant question texts in the question-answer text library according to the interaction keywords and the user identity number.

[0097] S507: The processor sorts the relevant question texts by relevance to obtain a relevant question text sequence.

[0098] S508: The processor generates a recommended topic text according to the relevant question text sequence and outputs it to the user.

[0099] By adopting the voice interaction method provided in the embodiment of the present disclosure, the user's first round of question voice can be collected through a voice receiver, and the voice receiver can be a pickup, a microphone or a mobile phone. The first round of question voice is preprocessed, and the preprocessing includes one or more operations of noise reduction, echo removal, and reverberation removal. By preprocessing the first round of question voice, the quality of the first round of question voice is improved, and then the accuracy of converting the first round of question voice into the text to be queried is improved. The first round of question voice after preprocessing can be converted into the text to be queried through the Whisper model, and the first round of question voice is converted into the text to be queried by using the Whisper model, so as to further improve the accuracy of the conversion between the first round of question voice and the text to be queried.

[0100] Combination Figure 6 As shown, the embodiment of the present disclosure provides another voice interaction method, including:

[0101] S601: The processor collects the user's first round of question voice.

[0102] S602: The processor pre-processes the first round of question speech.

[0103] S603, the processor converts the pre-processed first round of question speech into text to be queried.

[0104] S604: The processor searches for answers in the question-answer text library based on the text to be queried and outputs the answers to the user, completing the first round of dialogue interaction.

[0105] S605: When the operation status of the home appliance is abnormal, the processor feeds back abnormal information to the user.

[0106] S606: When the operation state of the home appliance is stable, the processor parses the first round of dialogue interaction information to obtain interaction keywords.

[0107] S607: The processor searches for relevant question texts in the question-answer text library according to the interaction keywords and the user identity number.

[0108] S608: The processor sorts the relevant question texts by relevance to obtain a relevant question text sequence.

[0109] S609: The processor generates a recommended topic text according to the relevant question text sequence and outputs it to the user.

[0110] By adopting the voice interaction method provided by the embodiment of the present disclosure, after the first round of dialogue interaction is completed, when the operating state of the home appliance is abnormal, the abnormal information can be fed back to the user first. Feedback of abnormal state of home appliance is realized, so that the user can make timely adjustments when the operating state of the home appliance is abnormal, thereby improving the safety of the home appliance. In addition, by limiting the abnormal information feedback to after the first round of dialogue interaction is completed, it is avoided that the home appliance suddenly feeds back abnormal information, disturbs the user, and affects the user experience.

[0111] Optionally, the voice interaction method also includes: acquiring operating parameters of the household appliance in real time; comparing the operating parameters with parameter thresholds to determine the operating status of the household appliance.

[0112] In this embodiment, the operating parameters of the household appliance are acquired in real time, and then the operating status of the household appliance is determined by comparing the operating parameters with parameter thresholds, so as to facilitate real-time monitoring of the household appliance.

[0113] The parameter threshold can be set by the user or determined by the user's historical setting parameters. For example, when the home appliance is a refrigerator, the user's historical setting temperature of the refrigerator is 4°C, 5°C, and 6°C. It is determined that the user's historical setting temperature range of the refrigerator is 4-6°C, and the corresponding parameter thresholds are 4°C and 6°C. The refrigerator's refrigerator temperature is obtained in real time. When the temperature of the refrigerator is within the range of 4-6°C, it is determined that the operating state of the refrigerator is stable at this time. When the temperature of the refrigerator is 8°C, which is not within the range of 4-6°C, it is determined that the operating state of the refrigerator is abnormal at this time. After the user wakes up the refrigerator and completes the first round of dialogue interaction, the abnormal information "Your refrigerator's permanent temperature is 4-6°C, and the current refrigerator's temperature is 8°C. Do you want to re-set the refrigerator's temperature?". For example, when the home appliance is an air conditioner, the user's set temperature is 26°C, and the corresponding parameter threshold is 26°C. The air outlet temperature of the air conditioner is obtained in real time. When the air outlet temperature reaches 26°C and the air conditioner continues to cool or heat, it is determined that the operating status of the air conditioner is abnormal. After the user wakes up the air conditioner and completes the first round of dialogue interaction, the abnormal information is fed back to the user, "Your set temperature is 26°C, and the current temperature deviates from 26°C. Do you want to reduce the wind speed?"

[0114] Combination Figure 7As shown, the embodiment of the present disclosure provides a voice interaction 70, including a processor (processor) 710 and a memory (memory) 720. Optionally, the device 70 may also include a communication interface (Communication Interface) 730 and a bus 740. Among them, the processor 710, the communication interface 730, and the memory 720 can communicate with each other through the bus 740. The communication interface 730 can be used for information transmission. The processor 710 can call the logic instructions in the memory 720 to execute the voice interaction method of the above embodiment.

[0115] In addition, the logic instructions in the memory 720 described above may be implemented in the form of software functional units and when sold or used as independent products, may be stored in a computer-readable storage medium.

[0116] The memory 720 is a computer-readable storage medium that can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 710 executes the function application and data processing by running the program instructions / modules stored in the memory 720, that is, implementing the voice interaction method in the above embodiment.

[0117] The memory 720 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 720 may include a high-speed random access memory and may also include a non-volatile memory.

[0118] Combination Figure 1 As shown, the embodiment of the present disclosure provides a household appliance 10, including: a device body 110, and the above-mentioned voice interaction device 70. The voice interaction device 70 is installed on the device body 110. The installation relationship described here is not limited to being placed inside the device body 110, but also includes the installation connection with other components of the household appliance 10, including but not limited to physical connection, electrical connection or signal transmission connection, etc. It can be understood by those skilled in the art that the voice interaction device 70 can be adapted to a feasible household appliance 10, thereby realizing other feasible embodiments.

[0119] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above-mentioned voice interaction method.

[0120] The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiment of the present disclosure. The aforementioned storage medium may be a non-transient storage medium, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes.

[0121] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible changes. Unless explicitly required, separate components and functions are optional, and the order of operation may vary. The parts and features of some embodiments may be included in or replace the parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of listings containing one or more associated ones. In addition, when used in the present application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the presence of other identical elements in the process, method or device comprising the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.

[0122] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods for each specific application to implement the described functions, but such implementations should not be considered to exceed the scope of the embodiments of the present disclosure. The technicians may clearly understand that, for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above may refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.

[0123] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.

[0124] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A voice interaction method, applied to home appliances, characterized in that: include: Parse the first round of dialogue interaction information and obtain interaction keywords; According to the interaction keywords and user identification number, relevant question texts are searched in the question-answer pair text library; Sort the relevant question texts by relevance to obtain a relevant question text sequence; Based on the sequence of relevant question texts, the recommended topic text is generated and output to the user.

2. The voice interaction method according to claim 1, characterized in that: The steps of searching for relevant question texts in the question-answer pair text library according to the interaction keywords and the user identification number include: Determine, according to the user identification number, a user question-answer pair sub-library corresponding to the user in the question-answer pair text library; Extract keywords from the question text in the user question-answer pair database to obtain topic keywords; Perform relevance matching on interactive keywords and topic keywords to determine relevant question texts.

3. The voice interaction method according to claim 1 or 2, characterized in that: The steps of generating a recommended topic text output to the user according to the relevant question text sequence include: Calculate the similarity of the relevant question texts in the relevant question text sequence in sequence; When the similarity is less than the similarity threshold and the relevance is greater than the relevance threshold, the recommended topic text is determined and output to the user.

4. The voice interaction method according to claim 1 or 2, characterized in that: The method further comprises: Convert the collected user voice into text to obtain a first text; According to the first text, a second text is generated based on a pre-trained decision model and knowledge graph; A question-answer pair text library for each user is generated according to the first text and the second text.

5. The voice interaction method according to claim 4, characterized in that: The step of generating a second text according to the first text based on a pre-trained decision model and a knowledge graph includes: Extracting subject and entity information of the first text; According to the subject and entity information of the first text, the second text is generated based on the pre-trained decision model and knowledge graph.

6. The voice interaction method according to claim 4, characterized in that: The method further comprises: extracting a topic for the first text; The step of generating a question-answer pair text library according to the first text and the second text comprises: Matching the first text, the subject of the first text, the user identification number, and the interaction time to generate a historical question-answer pair database for each user; Generate a user question-answer pair database based on each user's historical question-answer pair database and the second text; Based on each user's historical question-answer pair database and the user question-answer pair database, a question-answer pair text database is obtained.

7. The voice interaction method according to claim 1 or 2, characterized in that: After completing the first round of dialogue interaction, the method further includes: When the operating status of the home appliance is abnormal, abnormal information is fed back to the user.

8. A voice interaction device, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the voice interaction method according to any one of claims 1 to 7 when running the program instructions.

9. A household appliance, characterized in that: include: Equipment body; The voice interaction device as described in claim 8 is installed on the device body.

10. A computer-readable storage medium storing program instructions, characterized in that: When the program instructions are executed, the computer is used to execute the voice interaction method as described in any one of claims 1 to 7.