Voice question answering method and device, refrigerator and computer readable storage medium

By using pre-trained code generation models and decision models in the refrigerator intelligent question-and-answer system, data related to user problems is directly queried in the database, which solves the problem that knowledge graphs in the prior art are difficult to cover user problems and frequent updates, and improves the robustness of the system.

CN119961494APending Publication Date: 2025-05-09QINDAO HAIER REFRIGERATOR CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311483524.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art realizes intelligent question-and-answer in refrigerators by building knowledge graphs. However, this method is difficult to cover all problems when facing users' flexible and diverse problems, and frequent updates to knowledge graphs may lead to errors, which is not conducive to robustness.

Method used

The pre-trained code generation model is used to convert user speech into query language, query data related to the question text in the database, and obtain answers through the decision model, avoiding the construction and updating of the knowledge graph.

Benefits of technology

It improves the robustness of the voice Q&A method, can effectively deal with user flexible and diverse problems, and avoids the error risk caused by knowledge graph updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961494A_ABST
    Figure CN119961494A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent household appliances, and discloses a voice question answering method and device, a refrigerator and a computer readable storage medium, the voice question answering method is applied to the refrigerator, and the method comprises the following steps: converting collected user voice into a question text; querying data related to the question text in a database through a plurality of query modes, the plurality of query modes including querying the data related to the question text in the database by using a pre-trained code generation model; and inputting the data related to the question text into a pre-trained decision model, and obtaining target data as an answer. According to the method, the construction of the knowledge graph is avoided, and the knowledge graph does not need to be frequently updated in the face of the problem of flexibility and diversity of users, so that the robustness is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of smart home appliances, for example, to a voice question-and-answer method and device, a refrigerator, and a computer-readable storage medium. Background Art

[0002] With the development of science and technology, household appliances are becoming more intelligent. For example, many refrigerators on sale can not only be used to refrigerate and freeze food, but also have intelligent question-and-answer functions. The refrigerators can answer users' questions and interact with users based on pre-stored data. However, as the amount of stored data increases, how to quickly and accurately obtain answers that meet user needs from a large amount of data has become a research hotspot.

[0003] In order to quickly and accurately retrieve answers to user questions from a large amount of knowledge content, the relevant technology discloses an intelligent question-answering method, device and storage medium for refrigerators. The intelligent question-answering method includes: performing data cleaning on existing question-answer pairs; expanding the cleaned user question data based on a deep learning method, and establishing a corresponding relationship table between questions and answers; vectorizing the expanded user question data using a deep learning related model to obtain vectorized text; using a graph-based algorithm in the search field to store the vectorized text in a multi-level graph structure to obtain a question vector graph; using a deep learning-based model to classify the vectorized text into pre-defined categories; based on the multi-level question vector graph, using a graph-based algorithm in the search field to recall user questions; combining various similarity matching methods to accurately sort the recall results, select the question with the highest score as the match of the user question, and obtain the answer to the question according to the question and answer correspondence table and return it.

[0004] In the process of implementing the embodiments of the present disclosure, it is found that there are at least the following problems in the related art:

[0005] Although the related technology uses a graph-based algorithm to store refrigerator-related questions that users may ask in the form of a multi-level graph, which speeds up the answer retrieval speed. However, this method requires the pre-construction of a knowledge graph, and users' questions are flexible and diverse. It is difficult for the pre-constructed knowledge graph to cover all users' questions. In order to cope with the diversity of users, the knowledge graph needs to be frequently updated and integrated with previously stored data. Because the amount of knowledge stored in the knowledge graph is huge, frequent updates may lead to knowledge fusion errors, which is not conducive to the robustness of the knowledge graph. Therefore, the method of implementing knowledge question and answer by constructing a knowledge graph is less robust.

[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0007] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0008] The embodiments of the present disclosure provide a voice question-answering method and device, a refrigerator, and a computer-readable storage medium. When faced with flexible and diverse questions from users, there is no need to build and update a knowledge graph, thereby improving the robustness of the voice question-answering method.

[0009] In some embodiments, a voice question-answering method is provided, which is applied to a refrigerator, and the method includes: converting the collected user voice into question text; querying data related to the question text in a database through multiple query modes, and the multiple query modes include using a pre-trained code generation model to query data related to the question text in a database; inputting data related to the question text into a pre-trained decision model to obtain target data as an answer.

[0010] Optionally, the step of using a pre-trained code generation model to query data related to the question text in the database includes: using the pre-trained code generation model to convert the question text into a query language; and querying the database for related data according to the converted query language.

[0011] Optionally, the multiple query modes also include: using a pre-trained text vectorization model to query data related to the question text in the database; and / or using a search engine to query data related to the question text in the database.

[0012] Optionally, the step of using a pre-trained text vectorization model to query data related to the question text in the database includes: inputting the question text into the pre-trained text vectorization model to obtain vectorized text; and querying data related to the question text in the database based on the vectorized text.

[0013] Optionally, the step of pre-training the text vectorization model includes: selecting a vectorization model; pre-processing the training data; and inputting the pre-processed training data into the selected vectorization model for pre-training.

[0014] Optionally, the step of using a search engine to query data related to the question text in a database includes: vectorizing the data in the database through a pre-trained vectorization model to obtain vectorized data; storing the vectorized data in a vector database; packaging the question text with prompt language to limit the question text; vectorizing the question text after prompt language packaging through a pre-trained vectorization model to obtain vectorized question text; inputting the vectorized question text into a Transformer encoder for encoding to obtain a question encoding sequence; and querying relevant data in the vector database using a search field based on an LSH algorithm.

[0015] Optionally, the step of inputting data related to the question text into a pre-trained decision model to obtain target data as an answer includes: inputting data related to the question text into a pre-trained decision model to obtain data with the highest correlation with the question text as target data; and using a multimodal large model to perform stylized transformation on the target data to generate an answer.

[0016] Optionally, the voice question-answering method further includes: automatically sorting the training data according to the relevance of the training data to obtain a first sequence; manually correcting the first sequence according to the relevance of the training data to obtain a second sequence; and training a decision model through a reinforcement learning algorithm based on the first sequence and the second sequence.

[0017] Optionally, the voice question-answering method further includes: extracting information from the acquired initial data to obtain extracted data, wherein the extracted data includes text data and image data; re-layouting the extracted data; and saving the re-layouted extracted data to generate a database.

[0018] Optionally, the voice question-answering method further includes: acquiring original data; filtering the original data; classifying and saving the filtered original data to obtain initial data.

[0019] Optionally, the step of converting the collected user voice into question text includes: collecting the user voice; cleaning the user voice; and converting the cleaned user voice into question text.

[0020] In some embodiments, a voice question and answer device is provided, including: a conversion module, configured to convert the collected user voice into question text; a query module, configured to query data related to the question text in a database through multiple query modes, and the multiple query modes include using a pre-trained code generation model to query data related to the question text in the database; an analysis module, configured to input data related to the question text into a pre-trained decision model to obtain target data as an answer.

[0021] In some embodiments, a voice question-and-answer device is provided, comprising a processor and a memory storing program instructions, wherein the processor is configured to execute the voice question-and-answer method described in any of the above embodiments when running the program instructions.

[0022] In some embodiments, a refrigerator is provided, comprising: a refrigerator body; and a voice question-and-answer device as described in any of the above embodiments, installed on the refrigerator body.

[0023] In some embodiments, a computer-readable storage medium is provided, storing program instructions, which, when executed, are used to cause a computer to execute the voice question-answering method as described in any of the above embodiments.

[0024] The voice question-answering method and device, refrigerator, and computer-readable storage medium provided by the embodiments of the present disclosure can achieve the following technical effects:

[0025] The voice question-answering method provided by the disclosed embodiment can directly retrieve data related to the user's question in the database using a pre-trained code generation model, and then obtain the answer to the user's question through the relevant data and decision model. Compared with the related technology, this method avoids the construction of a knowledge graph, and when facing flexible and diverse user questions, it does not need to frequently update the knowledge graph, so it is more robust.

[0026] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein:

[0028] Figure 1 A voice question-answering method provided by an embodiment of the present disclosure;

[0029] Figure 2 It is another voice question and answer method provided by the embodiment of the present disclosure;

[0030] Figure 3 It is another voice question and answer method provided by the embodiment of the present disclosure;

[0031] Figure 4 It is another voice question and answer method provided by the embodiment of the present disclosure;

[0032] Figure 5 It is another voice question and answer method provided by the embodiment of the present disclosure;

[0033] Figure 6is a schematic diagram of a voice question-and-answer device provided by an embodiment of the present disclosure;

[0034] Figure 7 is a schematic diagram of another voice question-and-answer device provided by an embodiment of the present disclosure;

[0035] Figure 8 Schematic diagram of a refrigerator provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0036] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.

[0037] The terms "first", "second", etc. in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.

[0038] Unless otherwise stated, the term "plurality" means two or more.

[0039] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.

[0040] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.

[0041] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.

[0042] In the related art, a method, device and storage medium for intelligent question and answer of refrigerators are disclosed. The intelligent question and answer method includes: cleaning the existing question and answer pairs; expanding the cleaned user question data based on the deep learning method, and establishing a corresponding relationship table between the question and the answer; vectorizing the expanded user question data using the relevant model of deep learning to obtain the vectorized text; using the graph-based algorithm in the search field to store the vectorized text in a multi-level graph structure to obtain the question vector graph; using the model based on deep learning to classify the vectorized text into pre-defined categories; based on the multi-level question vector graph, using the graph-based algorithm in the search field to recall the user question; combining various similarity matching methods to accurately sort the recall results, select the question with the highest score as the match of the user question, and obtain the answer to the question according to the question and answer correspondence table and return it.

[0043] Although the related technology uses a graph-based algorithm to store refrigerator-related questions that users may ask in the form of a multi-level graph, which speeds up the answer retrieval speed. However, this method requires the pre-construction of a knowledge graph, and users' questions are flexible and diverse. It is difficult for the pre-constructed knowledge graph to cover all users' questions. In order to cope with the diversity of users, the knowledge graph needs to be frequently updated and integrated with previously stored data. Because the amount of knowledge stored in the knowledge graph is huge, frequent updates may lead to knowledge fusion errors, which is not conducive to the robustness of the knowledge graph. Therefore, the method of implementing knowledge question and answer by constructing a knowledge graph is less robust.

[0044] The disclosed embodiment provides a voice question-answering method, which does not need to build and update a knowledge graph when facing flexible and diverse questions from users, thereby improving the robustness of the voice question-answering method. The voice question-answering method can be run on a refrigerator, which generates corresponding answers based on the received user voice questions. The refrigerator includes a processor and a memory storing program instructions. It may also include a communication interface and a bus. The processor, the communication interface, and the memory communicate with each other via the bus. The communication interface is used for information transmission.

[0045] Combination Figure 1 As shown, the voice question answering method comprises:

[0046] S101, the processor converts the collected user voice into question text.

[0047] S102, the processor searches the database for data related to the question text through multiple query modes, and the multiple query modes include using a pre-trained code generation model to search the database for data related to the question text.

[0048] S103, the processor inputs the data related to the question text into a pre-trained decision model to obtain target data as an answer.

[0049] By using the voice question-answering method provided by the embodiment of the present disclosure, it is possible to use a pre-trained code generation model to directly retrieve data related to the user's question in the database, and then obtain the answer to the user's question through the relevant data and the pre-trained decision model. Compared with related technologies, this method has stronger adaptability and reliability when facing flexible and diverse user questions. At the same time, it avoids the construction of a knowledge graph and does not need to frequently update the knowledge graph, so it is more robust.

[0050] Optionally, the step of using a pre-trained code generation model to query data related to the question text in the database includes: using the pre-trained code generation model to convert the question text into a query language; and querying the database for related data according to the converted query language.

[0051] In this embodiment, the question text is converted into a corresponding query language by using a pre-trained code generation model, and then the data associated with the question text is queried in the database to improve the query speed and accuracy of the relevant data.

[0052] Specifically, the code generation model is pre-trained so that it has the ability to "extract entities and relations in sentences, standardize entity professional vocabulary, and generate database query language." The problem text to be searched is input into the trained code generation model to extract entities and relations and generate database query language. Relevant data is queried in the database according to the generated database query language. Since the pre-trained code generation model can automatically extract entity keywords in sentences and then splice them to generate query language, and then query the database for relevant data, it is concise and can improve the query rate of relevant data. In addition, since the pre-trained code generation model can perform entity professional vocabulary standardization on the entity keywords in the extraction, and splice the query language based on the entity keywords after the entity professional vocabulary standardization, it has higher accuracy.

[0053] Exemplarily, the question text to be queried is "What is the use of the smart freezing function in the refrigerator". The question text to be queried is input into a pre-trained code generation model, and entities and relationships are extracted for the question text to be queried. The extracted entities are "refrigerator" and "smart freezing", and the relationships are "function" and "effect", among which the standardized vocabulary of "smart freezing" is "deep cold smart freezing function". Standardize the entity professional vocabulary of "smart freezing" and convert "smart freezing" into "deep cold smart freezing function". Then execute the code "MATCH(n)-[:effect]->(fridge:refrigerator{function:'deep cold smart freezing function'});RETURN n" to obtain the query language corresponding to the question text to be queried. According to the obtained query language, complete the relevant data query in the database to obtain the relevant data of the question text to be queried.

[0054] Furthermore, the code generation model includes a NL2GQL (Natrual Language to Graph Query Language) model. The NL2GQL model is a set of models that use the code generation capabilities of a large model to convert natural language into a graph database query language. When the code generation model is an NL2GQL model, the pre-trained NL2GQL model can convert the question text into the GQL language. GQL (Graph Query Language) is a query language that is particularly advantageous for querying graph databases. Compared with related technologies, it does not require the pre-construction of a question vector graph. GQL is simple in implementing queries on association relationships, and therefore has a faster query rate.

[0055] Furthermore, the code generation model includes a NL2SQL (Natural Language to Structured Query Language) model. NL2SQL is a technology that converts a user's natural statements into executable SQL statements. The NL2SQL model has high accuracy and robustness in achieving natural language to SQL conversion. It can process various types of natural language inputs and generate SQL query statements that conform to grammatical rules. SQL (Structured Query Language) has the advantages of being fast, concise, easy to learn and use, and can therefore improve data management and analysis efficiency.

[0056] It should be noted that the selection of the code generation model needs to be made by technical personnel based on the actual application scenario and purpose of use, and is not specified here.

[0057] Optionally, the multiple query modes also include: using a pre-trained text vectorization model to query data related to the question text in the database; and / or using a search engine to query data related to the question text in the database.

[0058] In this embodiment, the multiple query modes include, in addition to using a pre-trained code generation model to query data related to the question text in the database, using a pre-trained text vectorization model to query data related to the question text in the database, and / or using a search engine to query data related to the question text in the database. Since the pre-trained text vectorization model can be used to query text data related to the question text in the database, and the search engine and the pre-trained code generation model can be used to query text data and image data related to the question text in the database, the combination of multiple query modes can increase the types and quantities of data related to the question text, thereby improving the diversity and accuracy of answers generated based on the relevant data.

[0059] Optionally, the step of using a pre-trained text vectorization model to query data related to the question text in the database includes: inputting the question text into the pre-trained text vectorization model to obtain vectorized text; and querying data related to the question text in the database based on the vectorized text.

[0060] In this embodiment, the purpose of text vectorization is to represent the text as a series of vectors that can express the semantics of the text. By pre-training the text vectorization model, the vectorized text obtained by the text vectorization model expresses the text semantics under the vertical field. A vertical field refers to a small field that is vertically subdivided under a large field, and vertical refers to vertical extension rather than horizontal expansion. Exemplarily, the question text to be queried is "What is maternal and child management", and the text queried after semantic parsing in a non-vertical field is "maternal and child management", which is a management model, that is, the management of maternal and child life. If the specific field is limited to a refrigerator, then the text queried after semantic parsing in its vertical field is "maternal and child management", and the semantic parsing is the maternal and child area in the refrigerator. Since the pre-trained text vectorization model has a stronger semantic parsing ability for a specific field, the accuracy of querying data related to the question text in the database through the pre-trained text vectorization model is higher.

[0061] Optionally, the step of pre-training the text vectorization model includes: selecting a vectorization model; pre-processing the training data; and inputting the pre-processed training data into the selected vectorization model for pre-training.

[0062] In this embodiment, the selected vectorization model can be an LSTM (long-short term memory) model or a Transformer model. The training data is all texts in a specific field. Preprocessing the training data includes removing noise and / or removing duplicate texts. The quality and reliability of the training data are improved by preprocessing the training data. At the same time, duplicate data is reduced to improve the training efficiency of the vectorization model. By pre-training the text vectorization model and limiting the training data to data in a specific field, it has a stronger semantic parsing ability for the specific field.

[0063] It should be noted that the selection of the vectorized model to be trained needs to be made by technical personnel based on the actual application scenario and purpose of use, and is not specified here.

[0064] Optionally, the step of using a search engine to query data related to the question text in a database includes: vectorizing the data in the database through a pre-trained vectorization model to obtain vectorized data; storing the vectorized data in a vector database; packaging the question text with prompt language to limit the question text; vectorizing the question text after prompt language packaging through a pre-trained vectorization model to obtain vectorized question text; inputting the vectorized question text into a Transformer encoder for encoding to obtain a question encoding sequence; and querying relevant data in the vector database using a search field based on an LSH algorithm.

[0065] In this embodiment, the data in the database is vectorized to obtain a vector database. Then the question text is packaged with a prompt language to limit the question text. For example, the question text is "What is the cleaning method of the refrigerator?" and the prompt language is "If you are a professional planner, rigorously search for the text content related to the question from a professional perspective." Then the question text is packaged with a prompt language to obtain "If you are a professional, rigorously search for the text content related to the question from a professional perspective. Question: What is the cleaning method of the refrigerator?" The packaged question text is vectorized, and then the vectorized question text is input into the Transformer encoder to obtain a question encoding sequence. The role of the LSH (Locality Sensitive Hashing) algorithm is to mine similar data from a large amount of data. For example, the question text is "What is the cleaning method of the refrigerator?", and the search field is based on the LSH algorithm to query the relevant data in the vector database, such as "It is best to clean the refrigerator shell every day, wipe the refrigerator shell and handle every day with a slightly damp soft cloth", "The air-cooled refrigerator freezer does not frost, and it is very easy to clean." By packaging the question text with prompt language, the question text is limited to improve the accuracy of querying related data based on the prompt language. At the same time, by using the search field based on the LSH algorithm to query related data in the vector database, multiple data related to the question text are obtained, further increasing the amount of related data obtained.

[0066] It should be noted that, in the step of using a search engine to query data related to the question text in the database, the pre-trained vectorization model used can be the same model as the pre-trained text vectorization model used in the step of using a pre-trained text vectorization model to query data related to the question text in the database, and the pre-trained text vectorization model can also be other vectorization models, as long as the data in the database can be vectorized to obtain vectorized data.

[0067] Combination Figure 2 As shown, another voice question-answering method provided by an embodiment of the present disclosure includes:

[0068] S201, the processor converts the collected user voice into question text.

[0069] S202, the processor searches the database for data related to the question text through multiple query modes, and the multiple query modes include using a pre-trained code generation model to search the database for data related to the question text.

[0070] S203, the processor inputs the data related to the question text into a pre-trained decision model to obtain the data with the highest correlation with the question text as the target data.

[0071] S204: The processor uses the multimodal large model to perform stylized transformation on the target data and generate an answer.

[0072] The voice question-answering method provided by the disclosed embodiment can directly input multiple data related to the question text queried in the database through multiple query modes into a pre-trained decision model, and obtain the data with the highest correlation with the question text as the target data through the pre-trained decision model. The target data is then input into the multimodal large model, and the target data is stylized and converted by the multimodal large model to generate an answer. Among them, the convertible style is pre-designed by the technician and forms a corresponding conversion template, and the user selects the style according to the conversion template. The multimodal large model performs stylized conversion on the target data according to the style selected by the user to generate a stylized answer that meets the user's needs, so as to enhance the personalization of the generated answer. Exemplarily, the multimodal large model can be a KOSMOS-1 or BLIP (Bootstrapping Language-Image Pre-training, classic visual language pre-training) model.

[0073] In addition, in the process of natural language processing, the large model can generate output text based on the given text. In the process of answer generation through the large model, since the content of the answer generated by the generative model is uncontrollable, the answer generation through the large model is prone to hallucination problems, that is, generating false information. The method provided by the present disclosure alleviates the large model hallucination problem by limiting the input data input to the multimodal large model to the data with the highest correlation with the question text before using the multimodal large model to generate answers, thereby further improving the accuracy of the generated answers.

[0074] Optionally, the voice question-answering method further includes: automatically sorting the training data according to the relevance of the training data to obtain a first sequence; manually correcting the first sequence according to the relevance of the training data to obtain a second sequence; and training a decision model through a reinforcement learning algorithm based on the first sequence and the second sequence.

[0075] In this embodiment, the training data can be automatically sorted according to the relevance of the training data to obtain a first sequence, and then the first sequence is manually corrected to obtain a second sequence. The first sequence and the second sequence are input into the training model, and the decision model is trained by the reinforcement learning algorithm, so that the trained decision model can be sorted according to the relevance of the input data. By pre-training the decision model, by inputting the relevant data into the decision model, the data with the highest relevance to the question text can be directly obtained as the target data, so as to simplify the question-answering method answer generation process and improve the question-answering method answer generation rate.

[0076] Combination Figure 3 As shown, another voice question-answering method provided by an embodiment of the present disclosure includes:

[0077] S301, the processor extracts information from the acquired initial data to obtain extracted data, where the extracted data includes text data and image data.

[0078] S302: The processor rearranges the extracted data.

[0079] S303: The processor saves the rearranged extracted data and generates a database.

[0080] S304: The processor converts the collected user voice into question text.

[0081] S305, the processor searches the database for data related to the question text through multiple query modes, and the multiple query modes include using a pre-trained code generation model to search the database for data related to the question text.

[0082] S306, the processor inputs the data related to the question text into a pre-trained decision model to obtain target data as an answer.

[0083] The voice question-answering method provided by the disclosed embodiment can extract information from the acquired initial data to obtain the extracted data. Since the acquired initial data is diverse, the extraction of the initial data can be performed by extracting information from the image data through the Textract tool, and by extracting information from the initial data containing mixed information such as images, text, and tables through OCR (Optical Character Recognition) or Spatial-Aware disentangled Attention (spatial-aware decomposition attention mechanism) to obtain diverse extracted data, thereby increasing the diversity of data in the database. Re-layout of the extracted extracted data refers to text-image alignment of the extracted data. Text-image alignment refers to matching according to text semantics and image meaning, so that the text and the image related to the text are matched together. The storage form for saving the rearranged extracted data can be distributed storage. Distributed storage is achieved by slicing the data before saving the data. By slicing the data, the data processing efficiency and the utilization rate of the storage space can be improved.

[0084] Furthermore, the step of rearranging the extracted data includes: performing feature extraction on the text data and image data in the extracted data respectively to obtain text features and image features; using a binary classification function to determine whether the text features and image features match; and pairing the text data and image data whose text features and image features match.

[0085] In this embodiment, a binary classification function is used to determine whether the text features and image features match, thereby realizing pairing processing of text data and image data. By pairing text data and image data, the text and the image related to the text are matched together, so that when searching for data related to the question text in the database, the relevant image can be directly found according to the text semantics, or the relevant text can be directly found according to the image.

[0086] Combination Figure 4 As shown, another voice question-answering method provided by an embodiment of the present disclosure includes:

[0087] S401: The processor obtains original data.

[0088] S402: The processor filters the original data.

[0089] S403: The processor classifies and saves the filtered original data to obtain initial data.

[0090] S404, the processor extracts information from the acquired initial data to obtain extracted data, where the extracted data includes text data and image data.

[0091] S405: The processor rearranges the extracted data.

[0092] S406: The processor saves the rearranged extracted data and generates a database.

[0093] S407, the processor converts the collected user voice into question text.

[0094] S408, the processor searches the database for data related to the question text through multiple query modes, and the multiple query modes include using a pre-trained code generation model to search the database for data related to the question text.

[0095] S409, the processor inputs the data related to the question text into a pre-trained decision model to obtain target data as an answer.

[0096] The voice question-and-answer method provided by the embodiment of the present disclosure can obtain original data through one or more channels including a small program APP, a WeChat public account, a crawler website, and an operation and maintenance customer service. Exemplarily, the relevant materials of refrigerators of different models under Company A are obtained as original data through the small program APP, the WeChat public account, the crawler website, and the operation and maintenance customer service. The obtained original data is filtered, and the filtering content includes duplicate data, harmful data, garbled code, and URLs. Since the source of the original data is wide and the format is diverse, while maintaining the diversified original data, by filtering out unavailable data and removing redundant and duplicate information in the data, the burden of data storage and processing can be reduced, and the quality and reliability of the original data can be improved. The filtered original data is classified and saved by screening. Exemplarily, the original data is classified and saved according to different models of refrigerators. For the original data sets that cannot be distinguished, manual screening is performed and classified and saved. By classifying and saving the filtered original data, the initial data is obtained to facilitate the subsequent management and call of the initial data.

[0097] Combination Figure 5 As shown, another voice question-answering method provided by an embodiment of the present disclosure includes:

[0098] S501, the processor collects user voice.

[0099] S502: The processor cleans the user's voice.

[0100] S503, the processor converts the cleaned user voice into question text.

[0101] S504, the processor searches the database for data related to the question text through multiple query modes, and the multiple query modes include using a pre-trained code generation model to search the database for data related to the question text.

[0102] S505, the processor inputs the data related to the question text into a pre-trained decision model to obtain target data as an answer.

[0103] The embodiment of the present disclosure provides a voice question-answering method, which can collect user voice, clean the user voice, convert the cleaned user voice into question text, and then query data related to the question text and generate answers based on the question text, thereby realizing voice interaction. Among them, the collection of user voice can be performed in real time through a microphone, a mobile phone, a pickup or an associated home appliance. Cleaning the user voice includes one or more operations of noise reduction, echo removal, and reverberation removal on the user voice. By cleaning the user voice, the quality of the user voice is improved, and then the accuracy of converting the user voice into question text is improved. The cleaned user voice can be converted into question text through the Whisper model, and the user voice is converted into question text by using the Whisper model to further improve the accuracy of the conversion between the user voice and the question text.

[0104] Combination Figure 6 As shown, the embodiment of the present disclosure provides a voice question-answering device 60, including a conversion module 610, a query module 620 and an analysis module 630. The conversion module 610 is configured to convert the collected user voice into a question text; the query module 620 is configured to query the data related to the question text in the database through a variety of query modes, and the variety of query modes include querying the data related to the question text in the database using a pre-trained code generation model; the analysis module 630 is configured to input the data related to the question text into a pre-trained decision model to obtain the target data as an answer.

[0105] By using the voice question-answering device 600 provided in the embodiment of the present disclosure, it is possible to directly retrieve data related to the user's question in the database using the pre-trained code generation model, and then obtain the answer to the user's question through the relevant data and the pre-trained decision model. Compared with the related technology, it avoids the construction of the knowledge graph and does not need to update the knowledge graph frequently, so it is more robust.

[0106] Combination Figure 7As shown, the embodiment of the present disclosure provides a voice question-answering device 70, including a processor (processor) 700 and a memory (memory) 701. Optionally, the device 70 may also include a communication interface (CommunicationInterface) 702 and a bus 703. Among them, the processor 700, the communication interface 702, and the memory 701 can communicate with each other through the bus 703. The communication interface 702 can be used for information transmission. The processor 700 can call the logic instructions in the memory 701 to execute the voice question-answering method described in any of the above embodiments.

[0107] In addition, the logic instructions in the memory 701 described above may be implemented in the form of software functional units and when sold or used as independent products, may be stored in a computer-readable storage medium.

[0108] The memory 701 is a computer-readable storage medium that can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present disclosure. The processor 700 executes the functional application and data processing by running the program instructions / modules stored in the memory 701, that is, implementing the voice question-answering method described in any of the above embodiments.

[0109] The memory 701 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 701 may include a high-speed random access memory and may also include a non-volatile memory.

[0110] Combination Figure 8 As shown, the embodiment of the present disclosure provides a refrigerator 80, including: a product body, and the above-mentioned voice question and answer device 60 (70). The voice question and answer device 60 (70) is installed on the refrigerator body. The installation relationship described here is not limited to placement inside the refrigerator body, but also includes installation connections with other components of the refrigerator 80, including but not limited to physical connections, electrical connections or signal transmission connections. It can be understood by those skilled in the art that the voice question and answer device 60 (70) can be adapted to a feasible refrigerator body, thereby realizing other feasible embodiments.

[0111] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the voice question-and-answer method described in any of the above embodiments.

[0112] The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiment of the present disclosure. The aforementioned storage medium may be a non-transient storage medium, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes.

[0113] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible changes. Unless explicitly required, separate components and functions are optional, and the order of operation may vary. The parts and features of some embodiments may be included in or replace the parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of listings containing one or more associated ones. In addition, when used in the present application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the presence of other identical elements in the process, method or device comprising the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.

[0114] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods for each specific application to implement the described functions, but such implementations should not be considered to exceed the scope of the embodiments of the present disclosure. The technicians may clearly understand that, for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above may refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.

[0115] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.

[0116] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A voice question-answering method, applied to a refrigerator, characterized in that: Methods include: Convert the collected user voice into question text; Querying data related to the question text in the database through multiple query modes, wherein the multiple query modes include querying data related to the question text in the database using a pre-trained code generation model; Input the data related to the question text into the pre-trained decision model and obtain the target data as the answer.

2. The voice question-answering method according to claim 1, characterized in that: The steps of using the pre-trained code generation model to query the database for data related to the question text include: Use a pre-trained code generation model to convert question text into query language; According to the converted query language, relevant data is queried in the database.

3. The voice question-answering method according to claim 1 or 2, characterized in that: Multiple query modes also include: Use a pre-trained text vectorization model to query the database for data related to the question text; and / or Use a search engine to query the database for data related to the question text.

4. The voice question-answering method according to claim 1 or 2, characterized in that: The steps of inputting data related to the question text into the pre-trained decision model to obtain the target data as the answer include: Input the data related to the question text into the pre-trained decision model to obtain the data with the highest correlation with the question text as the target data; Use a large multimodal model to stylize the target data and generate answers.

5. The voice question-answering method according to claim 1 or 2, characterized in that: Also includes: Automatically sorting the training data according to the relevance of the training data to obtain a first sequence; According to the correlation of the training data, the first sequence is manually corrected to obtain the second sequence; According to the first sequence and the second sequence, the decision model is trained by a reinforcement learning algorithm.

6. The voice question-answering method according to claim 1 or 2, characterized in that: Also includes: Extracting information from the acquired initial data to obtain extracted data, where the extracted data includes text data and image data; Re-layout the extracted data; The extracted data after rearrangement is saved to generate a database.

7. A voice question-answering device, characterized in that: include: A conversion module, configured to convert the collected user voice into question text; A query module is configured to query the data related to the question text in the database through multiple query modes, wherein the multiple query modes include querying the data related to the question text in the database using a pre-trained code generation model; The analysis module is configured to input data related to the question text into a pre-trained decision model to obtain target data as an answer.

8. A voice question-answering device, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the voice question-and-answer method as described in any one of claims 1 to 6 when running the program instructions.

9. A refrigerator, characterized in that: include: Refrigerator body; The voice question and answer device as described in claim 7 or 8 is installed on the refrigerator body.

10. A computer-readable storage medium storing program instructions, characterized in that: When the program instructions are executed, the computer is used to execute the voice question-answering method according to any one of claims 1 to 6.