Voice dialogue retrieval method and device based on large model and medium
By adopting a large model-based voice dialogue search method in the database retrieval system, using the inverse index database and fuzzy matching algorithm, the problem of inflexible traditional database retrieval methods and high cold start cost of voice dialogue retrieval system is solved, and efficient, flexible and accurate retrieval effects are achieved.
Patent Information
- Application Number
- CN202411916744.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-24
AI Technical Summary
The traditional database search method is inflexible, the existing voice dialogue search system has high cold start cost, and there are limitations in semantic understanding and dialogue reply generation.
The speech dialogue search method based on the big model is adopted, and the database to be retrieved and its corresponding inverted index database is constructed, and the preset speech templates are used to interact with the user, combined with the speech recognition model, the user's search speech is converted into pinyin search information, and the fuzzy matching algorithm is used for searching and matching.
It improves search efficiency and accuracy, reduces cold start costs, enhances the flexibility and user experience of conversation reply, and realizes the flexibility and cost reduction of search methods.
Smart Images

Figure CN120011589A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of voice retrieval, and in particular to a large model-based voice dialogue retrieval method, device and medium. Background Art
[0002] In today's information society, databases are important tools for storing and managing large amounts of data. Their retrieval efficiency and convenience are directly related to the efficiency of data utilization. Traditional database retrieval methods mostly rely on text input. Users need to enter retrieval keywords or sentences through devices such as keyboards. This method is not flexible and efficient in some scenarios. Especially in scenarios such as mobile devices and smart homes, users prefer to complete database retrieval operations through voice interaction.
[0003] However, existing voice dialogue retrieval systems often have some problems. For example, some systems rely on a large amount of training data to build models, resulting in high cold start costs, and the generalization ability of the model may be insufficient for database retrieval tasks in specific fields. In addition, some systems have limitations in semantic understanding and have difficulty accurately understanding users' complex intentions and expressions, resulting in low accuracy and recall of retrieval results. Some systems also lack flexibility in the generation of dialogue responses and cannot be dynamically adjusted according to the actual needs of users and the status of the dialogue.
[0004] Therefore, a large model-based speech dialogue retrieval method is proposed to solve the above problems. Summary of the invention
[0005] The embodiments of the present application provide a large-model-based voice dialogue retrieval method, device, and medium to solve the following technical problems: the traditional database retrieval method is inflexible and the cold start cost of the existing voice dialogue retrieval system is high.
[0006] In the first aspect, an embodiment of the present application provides a voice dialogue retrieval method based on a large model, characterized in that the method includes: constructing a database to be retrieved and an inverted index database associated with the database to be retrieved; wherein, the database to be retrieved is a database including multiple text field values, and the inverted index database is a database that converts multiple text field values in the database to be retrieved into multiple pinyin field values; communicating with the user based on a preset speech template to collect the user's search voice; wherein the speech template can guide the user on how to ask questions and communicate with the user based on the field value; processing the search voice based on a preset speech recognition large model to convert the search voice into pinyin search information; wherein the pinyin search information is information including multiple pinyin field values; matching the pinyin search information retrieval and the inverted index database based on a preset fuzzy matching algorithm, and collecting pinyin field values with a matching degree greater than a preset matching degree threshold to obtain a first number of field values to be processed. ; Wherein, the value range of the matching degree is 0 to 1; determine whether the first number of field values to be processed is greater than the preset quantity threshold; if so, process the first number of field values to be processed based on the speech template to generate a screening voice, and collect the user's search voice based on the screening voice; if not, process the pinyin search information based on the speech template to generate voice information, and remind the user to input the search voice based on the voice information; if not, and the first number is 0, remind the user to re-enter the search voice based on the speech template; if yes, and there is a pinyin field value with a matching degree of 1, determine the pinyin field value as the search field value; search the database to be searched based on the search field value to obtain the search result; when the search result is unique, process the search result based on the speech template and reply to the user; when the search result is not unique, compare the field value to be processed with the search result to determine the missing field value, and process the missing field value based on the speech template to collect the user's search voice.
[0007] In one implementation of the present application, a database to be searched and an inverted index database associated with the database to be searched are constructed, specifically including: constructing the database to be searched; performing pinyin conversion on each text field value in the database to be searched to generate a corresponding pinyin field value; establishing an association relationship between the pinyin field value and the original text field value to form an inverted index database.
[0008] In one implementation of the present application, the search voice is processed based on a preset speech recognition large model to convert the search voice into pinyin search information, specifically including: processing the search voice based on the speech recognition large model to convert the search voice into pinyin search information; and / or processing the search voice based on a preset speech recognition algorithm to convert the search voice into search text; when the user confirms the search text, processing the search text information based on a preset information extraction algorithm to generate a field value to be queried; processing the field value to be queried based on a preset natural language processing algorithm to generate pinyin search information.
[0009] In one implementation of the present application, the pinyin retrieval information is retrieved and the inverted index database is matched based on a preset fuzzy matching algorithm, and the pinyin field values whose matching degree is greater than a preset matching degree threshold are collected to obtain a first number of field values to be processed, specifically including: calculating the matching degree between the pinyin retrieval information and each pinyin field value in the inverted index database based on the fuzzy matching algorithm; and taking the pinyin field values whose matching degree is greater than the preset matching degree threshold as the field values to be processed.
[0010] In one implementation of the present application, if yes, the first number of field values to be processed are processed based on the speech template to generate a screening voice, and the user's search voice is collected based on the screening voice, specifically including: determining the processing flow; processing the first number of field values to be processed according to the processing flow to generate a screening voice prompt containing the first number of field values to be processed; converting the screening voice prompt into a voice signal through speech synthesis technology, and playing it to the user; collecting the search voice given by the user according to the screening voice prompt.
[0011] In one implementation of the present application, if not, the pinyin retrieval information is processed based on the speech template to generate voice information, and the user is reminded to re-enter the retrieval voice based on the voice information, specifically including: determining the processing flow; when the first quantity is less than the quantity threshold, processing the field value to be retrieved based on the speech template to generate a confirmation voice prompt including the field value to be retrieved; converting the confirmation voice prompt into a voice signal through speech synthesis technology, and playing it to the user; collecting the retrieval voice re-entered by the user.
[0012] In one implementation of the present application, when the retrieval results are not unique, the field values to be processed are compared with the retrieval results to determine the missing field values, and the missing field values are processed based on the speech template to collect the user's retrieval voice, specifically including: when the retrieval results are not unique, the field values to be processed are compared with the field values in the retrieval results to determine the missing field values; an inquiry voice containing the missing field value is generated based on the speech template and played to the user; and the retrieval voice given by the user based on the inquiry voice is collected.
[0013] In one implementation of the present application, the method further includes: recording the user's search history during the search process to update the speech recognition large model.
[0014] In the second aspect, the embodiment of the present application also provides a voice dialogue retrieval device based on a large model, the device comprising: at least one processor; and, a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: construct a database to be retrieved and an inverted index database associated with the database to be retrieved; wherein the database to be retrieved is a database including multiple text field values, and the inverted index database is a database that converts multiple text field values in the database to be retrieved into multiple pinyin field values; communicate with the user based on a preset speech template to collect the user's search speech; wherein the speech template can guide the user on how to ask questions, and communicate with the user according to the field value; process the search speech based on the preset speech recognition large model to convert the search speech into pinyin search information; wherein the pinyin search information is information including multiple pinyin field values; match the pinyin based on a preset fuzzy matching algorithm Retrieve information retrieval and inverted index database, collect Pinyin field values with a matching degree greater than a preset matching degree threshold to obtain a first number of field values to be processed; wherein, the value range of the matching degree is 0 to 1; determine whether the first number of field values to be processed is greater than the preset number threshold; if so, process the first number of field values to be processed based on the speech template to generate a screening voice, and collect the user's search voice based on the screening voice; if so, and there is a Pinyin field value with a matching degree of 1, determine the Pinyin field value as the search field value; if not, process the Pinyin search information based on the speech template to generate voice information, and remind the user to re-enter the search voice based on the voice information; search the database to be searched based on the search field value to obtain the search result; when the search result is unique, process the search result based on the speech template, and reply to the user; when the search result is not unique, compare the field value to be processed with the search result to determine the missing field value, and process the missing field value based on the speech template to collect the user's search voice.
[0015] In the third aspect, the embodiment of the present application also provides a non-volatile computer storage medium for voice dialogue retrieval based on a large model, storing computer executable instructions, characterized in that the computer executable instructions are configured to: construct a database to be retrieved and an inverted index database associated with the database to be retrieved; wherein the database to be retrieved is a database including multiple text field values, and the inverted index database is a database that converts multiple text field values in the database to be retrieved into multiple pinyin field values; communicate with the user based on a preset speech template to collect the user's search voice; wherein the speech template can guide the user on how to ask questions, and communicate with the user according to the field value; process the search voice based on a preset speech recognition large model to convert the search voice into pinyin search information; wherein the pinyin search information is information including multiple pinyin field values; match the pinyin search information retrieval and the inverted index database based on a preset fuzzy matching algorithm, and collect the search voice with a large matching degree. a pinyin field value with a preset matching degree threshold to obtain a first number of field values to be processed; wherein the matching degree ranges from 0 to 1; it is determined whether the first number of field values to be processed is greater than the preset number threshold; if so, the first number of field values to be processed are processed based on the speech template to generate a screening voice, and the user's search voice is collected based on the screening voice; if so, and there is a pinyin field value with a matching degree of 1, the pinyin field value is determined to be the search field value; if not, the pinyin search information is processed based on the speech template to generate a voice message, and the user is reminded to re-enter the search voice based on the voice message; the database to be searched is searched based on the search field value to obtain the search result; when the search result is unique, the search result is processed based on the speech template, and the user is replied; when the search result is not unique, the field value to be processed is compared with the search result to determine the missing field value, and the missing field value is processed based on the speech template to collect the user's search voice.
[0016] The embodiments of the present application provide a method, device and medium for voice dialogue retrieval based on a large model, which at least include the following technical effects:
[0017] By constructing a database to be searched and its corresponding inverted index database, the rapid conversion of text field values and pinyin field values is achieved, greatly improving the search efficiency. Preset speech templates are used to interact with users, guide users to ask search questions, and convert users' search voice into pinyin search information based on the speech recognition model. Fuzzy matching algorithms are used to match in the inverted index database to collect field values that are highly relevant to the search voice. Based on the number and quality of matching results, users are intelligently guided to further filter or supplement search information to ensure the accuracy of search results. When the search result is unique, the user is directly replied; when the result is not unique, the missing field value is determined by comparison, and the user is guided to supplement it, so as to accurately meet the user's search needs and improve the user experience, thus achieving the effect of flexible search methods and low cold start costs to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0019] Figure 1 A flow chart of a voice dialogue retrieval method based on a large model provided in an embodiment of the present application;
[0020] Figure 2 A schematic diagram of the internal structure of a large-model-based voice dialogue retrieval device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0022] The embodiments of the present application provide a large-model-based voice dialogue retrieval method, device, and medium to solve the following technical problems: the traditional database retrieval method is inflexible and the cold start cost of the existing voice dialogue retrieval system is high.
[0023] The technical solution proposed in the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0024] Figure 1 A flow chart of voice dialogue retrieval based on a large model is provided in the embodiment of the present application. Figure 1As shown, a large model-based voice dialogue retrieval method provided in an embodiment of the present application specifically includes the following steps:
[0025] Step 100, construct a database to be searched and an inverted index database associated with the database to be searched; wherein the database to be searched is a database including multiple text field values, and the inverted index database is a database that converts multiple text field values in the database to be searched into multiple pinyin field values.
[0026] A to-be-searched database containing multiple text field values is constructed, and an inverted index database is constructed based on the database, wherein the inverted index database converts the text field values into pinyin field values.
[0027] Step 101: construct a database to be searched.
[0028] Collect and organize the raw data to be retrieved to form a database containing multiple text field values. The text field value can be any form of Chinese text, such as article titles, content summaries, keywords, etc.
[0029] For example, to build a database to be searched about ancient Chinese poetry, you can collect the titles and contents of the ancient poetry from various sources and store the titles and contents as text field values in the database.
[0030] That is, a record in the database contains the fields "title" and "content", which correspond to the title and full text of an ancient poem respectively.
[0031] Step 102: Perform pinyin conversion on each text field value in the database to be searched to generate a corresponding pinyin field value.
[0032] Use a pinyin conversion tool or algorithm to convert each text field value in the database to be searched into a corresponding pinyin field value. The pinyin field value is a string consisting of the pinyin corresponding to the Chinese character, which is used for subsequent search operations.
[0033] Referring to step 101, continuing to take the ancient poetry database as an example, for the title of an ancient poem "Quiet Night Thoughts", a pinyin conversion tool can be used to convert it into "jing4ye4si1". Similarly, the content of the ancient poem also needs to be converted into pinyin word by word.
[0034] Step 103: Establish an association relationship between the pinyin field value and the original text field value to form an inverted index database.
[0035] An association relationship between the pinyin field value and the original text field value is established in the inverted index database, so that when a user searches by pinyin, the corresponding text field value and the document in which it is located can be quickly found according to the pinyin field value.
[0036] In the inverted index of the ancient poetry database, we can create an index entry for each pinyin field value (such as "jing4ye4si1"), and associate this index entry with the original text field value (such as "Jing Ye Si") and the document containing this field value. When the user inputs the pinyin "jing4ye4si1" for retrieval, the ancient poem title and content corresponding to it can be immediately found.
[0037] Step 200: Communicate with the user based on a preset conversation template to collect the user's retrieval voice; among them, the conversation template can guide the user on how to ask questions and communicate with the user according to the field value.
[0038] The conversation template is a series of templates for questions or responses preset according to common retrieval scenarios and user needs, which can guide the user on how to ask questions and make it easier to understand the user's intention.
[0039] First, design a series of conversation templates according to the target application scenario and user needs. These templates can include an opening statement to guide the user to ask questions, a way of asking questions for specific field values, and possible responses from the user, etc. For example, in the book retrieval scenario, the conversation templates include "May I ask which book you want to search for?" and "Please tell me the author or title of the book".
[0040] After receiving the user's voice input or triggering an interaction, select a suitable conversation template to interact with the user according to the current context or the user's historical behavior. Convert the conversation template into a voice output through text-to-speech technology (TTS) to guide the user on how to ask questions.
[0041] After the user asks questions according to the guidance of the conversation template, use automatic speech recognition technology (ASR) to collect and recognize the user's retrieval voice.
[0042] For example, in an intelligent book retrieval system, the user can interact with the system through voice to find the books they are interested in.
[0043] Preset conversation template:
[0044] Opening statement: "Hello, welcome to use the intelligent book retrieval system. May I ask which book you want to search for?"
[0045] Guiding question: "You can tell me the author, title or keyword of the book."
[0046] Example of the user's response: "I want to find the book 'Dream of the Red Chamber'."
[0047] Communicate with the user:
[0048] System (output through TTS): "Hello, welcome to use the intelligent book retrieval system. May I ask which book you want to search for?"
[0049] User (input via voice recognition): "I want to find a book about Dream of the Red Chamber."
[0050] Step 300: Process the search speech based on a preset speech recognition model to convert the search speech into pinyin search information; wherein the pinyin search information is information including multiple pinyin field values.
[0051] Pinyin search information is composed of multiple Pinyin field values and is used for subsequent search operations.
[0052] Step 301: Process the search speech based on the speech recognition large model to convert the search speech into pinyin search information.
[0053] Use a large speech recognition model (such as a deep learning model) to recognize and process the user's search voice. The large model can use a large amount of training data and a complex network structure to more accurately recognize the phonemes, syllables, and words in the speech and convert them into phonetic representation.
[0054] Even if the user's pronunciation is different or ambiguous, the large model can make intelligent corrections through context and voice features to generate relatively accurate pinyin content, and extract keywords in the pinyin content according to actual needs and training data to generate pinyin retrieval information.
[0055] It is understandable that the speech recognition large model is trained based on the existing large model according to actual needs, and will not be elaborated again.
[0056] However, this processing mode is invisible to the user, and if an error occurs, it cannot be corrected in time. Therefore, steps 302 to 304 can also be adopted to accurately input the user's voice.
[0057] Step 302: Process the search speech based on a preset speech recognition algorithm to convert the search speech into search text.
[0058] The search speech is converted into search text through the traditional speech recognition algorithm, and the converted search text can be used for the subsequent information extraction and pinyin generation steps.
[0059] Step 303: When the user confirms the search text, the search text information is processed based on a preset information extraction algorithm to generate a field value to be queried.
[0060] The converted search text will be displayed to the user through the display. The user can modify the search text. After the user confirms the search text, the information extraction algorithm is used to extract key information from the search text to generate the field value to be queried.
[0061] For example, the value of the field to be queried can be a book title, author name, keyword, etc., which is used for subsequent search operations. This step ensures the accurate communication of the user's intention and improves the accuracy of the search.
[0062] Step 304: Calculate and process the field value to be queried based on the preset natural language processing to generate pinyin search information.
[0063] The natural language processing algorithm is used to further process the query field value and convert it into pinyin search information. The natural language processing algorithm can analyze the syntax, semantics and context information of the field value and generate the corresponding pinyin representation more accurately. In this way, even if the field value contains complex situations such as rare characters and polyphonic characters, accurate pinyin search information can be generated.
[0064] In one example, the user says the search voice: "I want to find the book Dream of the Red Chamber."
[0065] The search speech is processed using a large speech recognition model and directly converted into pinyin search information. Due to the intelligent correction capability of the large model, even if there are slight differences in the user's pronunciation, accurate pinyin search information can be generated: "hong2lou2meng4."
[0066] Alternatively, select and execute step 302 to step 304 .
[0067] Use speech recognition algorithms to convert search voice into search text: "I want to find the book Dream of the Red Chamber.
[0068] After the user confirms that the search text is correct, an information extraction algorithm is used to extract the field value to be queried from the search text: "Dream of Red Mansions".
[0069] Use natural language processing algorithm to convert the field value to be queried into pinyin retrieval information: "hong2lou2meng4".
[0070] Step 400, based on a preset fuzzy matching algorithm, match the pinyin retrieval information retrieval and the inverted index database, collect the pinyin field values whose matching degree is greater than a preset matching degree threshold to obtain a first number of field values to be processed; wherein the matching degree ranges from 0 to 1.
[0071] The fuzzy matching algorithm can consider the similarities and differences between pinyin and give more flexible matching results. The matching threshold is a value between 0 and 1, which is used to control the strictness of the match. The first number of field values to be processed refers to the number of pinyin field values that meet the matching threshold condition.
[0072] Step 401: Calculate the matching degree between the pinyin search information and each pinyin field value in the inverted index database based on a fuzzy matching algorithm.
[0073] Use the fuzzy matching algorithm to calculate the degree of match between the pinyin search information and each pinyin field value in the inverted index database. The fuzzy matching algorithm can use a variety of methods, such as edit distance, cosine similarity, and Jaccard similarity, to calculate the similarity between two pinyin strings. For each pinyin field value, the algorithm will give a matching value with the pinyin search information. The closer the value is to 1, the higher the matching degree.
[0074] Step 402: The pinyin field value whose matching degree is greater than a preset matching degree threshold is used as the field value to be processed.
[0075] According to the preset matching threshold, the pinyin field values that meet the conditions are filtered. The matching value of each pinyin field value is compared with the pinyin search information. If the matching degree is greater than the matching threshold, the pinyin field value is used as the field value to be processed. In this way, a group of pinyin field values with high similarity to the pinyin search information can be collected for subsequent search processing.
[0076] For example, referring to the above example, an inverted index database containing the pinyin of book titles is constructed.
[0077] The user speaks the search voice, and the embodiment of the present application converts it into pinyin search information: "hong2long2meng4".
[0078] The fuzzy matching algorithm is used to calculate the matching degree between the pinyin search information and each pinyin field value in the inverted index database. For the pinyin field value "hong2lou2meng4" in the database (corresponding to the correct pinyin of the book "Dream of Red Mansions"), the matching degree calculated by the algorithm is 0.85.
[0079] The preset matching degree threshold is 0.8. Therefore, for the pinyin field value "hong2lou2meng4" with a matching degree of 0.85, it is used as the field value to be processed.
[0080] Step 500: Determine whether a first number of field values to be processed is greater than a preset number threshold.
[0081] The search results can be further optimized by determining whether the first number of field values to be processed obtained by the fuzzy matching algorithm exceeds the preset number threshold. The number threshold is a preset integer value used to control the number range of search results to avoid too many or too few results leading to poor user experience.
[0082] If the first number is greater than the quantity threshold, it means that there are too many search results, which may contain a large amount of irrelevant or redundant information, or the search conditions are relatively broad and require further processing or screening; if the first number is less than or equal to the quantity threshold, it means that the number of search results is moderate and can be directly processed or displayed.
[0083] For example, an inverted index database containing the pinyin of book names is constructed, and the quantity threshold is set to 10.
[0084] Pinyin search information: "ke1huan4xiao3shuo1".
[0085] A fuzzy matching algorithm is used to calculate the matching degree between the pinyin search information and each pinyin field value in the inverted index database, and to obtain the field values to be processed that meet the matching degree threshold condition.
[0086] The number of field values to be processed, that is, the first number is 15, which is greater than the preset number threshold of 10.
[0087] Because the first quantity is greater than the quantity threshold, the system determines that there are too many search results and further processing or screening is required.
[0088] Step 600: If yes, the first number of field values to be processed are processed based on the speech template to generate a screening voice, and the user's search voice is collected based on the screening voice.
[0089] When the system determines that the first number of field values to be processed is greater than the preset number threshold, it means that there are too many initial search results and the user needs to be guided to perform a more precise search. At this time, the system will process the field values to be processed based on the preset speech template, generate a filtering voice prompt containing the filtering conditions, and play it to the user so that the user can give a more specific search voice according to the prompt.
[0090] Step 601: Determine the processing flow.
[0091] It can be understood that the processing flow is the process of executing step 600, which has determined which speech template to use.
[0092] Step 602: Process the first number of field values to be processed according to the processing flow, and generate a screening voice prompt containing the first number of field values to be processed.
[0093] According to the determined processing flow, the most representative field value is selected from the first number of field values to be processed as the screening condition. Then, the screening conditions are combined into a clear screening voice prompt using a preset speech template. It is understood that this prompt should contain enough information so that the user can perform a secondary search based on the prompt.
[0094] Step 603: Convert the screened voice prompt into a voice signal through speech synthesis technology and play it to the user.
[0095] By using speech synthesis technology, the screening voice prompt is converted into a voice signal and played to the user through a speaker or other audio output device. In this way, the user can hear a voice prompt containing the screening conditions and perform a secondary search based on the prompt.
[0096] Step 604: Collect the search voice given by the user according to the screening voice prompt.
[0097] After the user hears the screening voice prompt, wait for the user to give a second search voice. Collect the user's search voice through a microphone or other audio input device.
[0098] For example, referring to the above example, it is determined that the initial search result (ie, the first number of field values to be processed) is greater than a preset number threshold.
[0099] Determine the processing flow and select the speech template about screening in the speech template (please provide more specific screening information about xx).
[0100] The most representative author, publication year and rating information are selected from the initial search results, and a filtering voice prompt is generated using a speech template: "Which author's science fiction novel are you looking for? Or which publication year are you looking for? Or do you have any requirements for the rating of the book?"
[0101] The screening voice prompts are converted into voice signals through speech synthesis technology and played to the user.
[0102] After hearing the voice prompt for the selection, the user gave a second search voice: "I want to find science fiction novels written by Liu Cixin, published after 2010, and with a score higher than 8 points."
[0103] Collect the user's secondary search voice and perform subsequent processing and analysis to obtain more accurate search results.
[0104] Step 700: If not, process the pinyin search information based on the speech template to generate voice information, and prompt the user to input the search voice based on the voice information.
[0105] Step 701: Determine the processing flow.
[0106] Refer to step 601, which will not be described in detail again.
[0107] Step 702: When the first quantity is less than the quantity threshold, the value of the field to be retrieved is processed based on the speech template to generate a confirmation voice prompt including the value of the field to be retrieved.
[0108] When the number of search results is less than the threshold, it is highly likely that the first number of fields to be searched includes the search target required by the user, and a confirmation voice prompt is generated based on the speech template and pinyin search information. This prompt should contain the key information of the user's original search request and be expressed in a questioning or confirmation manner to guide the user to confirm.
[0109] For example, if the user's pinyin search information is "hong3lou3meng4, liu2lao3lao3" (Liu Laolao from Dream of the Red Chamber), and the searched chapters containing Liu Laolao in Dream of the Red Chamber include Chapters 6, 39, 40, 41 and 113, a confirmation voice prompt will be generated: "Are you looking for Chapter 6, when Liu Laolao came to Rongguo Mansion for the first time; or Chapters 39 to 41, when Liu Laolao visited Rongguo Mansion for the second time and visited the Grand View Garden; or Chapter 113, when the Jia family was ransacked and Liu Laolao visited Aunt Feng".
[0110] Step 703: Convert the confirmation voice prompt into a voice signal through speech synthesis technology and play it to the user.
[0111] The generated confirmation voice prompt is converted into a voice signal using speech synthesis technology and played to the user through a speaker or other audio output device.
[0112] Step 704: Collect the search voice re-entered by the user.
[0113] After the user hears the confirmation voice prompt, wait for the user to re-enter the search voice.
[0114] Step 800: If not, and the first number is 0, the user is reminded to re-enter the search voice based on the speech template.
[0115] When the first number is 0, that is, no matching items are found, it means that the user's search request may not be correctly understood or there is no relevant content in the database. At this time, a guiding voice prompt will be generated based on the preset speech template to remind the user to re-enter the search voice in order to obtain the required information more accurately.
[0116] Step 900: If yes, and there is a pinyin field value with a matching degree of 1, then determine the pinyin field value as the search field value.
[0117] When there is a pinyin field value with a matching degree of 1, it means that a field value that is completely consistent with the pinyin input by the user has been found. When such a field value exists, it can be directly determined as the user's search field value without further confirmation or asking the user. This can simplify the search process and improve search efficiency.
[0118] Step 1000: Search the database to be searched based on the search field value to obtain the search results.
[0119] After determining the search field value, search the database to be searched according to the field value. Find relevant information in the database according to the search field value and generate search results.
[0120] Step 1100: When the search result is unique, process the search result based on the speech template and reply to the user.
[0121] When the search result is unique, it means that the information that fully matches the user's search request has been found. At this time, based on the preset speech template, the search result is formatted and a clear and concise reply voice is generated, which is played to the user through the intelligent voice assistant. In this way, the user can directly obtain the required information without further screening or confirmation.
[0122] Step 1200: When the search result is not unique, compare the field value to be processed with the search result to determine the missing field value, and process the missing field value based on the speech template to collect the user's search voice.
[0123] When the search results are not unique, it means that the user's search request may be too broad or vague, causing the system to be unable to accurately determine the user's search intent. At this time, the system needs to compare the field values to be processed (that is, the field values converted from the search request entered by the user) with the field values in the search results to find the missing field values that can further clarify the user's intent. Then, the system will generate an inquiry voice containing these missing field values based on the preset speech template, and play it to the user to guide the user to supplement the information.
[0124] Step 1201: When the search result is not unique, compare the field value to be processed with the field value in the search result to determine the missing field value.
[0125] Compare the field values to be processed with the field values in the search results one by one to find those field values that are not explicitly mentioned in the user's search request but exist in the search results. These field values are usually key information that can further narrow the search scope and clarify the user's intention. For example, if the user searches for "science fiction novels" and there are multiple science fiction novels by different authors and different publishers in the search results, then "author" and "publisher" are missing field values.
[0126] Step 1202: Generate an inquiry voice containing the missing field value based on the speech template and play it to the user.
[0127] After determining the missing field values, a voice query containing these missing field values is generated based on the preset speech template. For example, "Which author wrote the science fiction novel you want to find? Or which publisher published it?" This voice query is then played to the user through a speaker or other audio output device.
[0128] Step 1203: Collect the search voice given by the user according to the inquiry voice.
[0129] After the user hears the inquiry voice, the system will wait for the user to give further search voice according to the prompt.
[0130] The present application also includes the following method:
[0131] During the retrieval process, the user's retrieval history is recorded to update the speech recognition model. This content is prior art and will not be described in detail here.
[0132] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, the embodiment of this application also provides a large model-based voice dialogue retrieval device, whose structure is as follows: Figure 2 shown.
[0133] Figure 2 The internal structure diagram of a large model-based voice dialogue retrieval device provided in the embodiment of the present application is shown in FIG. Figure 2 As shown, the device includes:
[0134] at least one processor 201;
[0135] and, a memory 202 communicatively connected to the at least one processor;
[0136] The memory 202 stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor 201 to enable at least one processor 201 to:
[0137] Construct a database to be searched and an inverted index database associated with the database to be searched; wherein, the database to be searched is a database including multiple text field values, and the inverted index database is a database that converts multiple text field values in the database to be searched into multiple pinyin field values; communicate with the user based on a preset speech template to collect the user's search voice; wherein, the speech template can guide the user on how to ask questions, and communicate with the user according to the field value; process the search voice based on a preset speech recognition large model to convert the search voice into pinyin search information; wherein the pinyin search information is information including multiple pinyin field values; match the pinyin search information retrieval and the inverted index database based on a preset fuzzy matching algorithm, collect the pinyin field values with a matching degree greater than a preset matching degree threshold, to obtain a first number of field values to be processed; wherein, matching The value range of the degree is 0 to 1; determine whether the first number of field values to be processed is greater than the preset quantity threshold; if so, process the first number of field values to be processed based on the speech template to generate a screening voice, and collect the user's search voice based on the screening voice; if so, and there is a Pinyin field value with a matching degree of 1, determine the Pinyin field value as the search field value; if not, process the Pinyin search information based on the speech template to generate voice information, and remind the user to re-enter the search voice based on the voice information; search the database to be searched based on the search field value to obtain the search result; when the search result is unique, process the search result based on the speech template and reply to the user; when the search result is not unique, compare the field value to be processed with the search result to determine the missing field value, and process the missing field value based on the speech template to collect the user's search voice.
[0138] Some embodiments of the present application provide corresponding Figure 1 A non-volatile computer storage medium for speech dialogue retrieval based on a large model, storing computer executable instructions, wherein the computer executable instructions are set to:
[0139] Construct a database to be searched and an inverted index database associated with the database to be searched; wherein, the database to be searched is a database including multiple text field values, and the inverted index database is a database that converts multiple text field values in the database to be searched into multiple pinyin field values; communicate with the user based on a preset speech template to collect the user's search voice; wherein, the speech template can guide the user on how to ask questions, and communicate with the user according to the field value; process the search voice based on a preset speech recognition large model to convert the search voice into pinyin search information; wherein the pinyin search information is information including multiple pinyin field values; match the pinyin search information retrieval and the inverted index database based on a preset fuzzy matching algorithm, collect the pinyin field values with a matching degree greater than a preset matching degree threshold, to obtain a first number of field values to be processed; wherein, matching The value range of the degree is 0 to 1; determine whether the first number of field values to be processed is greater than the preset quantity threshold; if so, process the first number of field values to be processed based on the speech template to generate a screening voice, and collect the user's search voice based on the screening voice; if so, and there is a Pinyin field value with a matching degree of 1, determine the Pinyin field value as the search field value; if not, process the Pinyin search information based on the speech template to generate voice information, and remind the user to re-enter the search voice based on the voice information; search the database to be searched based on the search field value to obtain the search result; when the search result is unique, process the search result based on the speech template and reply to the user; when the search result is not unique, compare the field value to be processed with the search result to determine the missing field value, and process the missing field value based on the speech template to collect the user's search voice.
[0140] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the IoT device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0141] The system and medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the system and medium also have similar beneficial technical effects to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the system and medium will not be repeated here.
[0142] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0143] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0144] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0146] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0147] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0148] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0149] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0150] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A speech dialogue retrieval method based on a large model, characterized in that: The method comprises: Constructing a database to be searched and an inverted index database associated with the database to be searched; wherein the database to be searched is a database including a plurality of text field values, and the inverted index database is a database that converts a plurality of text field values in the database to be searched into a plurality of pinyin field values; Communicate with the user based on a preset speech template to collect the user's search voice; wherein the speech template can guide the user on how to ask questions and communicate with the user according to the field value; Processing the search speech based on a preset speech recognition model to convert the search speech into pinyin search information; wherein the pinyin search information is information including multiple pinyin field values; Based on a preset fuzzy matching algorithm, the pinyin search information retrieval and the inverted index database are matched, and the pinyin field values whose matching degree is greater than a preset matching degree threshold are collected to obtain a first number of field values to be processed; wherein the matching degree ranges from 0 to 1; Determine whether the first number of to-be-processed field values is greater than a preset number threshold; If yes, the first number of to-be-processed field values are processed based on the speech template to generate a screening speech, and the user's search speech is collected based on the screening speech; If not, processing the pinyin search information based on the speech template to generate voice information, and prompting the user to input a search voice based on the voice information; If not, and the first number is 0, prompting the user to re-enter the search voice based on the speech template; If yes, and there is a pinyin field value with a matching degree of 1, then the pinyin field value is determined to be the search field value; Search the database to be searched based on the search field value to obtain search results; When the search result is unique, processing the search result based on the speech template and replying to the user; When the search result is not unique, the to-be-processed field value is compared with the search result to determine the missing field value, and the missing field value is processed based on the speech template to collect the search voice of the user.
2. A method for voice dialogue retrieval based on a large model according to claim 1, characterized in that: Constructing a database to be searched and an inverted index database associated with the database to be searched, specifically including: Construct a database to be searched; Performing pinyin conversion on each text field value in the database to be searched to generate a corresponding pinyin field value; An association relationship is established between the pinyin field value and the original text field value to form an inverted index database.
3. The method for voice dialogue retrieval based on a large model according to claim 1, characterized in that: Processing the search speech based on a preset speech recognition model to convert the search speech into pinyin search information specifically includes: Processing the search speech based on the speech recognition big model to convert the search speech into pinyin search information; and / or Processing the search speech based on a preset speech recognition algorithm to convert the search speech into search text; When the user confirms the search text, the search text information is processed based on a preset information extraction algorithm to generate a field value to be queried; The field value to be queried is calculated and processed based on the preset natural language processing to generate pinyin retrieval information.
4. A method for voice dialogue retrieval based on a large model according to claim 3, characterized in that: Matching the pinyin search information retrieval and the inverted index database based on a preset fuzzy matching algorithm, collecting pinyin field values whose matching degree is greater than a preset matching degree threshold, to obtain a first number of field values to be processed, specifically includes: Calculate the matching degree between the pinyin search information and each pinyin field value in the inverted index database based on the fuzzy matching algorithm; The pinyin field value whose matching degree is greater than the preset matching degree threshold is used as the field value to be processed.
5. The method for voice dialogue retrieval based on a large model according to claim 1, characterized in that: If yes, the first number of to-be-processed field values are processed based on the speech template to generate a screening voice, and the user's search voice is collected based on the screening voice, specifically including: Determine the processing flow; Processing the first number of to-be-processed field values according to the processing flow, and generating a screening voice prompt including the first number of to-be-processed field values; The screening voice prompts are converted into voice signals through speech synthesis technology and played to the user; Collect the search voice given by the user according to the screening voice prompt.
6. The method for voice dialogue retrieval based on a large model according to claim 1, characterized in that: If not, the pinyin search information is processed based on the speech template to generate voice information, and the user is prompted to re-enter the search voice based on the voice information, specifically including: Determine the processing flow; When the first number is less than the number threshold, processing the to-be-searched field value based on the speech template, and generating a confirmation voice prompt including the to-be-searched field value; The confirmation voice prompt is converted into a voice signal by using speech synthesis technology, and played to the user; The search voice re-entered by the user is collected.
7. The method for voice dialogue retrieval based on a large model according to claim 1, characterized in that: When the search result is not unique, comparing the field value to be processed with the search result to determine the missing field value, and processing the missing field value based on the speech template to collect the search voice of the user, specifically including: When the search results are not unique, compare the field value to be processed with the field value in the search results to determine the missing field value; Generate a voice query containing missing field values based on the speech template and play it to the user; Collect the search voice given by the user based on the query voice.
8. The method for voice dialogue retrieval based on a large model according to claim 1, characterized in that: The method further comprises: The user's search history is recorded during the search process to update the speech recognition model.
9. A speech dialogue retrieval device based on a large model, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Constructing a database to be searched and an inverted index database associated with the database to be searched; wherein the database to be searched is a database including a plurality of text field values, and the inverted index database is a database that converts a plurality of text field values in the database to be searched into a plurality of pinyin field values; Communicate with the user based on a preset speech template to collect the user's search voice; wherein the speech template can guide the user on how to ask questions and communicate with the user according to the field value; Processing the search speech based on a preset speech recognition model to convert the search speech into pinyin search information; wherein the pinyin search information is information including multiple pinyin field values; Based on a preset fuzzy matching algorithm, the pinyin search information retrieval and the inverted index database are matched, and the pinyin field values whose matching degree is greater than a preset matching degree threshold are collected to obtain a first number of field values to be processed; wherein the matching degree ranges from 0 to 1; Determine whether the first number of to-be-processed field values is greater than a preset number threshold; If yes, the first number of to-be-processed field values are processed based on the speech template to generate a screening speech, and the user's search speech is collected based on the screening speech; If yes, and there is a pinyin field value with a matching degree of 1, then the pinyin field value is determined to be the search field value; If not, processing the pinyin search information based on the speech template to generate voice information, and prompting the user to re-enter the search voice based on the voice information; Search the database to be searched based on the search field value to obtain search results; When the search result is unique, processing the search result based on the speech template and replying to the user; When the search result is not unique, the to-be-processed field value is compared with the search result to determine the missing field value, and the missing field value is processed based on the speech template to collect the search voice of the user.
10. A non-volatile computer storage medium for speech dialogue retrieval based on a large model, storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Constructing a database to be searched and an inverted index database associated with the database to be searched; wherein the database to be searched is a database including a plurality of text field values, and the inverted index database is a database that converts a plurality of text field values in the database to be searched into a plurality of pinyin field values; Communicate with the user based on a preset speech template to collect the user's search voice; wherein the speech template can guide the user on how to ask questions and communicate with the user according to the field value; Processing the search speech based on a preset speech recognition model to convert the search speech into pinyin search information; wherein the pinyin search information is information including multiple pinyin field values; Based on a preset fuzzy matching algorithm, the pinyin search information retrieval and the inverted index database are matched, and the pinyin field values whose matching degree is greater than a preset matching degree threshold are collected to obtain a first number of field values to be processed; wherein the matching degree ranges from 0 to 1; Determine whether the first number of to-be-processed field values is greater than a preset number threshold; If yes, the first number of to-be-processed field values are processed based on the speech template to generate a screening speech, and the user's search speech is collected based on the screening speech; If yes, and there is a pinyin field value with a matching degree of 1, then the pinyin field value is determined to be the search field value; If not, processing the pinyin search information based on the speech template to generate voice information, and prompting the user to re-enter the search voice based on the voice information; Search the database to be searched based on the search field value to obtain search results; When the search result is unique, processing the search result based on the speech template and replying to the user; When the search result is not unique, the to-be-processed field value is compared with the search result to determine the missing field value, and the missing field value is processed based on the speech template to collect the search voice of the user.
Citation Information
Patent Citations
Voice recognition method, device and equipment and readable storage medium
CN111554297A
Voice query method and device, computer equipment and storage medium
CN111611349A
Intelligent query method and device based on user voice, equipment and storage medium
CN115588430A
Speech recogniton method, apparatus, device and readable storage medium
US20210193143A1
Speech recognition method, device, and apparatus, and computer-readable storage medium
WO2020215554A1