Method, apparatus, storage medium and voice device for reducing voice response time

CN115294979BActive Publication Date: 2026-08-11QINGDAO HAIER AIR CONDITIONER GENERAL CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

但并没有考虑到有些问题的答案是固定的,而有些问题的答案是变化的

Benefits of technology

[0043]This application pre-stores the audio answers to user-asked questions in storage. Then, when a similar audio question is encountered, the corresponding audio answer is directly retrieved from storage, saving time spent searching for answers on the network and re-synthesizing the audio. Essentially, the solution adds a caching mechanism for some recurring questions and their audio answers, directly recording the audio answers to frequently asked questions, thereby reducing the audio response time for some questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115294979B_ABST
    Figure CN115294979B_ABST
Patent Text Reader

Abstract

This application relates to the field of smart home appliance technology. It discloses a method, apparatus, storage medium, and voice device for reducing voice response time. The method for reducing voice response time includes: recognizing the user's voice information to obtain the user's question; searching for the answer voice in the device's storage space; and outputting the answer voice to the user. By pre-storing the answer voices of previously asked questions in the storage space, when the same voice question is encountered, the corresponding answer voice is directly searched in the storage space. This saves the time spent searching for answers from the network side and resynthesizing the voice. Essentially, the solution reduces the voice response time for some questions by adding a caching mechanism for some repetitive questions and answer voices, thereby directly recording the answer voices of questions that the user may frequently ask.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart home appliance technology, such as a method, apparatus, storage medium, and voice device for reducing voice response time. Background Technology

[0002] Currently, voice devices can recognize users' voice commands. By parsing the voice commands, they determine the question the user wants to ask. The voice response process includes the entire process from receiving the user's voice question to outputting the answer to the user. However, in current voice recognition solutions, regardless of the user's question, the device first converts the speech into audio data and then searches for the corresponding answer. After finding the answer, it then uses TTS (Text to Speech) technology to convert the answer into speech and output it to the user. This process significantly increases the voice response time.

[0003] In related technologies, a data transmission method has been proposed to reduce voice response time. This method includes: acquiring a voice query request collected by a target device and determining the response text corresponding to the voice query request; acquiring multiple sub-audio data corresponding to the response text; in response to the voice query request, sequentially sending the multiple sub-audio data to the target device according to a sorting result, and instructing the target device to play the multiple sub-audio data sequentially according to the sorting result.

[0004] The existing technology reduces voice response time by acquiring multiple sub-audio data points from the response text and then sending these sub-audio data points sequentially to the target device according to their sorting order. However, it does not take into account that some questions have fixed answers while others have variable answers. Therefore, the response time for some questions can be reduced by classifying them. Summary of the Invention

[0005] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.

[0006] This disclosure provides a method, apparatus, storage medium, and voice device for reducing voice response time, which can reduce the response time for some problems.

[0007] In some embodiments, methods for reducing voice response time include:

[0008] The system identifies the user's voice information to determine the user's question.

[0009] Search for the answer to the question in the device's storage space using voice;

[0010] The answer will be output to the user via voice.

[0011] The types of problems include static problems and dynamic problems.

[0012] Optionally, the answer to the question can be found in the device's storage space via voice, including:

[0013] In the case of a static question, search for the audio answer in the first storage area;

[0014] In the case of a dynamic question, the answer to the question is searched for in the second storage area.

[0015] Among them, static questions are questions with fixed answers, while dynamic questions are questions with non-fixed answers or whose answers change over time.

[0016] The first storage area is the local storage area, used to store static questions and their corresponding audio answers.

[0017] The second storage area is a cache area, which is essentially also a storage area, but it is distinguished from the first storage area. It is used to store dynamic questions and their corresponding audio answers.

[0018] Optionally, the method for reducing voice response time further includes:

[0019] If the audio answer to the question is not found in the first or second storage area, the audio answer to the question is obtained from the network side.

[0020] Optionally, the answer to the question can be obtained via the network side, including:

[0021] Send the issue to the network side;

[0022] Receive answers to questions from the network side;

[0023] The answer to the question is processed by speech synthesis to obtain the audio version of the answer.

[0024] Optionally, after receiving the audio answer to the question, the following may also be included:

[0025] If the question is static, store the question and answer (in audio) in the first storage area;

[0026] If the question is dynamic, the question and answer (in audio) are stored in the second storage area.

[0027] Optionally, the method for reducing voice response time further includes:

[0028] Update the audio of the answers stored in the second storage area.

[0029] Optionally, the answer speech stored in the second storage area is updated, including:

[0030] If new question and answer audio is detected in the second storage area, obtain the time attribute of the question;

[0031] Based on the time attribute, the update cycle T of the problem is obtained according to the preset relationship;

[0032] The audio of the answer to the question is updated according to the update cycle T.

[0033] Alternatively, the answers stored in the second storage area can be updated in another way. This method specifically includes:

[0034] If new question and answer audio is detected in the second storage area, the validity period of the answer audio is determined.

[0035] If the audio answer to a user's question can be found in the second storage area, determine whether the audio answer is currently within its validity period.

[0036] If not, then send the issue to the network side;

[0037] Receive answers to questions from the network side and update them in the second storage area.

[0038] Optionally, if the storage capacity of questions and answer audio in the storage space has reached its limit and more needs to be added, delete the questions and corresponding answer audio that are called least frequently in the storage space.

[0039] In some embodiments, an apparatus for reducing voice response time includes a processor and a memory storing program instructions, the processor being configured to, when executing the program instructions, perform a method for reducing voice response time as described in any of the above embodiments.

[0040] In some embodiments, the voice device includes:

[0041] The apparatus for reducing voice response time as described in the above embodiments.

[0042] The method, apparatus, storage medium, and voice device for reducing voice response time provided in the embodiments of this disclosure can achieve the following technical effects:

[0043] This application pre-stores the audio answers to user-asked questions in storage. Then, when a similar audio question is encountered, the corresponding audio answer is directly retrieved from storage, saving time spent searching for answers on the network and re-synthesizing the audio. Essentially, the solution adds a caching mechanism for some recurring questions and their audio answers, directly recording the audio answers to frequently asked questions, thereby reducing the audio response time for some questions.

[0044] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description

[0045] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein:

[0046] Figure 1 This is a schematic diagram of a method for reducing voice response time provided in an embodiment of this disclosure;

[0047] Figure 2 This is a schematic diagram of another method for reducing voice response time provided in an embodiment of this disclosure;

[0048] Figure 3 This is a schematic diagram of a method for storing the voice of an answer obtained from the network side, provided by an embodiment of this disclosure;

[0049] Figure 4 This is a schematic diagram of a method for updating the voice of the answer in the second storage area according to an embodiment of this disclosure;

[0050] Figure 5 This is a schematic diagram of another method for updating the answer voice in the second storage area provided by an embodiment of this disclosure;

[0051] Figure 6 This is a schematic diagram of a device for reducing voice response time provided in an embodiment of this disclosure. Detailed Implementation

[0052] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.

[0053] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0054] Unless otherwise stated, the term "multiple" means two or more.

[0055] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0056] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0057] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.

[0058] Currently, voice devices can recognize users' voice commands. By parsing the voice commands, they determine the question the user wants to ask. The voice response process includes the entire process from receiving the user's voice question to outputting the answer to the user. However, in current voice recognition solutions, regardless of the user's question, the device first converts the voice into audio data and then searches for the corresponding answer. After finding the answer, it is then processed using TTS (Text to Speech) technology to output the answer to the user. This process significantly increases the voice response time. To address this issue, related technologies employ methods such as segmenting the audio into blocks and translating and transmitting them simultaneously, instead of waiting for the entire audio to be translated before sending it as a whole, thus reducing the voice response time.

[0059] This application considers that among the voice questions asked by users, some are frequently asked with fixed answers, while others have answers that change over time. Therefore, it introduces a caching mechanism for both question and answer voice recordings. Based on the question type, the storage space is divided into two areas: one area stores questions with fixed answers, directly storing the voice recordings of these answers. This reduces the time spent on multiple voice synthesis processes for these types of answers. The other area stores questions with variable answers. Although the voice recordings of these questions are also stored, they are not always valid. Therefore, the voice recordings of these questions will be updated subsequently. Updates can be done automatically on a regular basis or only when the question is asked again by the user. By adding a caching mechanism for some recurring questions and answer voice recordings, the voice recordings of answers to questions that users are likely to ask frequently are directly recorded, thereby reducing the voice response time for some questions.

[0060] These voice devices all include voice means for implementing control logic. For example, the means for reducing voice response time described in this application includes a processor and a memory. The processor, when executed, can implement a method for reducing voice response time.

[0061] The following is combined Figure 1 This application provides a method for reducing voice response time. The method for reducing voice response time includes:

[0062] S101, the processor recognizes the user's voice information and obtains the user's question.

[0063] S102, the processor searches for the answer to the question in the device's storage space via voice.

[0064] S103, the processor outputs the answer to the user in voice.

[0065] In this embodiment, the audio answers to user-asked questions are pre-stored in storage. When a similar audio question is encountered, the processor can directly retrieve the corresponding audio answer from storage, saving time spent searching for the answer from the network and re-synthesizing the audio. Essentially, the solution adds a caching mechanism for repetitive questions and their answers, directly recording the audio answers to frequently asked questions, thereby reducing the audio response time for some questions.

[0066] Optionally, the question type includes static questions and dynamic questions. The step of searching for the audio answer to the question in the device's storage space specifically includes: if the question is static, the processor searches for the audio answer in the first storage area; if the question is dynamic, the processor searches for the audio answer in the second storage area.

[0067] Static questions have fixed answers, while dynamic questions have non-fixed answers or answers that change over time. The first storage area is the local storage area, used to store static questions and their corresponding audio answers. The second storage area is a cache area, essentially also a storage area, but distinct from the first storage area, used to store dynamic questions and their corresponding audio answers.

[0068] In this embodiment, the storage space is divided into two areas based on the question type. One area stores questions with fixed answers, such as "What is the city flower of Shenyang?", directly storing the audio answer to this question. This reduces the time spent on multiple speech synthesis processes for such answers. The other area stores questions with variable answers, such as "How's the weather today?" We know that the weather changes daily, so the corresponding answer needs to be updated every day, requiring a search on the network. However, if a user asked the same question today, and another user asks the same question again today, the audio answer stored in the second storage area can be directly invoked, further reducing the time spent searching and synthesizing audio.

[0069] Combining the above solutions, such as Figure 2 As shown, another method for reducing voice response time is provided. This includes:

[0070] S201, the processor recognizes the user's voice information and obtains the user's question.

[0071] S202, the processor determines whether the problem is a static problem. If yes, proceed to S2031; otherwise, proceed to S2032.

[0072] S2031, the processor searches for the answer to the question in the first memory area via voice.

[0073] S2032, the processor searches for the answer to the question in the second memory area via voice.

[0074] S204, the processor outputs the answer to the user in voice.

[0075] This embodiment describes the process of retrieving the answer audio from storage space. Static questions are those with fixed answers, while dynamic questions are those with variable answers or answers that change over time. The first storage area is the local storage area, used to store static questions and their corresponding answer audio. The second storage area is a cache area, used to store dynamic questions and their corresponding answer audio. The second storage area is essentially also a storage area, but it is distinguished from the first storage area. After obtaining the user's audio question, the processor determines whether the question is static or dynamic. If the question is static, it searches for the corresponding answer audio in the first storage area; if the question is dynamic, it searches for the corresponding answer audio in the second storage area, thus saving time spent searching and synthesizing audio multiple times via the network.

[0076] Optionally, if the answer audio is not found in either the first or second storage area, the answer audio is obtained from the network side. The specific steps include: the processor sending the question to the network side; the processor receiving the answer from the network side; the processor performing speech synthesis on the answer to obtain the answer audio; and after obtaining the answer audio, if the question is static, the processor stores the question and answer audio in the first storage area; if the question is dynamic, the processor stores the question and answer audio in the second storage area.

[0077] like Figure 3 As shown, the method for searching for the answer voice message in the first and second storage areas before storing it in the network is explained. This method includes:

[0078] S301, the processor recognizes the user's voice information and obtains the user's question.

[0079] S302, the processor searches for the answer to the question in the device's storage space via voice.

[0080] S303: If the processor cannot find the answer voice for the question in the storage space, it obtains the answer voice for the question from the network side.

[0081] S304, if the processor does not find the answer voice for the question in the first or second storage area, it obtains the answer voice for the question from the network side.

[0082] S305, the processor determines whether the problem is a static problem. If yes, proceed to S3061; otherwise, proceed to S3062.

[0083] S3061, the processor stores the question and answer voice in the first memory area.

[0084] S3062, the processor stores the question and answer voice in the second memory area.

[0085] This embodiment describes the scenario where no answer audio is found in the storage space. This is essentially the case when a user asks a question for the first time. When a user asks a question for the first time, the corresponding answer audio is not stored in the storage space. Therefore, the answer is obtained from the network side. However, after obtaining the answer, the processor determines whether the question is static or dynamic based on its type. It then stores the question and answer in the corresponding storage area. This minimizes the audio response time when the user asks the same question a second time.

[0086] It's worth noting that since the answers in the first storage area generally don't change, you can pre-set the audio answers to some frequently asked questions when the product or device leaves the factory. Similarly, you can also pre-set the audio answers to some questions in the second storage area. If it wasn't set at the factory, the audio will be automatically generated via... Figure 3 The method described above adds the question and answer voice to either the first or second storage area. When a user asks the same question a second time, the time spent searching online and synthesizing speech is minimized, thus reducing the voice response time as much as possible.

[0087] In the above embodiment, since the answers to the dynamic questions in the second storage area are not completely constant, the spoken answers in the second storage area also need to be updated. Figure 4 As shown, a method for updating the answer speech stored in the second storage area is provided, including:

[0088] S401, when the processor detects that there are new questions and answer voices in the second storage area, it obtains the time attribute of the question.

[0089] S402, the processor obtains the problem update cycle T according to the time attribute and a preset relationship.

[0090] S403, the processor updates the audio of the answer to the question according to the update cycle T.

[0091] In this embodiment, when new questions and answer voices are added to the second storage area, the time attribute of the question is obtained and recorded in the memory. For example, when the new question is "What's the weather like?", the time attribute T can be marked as 12 hours or 24 hours, and the answer is updated every 12 or 24 hours. When the new question is "What's the stock market like?", a corresponding update time period can be set for the question. The last moment of the stock's trading period is used as the update time point, and no update is performed at other times. The answer is then obtained and converted into voice to replace the previous answer. In this way, basically any question asked by a user can have its voice response time greatly reduced for the second time.

[0092] Optionally, if the storage capacity of questions and answer audio in the storage space has reached its limit and more needs to be added, delete the questions and corresponding answer audio that are called least frequently in the storage space.

[0093] In this embodiment, when the first or second storage area is full and new content is added, the least frequently asked questions and their corresponding audio answers are deleted first. That is, questions that users don't often ask are deleted, thus freeing up more space. This is beneficial for storing other questions and audio answers.

[0094] The above methods can minimize voice response time; however, periodic updates for these dynamic issues will consume significant network resources. For example... Figure 5 As shown, another method for updating the answer speech stored in the second storage area is provided, including:

[0095] S501 determines the validity period of the answer voice when new question and answer voice are detected in the second storage area.

[0096] S502, if the answer voice of the user's question is stored in the second storage area, determine whether the answer voice is currently within its validity period.

[0097] S503: If the answer audio is not valid at the current time, obtain the answer audio from the network side and update the stored answer audio.

[0098] This embodiment proposes another method for updating the audio answers in the second storage area. When new questions and audio answers are added to the second storage area, the validity period of the audio answer is determined. Using weather as an example, the validity period can be set to 12 hours or 24 hours. When a user asks the same question again, it is first determined whether the current time is within the validity period. If the validity period has expired, the audio answer to the question needs to be retrieved from the network side and updated again. This method, where updates are only performed when the user asks the question again, consumes fewer network resources.

[0099] Optionally, if the storage capacity of questions and answer audio in the storage space has reached its limit and more needs to be added, delete the questions and corresponding answer audio that are called least frequently in the storage space.

[0100] In this embodiment, when the first or second storage area is full and new content is added, the least frequently asked questions and their corresponding audio answers are deleted first. That is, questions that users don't often ask are deleted, thus freeing up more space. This is beneficial for storing other questions and audio answers.

[0101] Combination Figure 6 As shown, this disclosure provides an apparatus for reducing voice response time, including a processor 600 and a memory 601. Optionally, the apparatus may further include a communication interface 602 and a bus 603. The processor 600, communication interface 602, and memory 601 can communicate with each other via the bus 603. The communication interface 602 can be used for information transmission. The processor 600 can call logical instructions in the memory 601 to execute the method for reducing voice response time described in the above embodiment.

[0102] Furthermore, the logic instructions in the aforementioned memory 601 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0103] The memory 601, as a storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 600 executes functional applications and data processing by running the program instructions / modules stored in the memory 601, thereby implementing the method for reducing voice response time in the above embodiments.

[0104] The memory 601 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 601 may include high-speed random access memory and may also include non-volatile memory.

[0105] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0106] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A method for reducing speech response time, characterized in that, include: The system recognizes the user's voice information to obtain the user's questions; the types of questions include static questions and dynamic questions; static questions are questions with fixed answers, while dynamic questions are questions with non-fixed answers or answers that change over time. If the question is a static question, search for the audio answer to the question in the first storage area; If the question is a dynamic question, the audio answer to the question is retrieved from the second storage area. If the audio answer to the question is not found in the first or second storage area, the audio answer to the question is obtained through the network side. If the question is a static question, the question and the answer (in audio) are stored in the first storage area; If the question is a dynamic question, the question and answer (in audio) will be stored in the second storage area; The answer will be output to the user via voice.

2. The method according to claim 1, characterized in that, The first storage area is the local storage area, used to store static questions and their corresponding audio answers; the second storage area is the cache area, used to store dynamic questions and their corresponding audio answers.

3. The method according to claim 1, characterized in that, The audio answer to the question obtained via the network side includes: Send the question to the network side; Receive the answer to the question from the network side; The answer to the question is processed by speech synthesis to obtain the speech of the answer to the question.

4. The method according to claim 1, characterized in that, The first storage area pre-sets voice answers to frequently asked questions when the product or device leaves the factory.

5. The method according to any one of claims 1 to 4, characterized in that, Also includes: Update the audio of the answer stored in the second storage area.

6. The method according to claim 5, characterized in that, Updating the audio of the answer stored in the second storage area includes: If new question and answer audio are detected in the second storage area, the time attribute of the question is obtained; Based on the time attribute, the update period T of the problem is obtained according to a preset relationship; The audio of the answer to the question is updated according to the update cycle T.

7. An apparatus for reducing voice response time, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to, when executing the program instructions, perform the method for reducing voice response time as described in any one of claims 1 to 6.

8. A storage medium storing program instructions, characterized in that, When the program instructions are executed, they perform the method for reducing voice response time as described in any one of claims 1 to 6.

9. A voice device, characterized in that, include: The apparatus for reducing voice response time as described in claim 7.

Citation Information

Patent Citations

  • Voice equipment, voice interaction method and device thereof and storage medium

    CN109960754A

  • Answer extraction method and input method for intelligent voice questions and answers and intelligent equipment

    CN111274360A