Intelligent voice interaction methods, systems, set-top boxes, and storage media
By combining microphone arrays and speech recognition models, the location of sound sources can be accurately determined and flexible operation commands can be generated, solving the problems of low speech recognition accuracy and cumbersome operation in noisy environments, and realizing the high efficiency and convenience of intelligent voice interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-09
- Publication Date
- 2026-06-30
AI Technical Summary
In noisy environments, existing voice interaction systems struggle to adapt to complex noise, resulting in low voice recognition accuracy and an inability to respond flexibly to user operating habits, leading to cumbersome operation and reduced willingness to use them.
A microphone array is used to collect voice information. The voice information in the target direction is enhanced by calculating the user's azimuth in real time. Combined with voice recognition and prediction models, the frequency and correlation of user operations are analyzed to generate flexible operation commands.
Improve voice recognition success rate in noisy environments, ensure smooth voice interaction, optimize device operation mode according to user habits, simplify operation process, and enhance ease of use and willingness to use.
Smart Images

Figure CN122313974A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent voice interaction technology, and in particular to intelligent voice interaction methods, systems, set-top boxes, and storage media. Background Technology
[0002] With the development of artificial intelligence technology, voice recognition and natural language processing technologies are increasingly widely used in the field of smart homes, providing users with a convenient voice control experience. The smart home ecosystem is gradually improving, and various smart devices are interconnected. Voice control has become a key link connecting these devices. Users can easily control various devices in their homes, such as lights, curtains, TVs, and robot vacuums, through simple voice commands. For example, on their way home from get off work, users can use a smart home application on their mobile phones to turn on the air conditioner and water heater in advance by voice, so that they can enjoy a comfortable environment and hot water as soon as they arrive home.
[0003] Currently, during voice interaction, fixed-directional microphone arrays are difficult to adapt to complex and varied noise environments. In noisy environments, they cannot dynamically adjust the beam direction according to the noise distribution, resulting in a large amount of environmental noise being collected, drowning out the voice signal and reducing the accuracy of voice recognition. At the same time, existing systems cannot sense the correlation with user operating habits, and the device can only operate according to a preset fixed mode. It cannot respond flexibly according to the actual user scenario, making the operation cumbersome and reducing the willingness to use it.
[0004] Therefore, we propose an intelligent voice interaction method, system, set-top box, and storage medium to solve the above problems. Summary of the Invention
[0005] The main objective of this invention is to provide an intelligent voice interaction method, system, set-top box, and storage medium to solve the problems mentioned in the background.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is: an intelligent voice interaction system, the system comprising a voice receiving end, a voice parsing end, and a voice prediction end; The voice receiver is used to collect voice information through a microphone array, and to calculate the microphone array output signal based on the user's azimuth angle in real time, thereby enhancing the voice information in the target direction, and transmitting the collected voice information to the voice parsing end. The voice parsing terminal converts speech into text using speech recognition technology, and uses a language training model to parse the text, extract user intent and semantic slots; and generates operation instructions based on user intent and semantic slots, and sends the operation instructions to the set-top box; The voice prediction terminal collects the number of times the user operates the device through a data acquisition device, calculates the operation frequency, analyzes the correlation between different user operations, calculates the support and confidence thresholds between different operations, and filters out strongly correlated operation combinations.
[0007] Preferably, the voice receiving end includes a collection unit, a positioning unit, an execution unit, and a transmission unit; The collection unit is used to collect voice information emitted by the user through a microphone array and to measure the frequency of the sound signal through a frequency sensor; The positioning unit is used to calculate the user's azimuth angle by the phase difference generated when sound reaches different microphones. The calculation formula is as follows: ; in, Represents the user's azimuth angle. This represents the spacing between different microphones. Represents the phase difference between different microphones. Represents the speed of sound. Represents frequency.
[0008] Preferably, the execution unit is used to calculate the output signal and weighting coefficients of the microphone array, and the calculation steps are as follows: Step 1: Calculate the output signal of the microphone array. The calculation formula is as follows: ; in, This represents the output signal of the microphone array at time t, where M represents the number of microphones in the array. Represents an index variable. Representing the The signal received by each element The corresponding weighting coefficients, Representing the The signal received by each array element at time t; Step 2, calculate the first The weighting coefficients corresponding to each array element are calculated using the following formula: ; in, Representing the The weighting coefficients corresponding to each array element Represents the imaginary unit. Represents the sound wave length, This represents the spacing between adjacent elements in a uniform linear array. Represents the user's azimuth angle. This represents the position index of the corresponding array element in the array; The transmission unit is used to transmit the collected information to the voice analysis terminal via the Wi-Fi module.
[0009] Preferably, the speech parsing terminal includes a conversion unit, a parsing unit, a generation unit, and a transmission unit; The conversion unit is used to convert speech into text using an HMM model; The parsing unit is used to parse text using the BERT model to extract user intent and semantic slots; The generation unit is used to generate operation instructions based on user intent and semantic slots using the BART model; The transmission unit is used to transmit instructions to the set-top box via a Wi-Fi module.
[0010] Preferably, the speech prediction terminal includes a frequency unit, an association unit, and a judgment unit; The frequency unit is used to collect the number of device operations and calculate the operation frequency. The calculation formula is as follows: ; in, Represents the operating frequency. Represents a time period. Represents time period The number of times a certain operation is performed.
[0011] Preferably, the association unit is used to calculate the probability of operation A and operation B occurring simultaneously, and the probability of operation B occurring when operation A occurs. The calculation steps are as follows: S1: Calculate the probability that operations A and B occur simultaneously, using the following formula: ; in, This represents the probability that operation A and operation B occur simultaneously. This represents the number of times operation A and operation B occur simultaneously. Represents the total number of operands; S2: Calculate the probability of operation B occurring when operation A occurs, using the following formula: ; in, This represents the probability that operation B will occur given that operation A has occurred. This represents the number of times operation A and operation B occur simultaneously. This represents the number of times operation A occurs.
[0012] Preferably, the comparison unit is used to set a probability threshold for the simultaneous occurrence of operation A and operation B. The probability threshold of operation B occurring when operation A occurs. ,when Greater than or equal to and Greater than When this occurs, it means that operation A and operation B are strongly correlated.
[0013] A smart voice interaction method includes the following steps: (1): Enter the voice receiving end, collect the voice information emitted by the user through the microphone array, measure the frequency of the sound signal through the frequency sensor, calculate the user's azimuth angle through the phase difference generated when the sound reaches different microphones, as well as the output signal and weighting coefficient of the microphone array, and transmit the collected information to the voice parsing end through the Wi-Fi module. (2): Enter the speech parsing end, which is used to convert speech into text through the HMM model, and use the BERT model to parse the text, extract user intent and semantic slots, and use the BART model to generate operation instructions based on user intent and semantic slots. Then the Wi-Fi module sends the operation instructions to the set-top box. (3): Enter the voice prediction terminal to collect the number of device operations, calculate the operation frequency, calculate the probability of operation A and operation B occurring simultaneously, and the probability of operation B occurring when operation A occurs, and compare it with the set threshold to determine the strongly correlated operation combination.
[0014] A set-top box, characterized in that it includes a memory for storing computer programs; A processor connected to the memory is used to execute the computer program stored in the memory.
[0015] A smart voice-interactive storage medium, wherein a computer-readable storage medium stores a computer program.
[0016] The present invention has the following beneficial effects: 1. In this invention, the collection unit collects voice information through a microphone array and measures the frequency of the sound signal with a frequency sensor, enabling more comprehensive voice acquisition and providing high-quality raw data for subsequent voice analysis. It also calculates the user's azimuth angle by utilizing the phase difference generated when the sound reaches different microphones, which can accurately determine the location of the sound source. Furthermore, it calculates the output signal of the microphone array based on the user's azimuth angle, enhancing the voice information in the target direction and suppressing noise and interference signals from other directions. Even in noisy environments, it can highlight the user's voice, improve the success rate of voice recognition, and ensure smooth voice interaction.
[0017] 2. In this invention, by modeling the collected voice information and calculating the probability of state transitions and output observations, a stable conversion from voice to text is achieved, ensuring the accuracy and reliability of text conversion. At the same time, considering the contextual information of the text, a deep understanding of semantics is achieved, which can accurately grasp the complex semantic relationships in the text and the user's true intentions, extract key information, and provide an accurate basis for the generation of subsequent operation instructions. In addition, by combining a bidirectional encoder and an autoregressive decoder, accurate operation instructions can be flexibly generated based on the input semantic information, thereby improving the level of intelligence of voice interaction.
[0018] 3. In this invention, by collecting the number of device operations and calculating the operation frequency, the frequency of user operations on the device can be quantified with specific values. At the same time, the probability of operation A and operation B occurring simultaneously and the probability of operation B occurring when operation A occurs can be calculated. This allows for in-depth exploration of the potential correlation between different operations. Strongly correlated operation combinations can be screened out through threshold comparison. Based on these strongly correlated combinations, the system can optimize the operating mode according to user operating habits, simplify the operation process, improve ease of use, and enhance user willingness to use the system. Attached Figure Description
[0019] Figure 1 This is a system flowchart of the intelligent voice interaction method, system, set-top box, and storage medium of the present invention; Figure 2 This is a flowchart illustrating the intelligent voice interaction method, system, set-top box, and storage medium of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Implementation 1: Please refer to Figure 1 and Figure 2 The present invention provides a technical solution: an intelligent voice interaction system, the system including a voice receiving end, a voice parsing end and a voice prediction end; The voice receiver is used to collect voice information through a microphone array, and to calculate the output signal of the microphone array based on the user's azimuth angle in real time, thereby enhancing the voice information in the target direction, and transmitting the collected voice information to the voice parsing end. The voice parsing end uses speech recognition technology to convert speech into text, and uses a language training model to parse the text, extract user intent and semantic slots; and generates operation instructions based on user intent and semantic slots, and sends the operation instructions to the set-top box; The voice prediction end collects the number of times users operate the device through a data acquisition device and calculates the operation frequency. At the same time, it analyzes the correlation between different user operations, calculates the support and confidence thresholds between different operations, and filters out strongly correlated operation combinations.
[0022] The voice receiver includes a collection unit, a positioning unit, an execution unit, and a transmission unit; The collection unit is used to collect voice information emitted by the user through a microphone array and to measure the frequency of the sound signal through a frequency sensor; The positioning unit is used to calculate the user's azimuth angle by measuring the phase difference generated when sound reaches different microphones. The calculation formula is as follows: ; in, Represents the user's azimuth angle. This represents the spacing between different microphones. Represents the phase difference between different microphones. Represents the speed of sound. Represents frequency.
[0023] The execution unit is used to calculate the output signal and weighting coefficients of the microphone array. The calculation steps are as follows: Step 1: Calculate the output signal of the microphone array. The calculation formula is as follows: ; in, This represents the output signal of the microphone array at time t, where M represents the number of microphones in the array. Represents an index variable. Representing the The signal received by each element The corresponding weighting coefficients, Representing the The signal received by each array element at time t; Step 2, calculate the first The weighting coefficients corresponding to each array element are calculated using the following formula: ; in, Representing the The weighting coefficients corresponding to each array element Represents the imaginary unit. Represents the sound wave length, This represents the spacing between adjacent elements in a uniform linear array. Represents the user's azimuth angle. This represents the position index of the corresponding array element in the array; The transmission unit is used to transmit the collected information to the voice analysis terminal via the Wi-Fi module.
[0024] In this embodiment, the collection unit collects voice information through a microphone array and measures the frequency of the sound signal with a frequency sensor, enabling more comprehensive voice acquisition and providing high-quality raw data for subsequent voice analysis. It also calculates the user's azimuth angle by utilizing the phase difference generated when the sound reaches different microphones, which can accurately determine the location of the sound source. Furthermore, it calculates the output signal of the microphone array based on the user's azimuth angle, enhancing the voice information in the target direction and suppressing noise and interference signals from other directions. Even in noisy environments, it can highlight the user's voice, improve the success rate of voice recognition, and ensure smooth voice interaction.
[0025] Implementation 2: Please refer to Figure 1 and Figure 2 The present invention provides a technical solution: based on the first embodiment, the speech parsing end includes a conversion unit, a parsing unit, a generation unit and a transmission unit; The conversion unit is used to convert speech into text using an HMM model; For example, in a smart home control scenario, when a user issues the voice command "turn on the living room lights," the voice receiver collects the voice information and transmits it to the voice parsing unit. The conversion unit then starts working, for example, separating the voice signals of different syllables such as "turn on," "living room," "of," and "lights." Then, it calculates the probability of each syllable corresponding to different texts using a pre-trained model. Combining this with the context information of the voice, it finally determines the most likely text content, "turn on the living room lights."
[0026] The parsing unit is used to parse text using the BERT model to extract user intent and semantic slots; For example, the BERT model understands each word in the sentence "turn on the living room light" by placing it in the context of the whole sentence. Through pre-training on a large amount of text, BERT can recognize that "turn on" is an action intention, "living room" is a location semantic slot, and "light" is a target device semantic slot.
[0027] The generation unit is used to generate operation instructions based on user intent and semantic slots using the BART model; Specifically, the BART model combines information from the parsing unit and the conversion unit to generate instructions that meet the device control requirements, such as "send an on signal to the living room light device".
[0028] The transmission unit is used to transmit commands to the set-top box via the Wi-Fi module.
[0029] In this embodiment, by modeling the collected voice information and calculating the probability of state transitions and output observations, a stable conversion from voice to text is achieved, ensuring the accuracy and reliability of text conversion. At the same time, considering the contextual information of the text, a deep understanding of semantics is achieved, which can accurately grasp the complex semantic relationships in the text and the user's true intentions, extract key information, and provide an accurate basis for the generation of subsequent operation instructions. In addition, by combining a bidirectional encoder and an autoregressive decoder, accurate operation instructions can be flexibly generated based on the input semantic information, thereby improving the level of intelligence of voice interaction.
[0030] Implementation 3: Please refer to Figure 1 and Figure 2 The present invention provides a technical solution: based on the first embodiment, the speech prediction end includes a frequency unit, an association unit and a judgment unit; The frequency unit is used to collect the number of device operations and calculate the operation frequency. The calculation formula is as follows: ; in, Represents the operating frequency. Represents a time period. Represents time period The number of times a certain operation is performed.
[0031] The correlation unit is used to calculate the probability of operation A and operation B occurring simultaneously, and the probability of operation B occurring when operation A occurs. The calculation steps are as follows: S1: Calculate the probability that operations A and B occur simultaneously, using the following formula: ; in, This represents the probability that operation A and operation B occur simultaneously. This represents the number of times operation A and operation B occur simultaneously. Represents the total number of operands; S2: Calculate the probability of operation B occurring when operation A occurs, using the following formula: ; in, This represents the probability that operation B will occur given that operation A has occurred. This represents the number of times operation A and operation B occur simultaneously. This represents the number of times operation A occurs.
[0032] The comparison unit is used to set the probability threshold for the simultaneous occurrence of operation A and operation B. The probability threshold of operation B occurring when operation A occurs. ,when Greater than or equal to and Greater than When this occurs, it means that operation A and operation B are strongly correlated.
[0033] For example, if a strong correlation is found between a user's actions of "turning off the bedroom light before bed" and "turning on sleep mode music", the system can automatically play sleep music when it detects the action of turning off the bedroom light to meet the user's personalized needs.
[0034] In this embodiment, by collecting the number of device operations and calculating the operation frequency, the frequency of user operations on the device can be quantified with specific values. At the same time, the probability of operation A and operation B occurring simultaneously and the probability of operation B occurring when operation A occurs can be calculated. This allows for in-depth exploration of the potential correlation between different operations. By comparing thresholds, strongly correlated operation combinations can be screened out. Based on these strongly correlated combinations, the system can optimize the operating mode according to user operating habits, simplify the operation process, improve ease of use, and enhance user willingness to use the system.
[0035] In this invention, the intelligent voice interaction method, system, set-top box, and storage medium utilize a collection unit that gathers voice information through a microphone array and measures the frequency of the sound signal using a frequency sensor. This allows for more comprehensive voice acquisition, providing high-quality raw data for subsequent voice analysis. Furthermore, by utilizing the phase difference generated when sound reaches different microphones, the user's azimuth angle can be calculated, accurately determining the sound source location. Based on the user's azimuth angle, the microphone array output signal is calculated, enhancing voice information in the target direction and suppressing noise and interference signals from other directions. Even in noisy environments, the user's voice can be highlighted, improving the success rate of voice recognition and ensuring smooth voice interaction. By modeling the collected voice information and calculating the probability of state transitions and output observations, a stable voice-to-text conversion is achieved, ensuring the accuracy and reliability of text conversion. Simultaneously, it considers… By considering text context information and deeply understanding semantics, it can accurately grasp complex semantic relationships and the user's true intentions in the text, extract key information, and provide an accurate basis for generating subsequent operation instructions. Combined with a bidirectional encoder and an autoregressive decoder, it can flexibly generate accurate operation instructions based on the input semantic information, improving the intelligence level of voice interaction. By collecting the number of device operations and calculating the operation frequency, it can quantify the frequency of user operations on the device with specific values. At the same time, it can calculate the probability of operation A and operation B occurring simultaneously, as well as the probability of operation B occurring when operation A occurs. It can deeply explore the potential correlation between different operations, and filter out strongly correlated operation combinations through threshold comparison. Based on the strongly correlated combination, the system can optimize the operation mode according to user operation habits, simplify the operation process, improve the ease of use, and enhance the user's willingness to use it. A smart voice interaction method includes the following steps: (1): Enter the voice receiving end, collect the voice information emitted by the user through the microphone array, measure the frequency of the sound signal through the frequency sensor, calculate the user's azimuth angle through the phase difference generated when the sound reaches different microphones, as well as the output signal and weighting coefficient of the microphone array, and transmit the collected information to the voice parsing end through the Wi-Fi module. (2): Enter the speech parsing end, which is used to convert speech into text through the HMM model, and use the BERT model to parse the text, extract user intent and semantic slots, and use the BART model to generate operation instructions based on user intent and semantic slots. Then the Wi-Fi module sends the operation instructions to the set-top box. (3): Enter the voice prediction terminal to collect the number of device operations, calculate the operation frequency, calculate the probability of operation A and operation B occurring simultaneously, and the probability of operation B occurring when operation A occurs, and compare it with the set threshold to determine the strongly correlated operation combination.
[0036] A set-top box, characterized in that it includes a memory for storing computer programs; A processor connected to the memory is used to execute the computer program stored in the memory.
[0037] A smart voice-interactive storage medium, wherein a computer-readable storage medium stores a computer program.
[0038] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.
[0039] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An intelligent voice interaction system, characterized in that, The system includes a voice receiver, a voice parser, and a voice prediction unit; The voice receiver is used to collect voice information through a microphone array, and to calculate the microphone array output signal based on the user's azimuth angle in real time, thereby enhancing the voice information in the target direction, and transmitting the collected voice information to the voice parsing end. The voice parsing terminal converts speech into text using speech recognition technology, and uses a language training model to parse the text, extract user intent and semantic slots; and generates operation instructions based on user intent and semantic slots, and sends the operation instructions to the set-top box; The voice prediction terminal collects the number of times the user operates the device through a data acquisition device, calculates the operation frequency, analyzes the correlation between different user operations, calculates the support and confidence thresholds between different operations, and filters out strongly correlated operation combinations.
2. The intelligent voice interaction system of claim 1, wherein, The voice receiver includes a collection unit, a positioning unit, an execution unit, and a transmission unit; The collection unit is used to collect voice information emitted by the user through a microphone array and to measure the frequency of the sound signal through a frequency sensor; The positioning unit is used to calculate the user's azimuth angle by the phase difference generated when sound reaches different microphones. The calculation formula is as follows: ; in, Represents the user's azimuth angle. This represents the spacing between different microphones. Represents the phase difference between different microphones. Represents the speed of sound. Represents frequency.
3. The intelligent voice interaction system according to claim 2, characterized in that, The execution unit is used to calculate the output signal and weighting coefficients of the microphone array. The calculation steps are as follows: Step 1: Calculate the output signal of the microphone array. The calculation formula is as follows: ; in, This represents the output signal of the microphone array at time t, where M represents the number of microphones in the array. Represents an index variable. Representing the The signal received by each element The corresponding weighting coefficients, Representing the The signal received by each array element at time t; Step 2, calculate the first The weighting coefficients corresponding to each array element are calculated using the following formula: ; in, Representing the The weighting coefficients corresponding to each array element Represents the imaginary unit. Represents the sound wave length, This represents the spacing between adjacent elements in a uniform linear array. Represents the user's azimuth angle. This represents the position index of the corresponding array element in the array; The transmission unit is used to transmit the collected information to the voice analysis terminal via the Wi-Fi module.
4. The intelligent voice interaction system according to claim 1, characterized in that, The speech parsing unit includes a conversion unit, a parsing unit, a generation unit, and a transmission unit; The conversion unit is used to convert speech into text using an HMM model; The parsing unit is used to parse text using the BERT model to extract user intent and semantic slots; The generation unit is used to generate operation instructions based on user intent and semantic slots using the BART model; The transmission unit is used to transmit instructions to the set-top box via a Wi-Fi module.
5. The intelligent voice interaction system according to claim 1, characterized in that, The speech prediction terminal includes a frequency unit, an association unit, and a judgment unit; The frequency unit is used to collect the number of device operations and calculate the operation frequency. The calculation formula is as follows: ; in, Represents the operating frequency. Represents a time period. Represents time period The number of times a certain operation is performed.
6. The intelligent voice interaction system according to claim 5, characterized in that, The association unit is used to calculate the probability of operation A and operation B occurring simultaneously, and the probability of operation B occurring when operation A occurs. The calculation steps are as follows: S1: Calculate the probability that operations A and B occur simultaneously, using the following formula: ; in, This represents the probability that operation A and operation B occur simultaneously. This represents the number of times operation A and operation B occur simultaneously. Represents the total number of operands; S2: Calculate the probability of operation B occurring when operation A occurs, using the following formula: ; in, This represents the probability that operation B will occur given that operation A has occurred. This represents the number of times operation A and operation B occur simultaneously. This represents the number of times operation A occurs.
7. The intelligent voice interaction system according to claim 6, characterized in that, The comparison unit is used to set a probability threshold for the simultaneous occurrence of operation A and operation B. The probability threshold of operation B occurring when operation A occurs. ,when Greater than or equal to and Greater than When this occurs, it means that operation A and operation B are strongly correlated.
8. An intelligent voice interaction method, referring to the intelligent voice interaction system according to any one of claims 1-7, characterized in that, Includes the following steps: (1): Enter the voice receiving end, collect the voice information emitted by the user through the microphone array, measure the frequency of the sound signal through the frequency sensor, calculate the user's azimuth angle through the phase difference generated when the sound reaches different microphones, as well as the output signal and weighting coefficient of the microphone array, and transmit the collected information to the voice parsing end through the Wi-Fi module. (2): Enter the speech parsing end, which is used to convert speech into text through the HMM model, and use the BERT model to parse the text, extract user intent and semantic slots, and use the BART model to generate operation instructions based on user intent and semantic slots. Then the Wi-Fi module sends the operation instructions to the set-top box. (3): Enter the voice prediction terminal to collect the number of device operations, calculate the operation frequency, calculate the probability of operation A and operation B occurring simultaneously, and the probability of operation B occurring when operation A occurs, and compare it with the set threshold to determine the strongly correlated operation combination.
9. A set-top box, characterized in that, Includes memory used to store computer programs; A processor connected to the memory is used to execute the computer program stored in the memory to implement the intelligent voice interaction system according to any one of claims 1-7.
10. An intelligent voice-interactive storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, performs the intelligent voice interaction system according to any one of claims 1-7.