Voice interaction method and system based on edge algorithm
By collecting and processing voice signals on local terminal devices, generating text sequences in real time and conducting database queries, the problem of voice interaction delay is solved, and low-latency and high-privacy voice interaction is achieved, which is suitable for smart homes, on-board systems and industrial control scenarios.
Patent Information
- Application Number
- CN202510382193.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-08
AI Technical Summary
The existing voice interaction technology is easily affected by the increase in network quality and number of users during the interaction between users and terminal devices, resulting in obvious interaction delay.
The original voice signal is collected through the microphone module of the local terminal device, and after preprocessing, the hardware-aware compressed speech recognition model is imported to generate text sequences in real time, and feature sequence matching and querying is performed in the local database, target instructions and response text are determined, and voice output is performed.
It realizes low latency and high privacy voice interaction, suitable for smart home, on-board systems and industrial control scenarios.
Smart Images

Figure CN120279915A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voice interaction, and particularly to a voice interaction method and system based on an edge algorithm. Background Art
[0002] Voice interaction usually exists between a user and a terminal device. The terminal device detects the voice emitted by the user and then gives a corresponding response, which is widely applicable to fields such as smart home and in-vehicle systems.
[0003] However, in the current voice interaction during the voice interaction process between the user and the terminal device, the generated data usually needs to be processed based on the background server of the terminal device. For example, when the network connection quality deteriorates, there will be an obvious delay in the voice interaction process, and all terminal devices perform data processing through the background server. With the increase in the number of users, the delay phenomenon becomes more obvious. Therefore, the current voice interaction is affected by objective factors and there is also a problem of easy interaction delay. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a voice interaction method and system based on an edge algorithm, aiming to solve the problem that voice interaction in the prior art is affected by objective factors and there is also a problem of easy interaction delay.
[0005] The first aspect of the present invention provides a voice interaction method based on an edge algorithm, and the method includes:
[0006] Collect an original voice signal through a microphone module of a local terminal device, and preprocess the original voice signal to obtain a target voice signal;
[0007] Import the target voice signal into a speech recognition model compressed by hardware perception to generate a text sequence corresponding to the target voice signal in real time;
[0008] According to the text features in the text sequence, extract the feature sequence in the text sequence, traverse the preset database according to the feature sequence, and determine the target instruction corresponding to the text sequence to obtain a voice parsing result;
[0009] According to the voice parsing result, traverse and query the response text corresponding to the target instruction in the database and perform voice output.
[0010] According to one aspect of the above technical solution, the step of extracting the feature sequence in the text sequence according to the text features in the text sequence, traversing the preset database according to the feature sequence, and determining the target instruction corresponding to the text sequence to obtain a voice parsing result includes:
[0011] Extract verbs, nouns, and location words from the text sequence according to the text features in the text sequence, and construct a feature sequence corresponding to the text sequence based on the verbs, nouns, and location words in the text sequence;
[0012] According to the feature sequence, perform a polling match of the feature sequence in a preset database to determine a target instruction corresponding to the text sequence, and obtain a voice parsing result.
[0013] According to one aspect of the above technical solution, the step of performing a polling match of the feature sequence in a preset database according to the feature sequence to determine a target instruction corresponding to the text sequence and obtain a voice parsing result includes:
[0014] According to the feature sequence and the state transition conditions of the finite state machine, perform a polling match of the feature sequence in the preset database through the finite state machine;
[0015] Determine a target instruction corresponding to the text sequence and obtain a voice parsing result.
[0016] According to one aspect of the above technical solution, the step of traversing and querying a response text corresponding to the target instruction in the database according to the voice parsing result and performing voice output includes:
[0017] Construct an index table for text indexing according to the target instruction in the voice parsing result, with the key being the instruction hash value and the value being the pre-stored response text;
[0018] According to the hash value corresponding to the target instruction, query the pre-stored response text corresponding to the hash value in the index table, determine the response text corresponding to the target instruction, and perform voice output according to the response text to complete voice interaction.
[0019] According to one aspect of the above technical solution, the step of querying the pre-stored response text corresponding to the hash value in the index table according to the hash value corresponding to the target instruction, determining the response text corresponding to the target instruction, and performing voice output according to the response text to complete voice interaction includes:
[0020] According to the hash value corresponding to the target instruction, query the pre-stored response text corresponding to the hash value in the index table, where the pre-stored response text is a response text extracted from historical interaction data;
[0021] When the key of the hash value matches the value of the pre-stored response text, determine the pre-stored response text as the response text corresponding to the target instruction;
[0022] Using the response text as voice data, the voice data of the response text is output through the speaker of the local terminal device to complete voice interaction.
[0023] According to one aspect of the above technical solution, the method further includes:
[0024] Obtain all historical instructions within a preset time period, obtain the trigger time of each historical instruction, and calculate the high-frequency trigger time range of each type of historical instruction;
[0025] When it is recognized that the current time is within the high-frequency trigger time range, automatically call the historical response text corresponding to the historical instruction in the database and store it locally;
[0026] When receiving a target instruction corresponding to the historical instruction, output a target response text according to the historical response text.
[0027] According to one aspect of the above technical solution, the step of outputting a target instruction according to the historical response instruction when receiving a target instruction corresponding to the historical instruction includes:
[0028] When receiving a target instruction corresponding to the historical instruction, generate a target response text according to the historical response text and the current weather information for voice output.
[0029] The second aspect of the present invention is to provide a voice interaction system based on an edge algorithm, which is applied to the method described in the above technical solution. The system includes:
[0030] An acquisition module, configured to acquire an original voice signal through a microphone module of a local terminal device, and preprocess the original voice signal to obtain a target voice signal;
[0031] A generation module, configured to import the target voice signal into a speech recognition model compressed by hardware perception, and generate a text sequence corresponding to the target voice signal in real time;
[0032] A traversal module, configured to extract a feature sequence from the text sequence according to the text features in the text sequence, traverse in a preset database according to the feature sequence, and determine a target instruction corresponding to the text sequence to obtain a voice parsing result;
[0033] An output module, configured to traverse and query a response text corresponding to the target instruction in the database according to the voice parsing result and perform voice output.
[0034] The third aspect of the present invention lies in providing a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the above technical solution is implemented.
[0035] The fourth aspect of the present invention lies in providing an electronic device, including a memory, a processor, and a computer program stored on the memory and operable on the processor, and when the processor executes the computer program, the method described in the above technical solution is implemented.
[0036] Compared with the prior art, the voice interaction method and system based on the edge algorithm shown in the present invention have the following beneficial effects:
[0037] In this embodiment, the original voice signal is collected by the microphone module of the local terminal device, and the original voice signal is preprocessed to obtain the target voice signal; the target voice signal is imported into the speech recognition model compressed by hardware perception to generate a text sequence corresponding to the target voice signal in real time; according to the text features in the text sequence, the feature sequence in the text sequence is extracted, and the feature sequence is traversed in a preset database to determine the target instruction corresponding to the text sequence to obtain the voice parsing result; according to the voice parsing result, the response text corresponding to the target instruction is traversed and queried in the database and voice output is performed. Then, in this embodiment, voice collection is performed and then text conversion is carried out, and text features are extracted to determine the target instruction, and finally the response text corresponding to the target instruction is locally queried in the database for voice output, so that a series of processes of voice data processing can be completed through the local terminal device, and low-latency and high-privacy voice interaction can be realized on the edge device. Typical applications include scenarios such as smart home, vehicle-mounted system, and industrial control. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:
[0039] Figure 1 is a schematic flow chart of the voice interaction method based on the edge algorithm shown in an embodiment of the present invention;
[0040] Figure 2 is a structural block diagram of the voice interaction system based on the edge algorithm shown in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To make the objectives, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0043] Embodiment 1
[0044] Please refer to Figure 1 , the first embodiment of the present invention provides a voice interaction method based on an edge algorithm, and the method includes steps S10 - S40:
[0045] Step S10, collect an original voice signal through a microphone module of a local terminal device, and preprocess the original voice signal to obtain a target voice signal.
[0046] In this embodiment, the local terminal device includes a mobile communication device, a voice interaction device, or other electronic devices equipped with a voice interaction system, etc. The method shown in this embodiment is a voice interaction method based on an edge algorithm. In specific implementation, an original voice signal of a user is collected through the microphone module of the local terminal device, and then the original voice signal is preprocessed such as data cleaning and denoising to obtain a target voice signal, so that the target voice signal is clearer for subsequent processing.
[0047] Step S20, import the target voice signal into a speech recognition model compressed by hardware perception, and generate a text sequence corresponding to the target voice signal in real time.
[0048] In this embodiment, after collecting the original voice signal through the microphone module and preprocessing the original voice signal to obtain a target voice signal, then the target voice signal is imported into a speech recognition model compressed by hardware perception, and the target voice signal is converted in real time through this speech recognition model to obtain a text sequence corresponding to the target voice signal.
[0049] Specifically, this text sequence is actually the conversion result of the user's output speech, specifically converting a segment of speech into a segment of text. This text sequence is, for example, "What's the temperature today" or "Set the air conditioner temperature to 26 degrees". In fact, it is equivalent to the control instruction of the user to the local terminal device. Through the edge algorithm, relevant data processing is completed within the local terminal device, which can effectively improve the data processing rate and avoid over-reliance on the background server.
[0050] Step S30: According to the text features in the text sequence, extract the feature sequence in the text sequence, traverse the preset database according to the feature sequence, and determine the target instruction corresponding to the text sequence to obtain the speech parsing result.
[0051] In this embodiment, the step of extracting the feature sequence in the text sequence according to the text features in the text sequence, traversing the preset database according to the feature sequence, and determining the target instruction corresponding to the text sequence to obtain the speech parsing result includes:
[0052] According to the text features in the text sequence, extract the verbs, nouns, and location words in the text sequence, and construct the feature sequence corresponding to the text sequence according to the verbs, nouns, and location words in the text sequence;
[0053] According to the feature sequence, perform polling matching of the feature sequence in the preset database, determine the target instruction corresponding to the text sequence, and obtain the speech parsing result.
[0054] Among them, the step of performing polling matching of the feature sequence in the preset database according to the feature sequence, determining the target instruction corresponding to the text sequence, and obtaining the speech parsing result includes:
[0055] According to the feature sequence and the state transition conditions of the finite state machine, perform polling matching of the feature sequence in the preset database through the finite state machine;
[0056] Determine the target instruction corresponding to the text sequence and obtain the speech parsing result.
[0057] Specifically, after importing the target speech signal into the speech recognition model with hardware compressed sensing to generate the text sequence corresponding to the target speech signal in real time, according to the text features in the text sequence, the feature sequence in the text sequence will be extracted according to the verbs, nouns, and location words in the text sequence. For example, the feature sequence "air conditioner", "temperature", and "twenty-six" in the text sequence "Set the air conditioner temperature to 26 degrees" is extracted, and then polling matching is performed in the pre-constructed database according to the extracted feature sequence to determine the target instruction corresponding to the text sequence and obtain the speech parsing result corresponding to the user.
[0058] Step S40: According to the voice parsing result, traverse and query in the database for the response text corresponding to the target instruction and perform voice output.
[0059] In this embodiment, the step of traversing and querying in the database for the response text corresponding to the target instruction according to the voice parsing result and performing voice output includes:
[0060] Construct an index table for text indexing according to the target instruction in the voice parsing result, with the key being the instruction hash value and the value being the pre-stored response text;
[0061] According to the hash value corresponding to the target instruction, query in the index table for the pre-stored response text corresponding to the hash value, determine the response text corresponding to the target instruction, and perform voice output according to the response text to complete the voice interaction.
[0062] Among them, the step of querying in the index table for the pre-stored response text corresponding to the hash value according to the hash value corresponding to the target instruction, determining the response text corresponding to the target instruction, and performing voice output according to the response text to complete the voice interaction includes:
[0063] According to the hash value corresponding to the target instruction, query in the index table for the pre-stored response text corresponding to the hash value, where the pre-stored response text is the response text extracted from historical interaction data;
[0064] When the key of the hash value matches the value of the pre-stored response text, determine that the pre-stored response text is the response text corresponding to the target instruction;
[0065] Use the response text as voice data and perform voice output of the voice data of the response text through the speaker of the local terminal device to complete the voice interaction.
[0066] Specifically, after determining the target instruction corresponding to the text sequence to obtain the voice parsing result, an index table for text indexing, that is, an index relationship table or a comparison relationship table, will be constructed first according to the target instruction in the voice parsing result, with the key being the hash value of the target instruction and the value being the pre-stored response text in the database. Then, according to the hash value corresponding to the target instruction (such as "Turn on the living room light" → 0x5E3A), retrieve the pre-stored response text corresponding to the hash value (such as "Okay", "The living room light has been turned on"), so as to determine the response text corresponding to the target instruction, that is, the voice content that the local terminal device needs to output. Then, perform voice output according to the retrieved response text to complete the voice interaction with the user.
[0067] More specifically, in the index table, the key is a hash value randomly generated according to a preset rule and corresponding to the target instruction. The instruction hash values corresponding to each target instruction are all different, and the pre-stored response text corresponding to the instruction hash value (or target instruction) is automatically generated through the training of the language model in the background. There is at least one, usually multiple, response text corresponding to each target instruction, and all are pre-stored in the database.
[0068] Moreover, in this embodiment, a user portrait is also constructed through the historical interaction data between the local terminal device and the user, so as to analyze user characteristics, such as living patterns, etc., in order to better serve all aspects of the user.
[0069] Compared with the prior art, the beneficial effects of adopting the voice interaction method based on the edge algorithm shown in this embodiment are as follows:
[0070] In this embodiment, the original voice signal is collected by the microphone module of the local terminal device, and the original voice signal is preprocessed to obtain the target voice signal; the target voice signal is imported into the speech recognition model compressed by hardware perception to generate a text sequence corresponding to the target voice signal in real time; according to the text features in the text sequence, the feature sequence in the text sequence is extracted, and the preset database is traversed according to the feature sequence to determine the target instruction corresponding to the text sequence to obtain the voice parsing result; according to the voice parsing result, the response text corresponding to the target instruction is traversed and queried in the database and voice output is performed. Then, in this embodiment, voice collection is performed and then text conversion is carried out, and text features are extracted to determine the target instruction, and finally the response text corresponding to the target instruction is locally queried in the database for voice output, so that a series of processes of voice data processing can be completed through the local terminal device, and low-latency and high-privacy voice interaction can be realized on edge devices. Typical applications include scenarios such as smart home, in-vehicle system, and industrial control.
[0071] Embodiment 2
[0072] The second embodiment of the present invention also provides a voice interaction method based on the edge algorithm. The method shown in this embodiment is basically the same as the method shown in the first embodiment, and the difference lies in:
[0073] In this embodiment, the method further includes:
[0074] Obtain all historical instructions within a preset time period, obtain the trigger time of each historical instruction, and calculate the high-frequency trigger time range of each historical instruction;
[0075] When it is recognized that the current time is within the high-frequency trigger time range, automatically call the historical response text corresponding to the historical instruction in the database and store it locally;
[0076] When receiving a target instruction corresponding to the historical instruction, output a target response text according to the historical response text.
[0077] Among them, the step of outputting a target instruction according to the historical response instruction when receiving a target instruction corresponding to the historical instruction includes:
[0078] When receiving a target instruction corresponding to the historical instruction, generate a target response text according to the historical response text and the current weather information for voice output.
[0079] Specifically, at any moment, in this embodiment, all historical instructions within a preset time period will also be obtained, that is, the instructions triggered at historical moments. Then, the trigger time corresponding to each historical instruction is obtained, and the high-frequency trigger time range of each type of historical instruction is calculated according to the trigger time of the historical instruction, which is equivalent to outputting user habits. For example, in a home application scenario, the high-frequency trigger time range of the historical instruction "open the curtain" is 7:00 - 7:30, and the high-frequency trigger time range of the historical instruction "turn on the living room light" is "17:30 - 18:30". Then, when it is detected in a future moment that the current time is within the high-frequency trigger time range, the historical response text corresponding to the historical instruction will be automatically called in the database and stored locally. Then, when receiving a target instruction corresponding to the historical instruction, combined with the historical response text and the weather information of the current day, a target response text corresponding to the current time is automatically generated for voice output.
[0080] It should be noted here that in this embodiment, the high-frequency trigger time range is calculated by outputting the trigger time of the historical instruction, which is equivalent to outputting user habits. Then, the local terminal device can quickly respond when receiving a target instruction similar to or the same as the historical instruction, and adaptively adjust the historical response text in combination with the actual weather information for voice output, thereby improving both the interaction efficiency and the user experience.
[0081] Embodiment III
[0082] Please refer to Figure 2 , the third embodiment of the present invention provides a voice interaction system based on an edge algorithm, which is applied to the method described in any of the above embodiments. The system includes:
[0083] An acquisition module 10, configured to collect an original voice signal through a microphone module of a local terminal device and preprocess the original voice signal to obtain a target voice signal;
[0084] A generation module 20 for importing the target voice signal into a hardware-aware compressed speech recognition model to generate a text sequence corresponding to the target voice signal in real time;
[0085] A traversal module 30 for extracting a feature sequence from the text sequence according to the text features in the text sequence, traversing in a preset database according to the feature sequence, and determining a target instruction corresponding to the text sequence to obtain a voice parsing result;
[0086] An output module 40 for traversing and querying a response text corresponding to the target instruction in the database according to the voice parsing result and performing voice output.
[0087] Compared with the prior art, the voice interaction system based on the edge algorithm shown in this embodiment has the beneficial effects that:
[0088] In this embodiment, the original voice signal is collected by the microphone module of the local terminal device, and the original voice signal is preprocessed to obtain a target voice signal; the target voice signal is imported into a hardware-aware compressed speech recognition model to generate a text sequence corresponding to the target voice signal in real time; according to the text features in the text sequence, a feature sequence in the text sequence is extracted, and traversing is performed in a preset database according to the feature sequence to determine a target instruction corresponding to the text sequence to obtain a voice parsing result; according to the voice parsing result, a response text corresponding to the target instruction is traversed and queried in the database and voice output is performed. Then, in this embodiment, voice collection is performed and then text conversion is performed, and text features are extracted to determine the target instruction, and finally a response text corresponding to the target instruction is locally queried in the database for voice output, so that a series of processes of voice data processing can be completed through the local terminal device, and low-latency and high-privacy voice interaction can be realized on edge devices. Typical applications include scenarios such as smart homes, vehicle-mounted systems, and industrial control.
[0089] Embodiment 4
[0090] The fourth embodiment of the present invention provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any of the above embodiments is implemented.
[0091] Embodiment 5
[0092] The fifth embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the method described in any of the above embodiments is implemented.
[0093] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0094] The above-described embodiments merely represent several implementation manners of the present invention. The descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A voice interaction method based on an edge algorithm, characterized in that, The method includes: Collecting an original voice signal through a microphone module of a local terminal device, and preprocessing the original voice signal to obtain a target voice signal; Importing the target voice signal into a speech recognition model compressed by hardware perception, and generating a text sequence corresponding to the target voice signal in real time; According to the text features in the text sequence, extracting a feature sequence from the text sequence, traversing in a preset database according to the feature sequence, and determining a target instruction corresponding to the text sequence to obtain a speech parsing result; According to the speech parsing result, traversing and querying in the database for a response text corresponding to the target instruction and performing voice output.
2. The voice interaction method based on an edge algorithm according to claim 1, wherein, The step of, according to the text features in the text sequence, extracting a feature sequence from the text sequence, traversing in a preset database according to the feature sequence, and determining a target instruction corresponding to the text sequence to obtain a speech parsing result, includes: According to the text features in the text sequence, extracting verbs, nouns, and location words from the text sequence, and constructing a feature sequence corresponding to the text sequence according to the verbs, nouns, and location words in the text sequence; According to the feature sequence, performing polling matching of the feature sequence in a preset database, and determining a target instruction corresponding to the text sequence to obtain a speech parsing result.
3. The voice interaction method based on an edge algorithm according to claim 2, wherein The step of, according to the feature sequence, performing polling matching of the feature sequence in a preset database, and determining a target instruction corresponding to the text sequence to obtain a speech parsing result, includes: According to the feature sequence and the state transition conditions of a finite state machine, performing polling matching of the feature sequence in a preset database through the finite state machine; Determining a target instruction corresponding to the text sequence to obtain a speech parsing result.
4. The voice interaction method based on an edge algorithm according to claim 1, characterized in that, The step of, according to the speech parsing result, traversing and querying in the database for a response text corresponding to the target instruction and performing voice output, includes: According to the target instruction in the speech parsing result, constructing an index table for text indexing, with the key being the instruction hash value and the value being the pre-stored response text; According to the hash value corresponding to the target instruction, querying in the index table for the pre-stored response text corresponding to the hash value, determining the response text corresponding to the target instruction, and performing voice output according to the response text to complete voice interaction.
5. The voice interaction method based on an edge algorithm according to claim 4, wherein, The step of, according to the hash value corresponding to the target instruction, querying in the index table for the pre-stored response text corresponding to the hash value, determining the response text corresponding to the target instruction, and performing voice output according to the response text to complete voice interaction, includes: According to the hash value corresponding to the target instruction, querying in the index table for the pre-stored response text corresponding to the hash value, where the pre-stored response text is a response text extracted from historical interaction data; When the key of the hash value matches the value of the pre-stored response text, determining the pre-stored response text as the response text corresponding to the target instruction; Using the response text as voice data, and performing voice output of the voice data of the response text through a speaker of the local terminal device to complete voice interaction.
6. The voice interaction method based on an edge algorithm according to any one of claims 1-5, characterized in that The method further includes: Obtaining all historical instructions within a preset time period, obtaining the triggering time of each historical instruction, and calculating the high-frequency triggering time range of each type of historical instruction; When it is recognized that the current time is within the high-frequency triggering time range, automatically calling the historical response text corresponding to the historical instruction in the database and storing it locally; When receiving a target instruction corresponding to the historical instruction, outputting a target response text according to the historical response text.
7. The voice interaction method based on an edge algorithm according to claim 6, wherein The step of outputting a target instruction according to the historical response instruction when receiving a target instruction corresponding to the historical instruction includes: When receiving a target instruction corresponding to the historical instruction, generating a target response text according to the historical response text and the current weather information for voice output.
8. A voice interaction system based on an edge algorithm, characterized in that, Applied to the method according to any one of claims 1-7, the system includes: A collection module, configured to collect an original voice signal through a microphone module of a local terminal device and preprocess the original voice signal to obtain a target voice signal; A generation module, configured to import the target voice signal into a speech recognition model compressed by hardware perception to generate a text sequence corresponding to the target voice signal in real time; A traversal module, configured to extract a feature sequence from the text sequence according to the text features in the text sequence, traverse in a preset database according to the feature sequence, and determine a target instruction corresponding to the text sequence to obtain a speech parsing result; An output module, configured to traverse and query a response text corresponding to the target instruction in the database according to the speech parsing result and perform voice output.
9. A readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1-7 is implemented.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1-7 is implemented.