Continuous interaction method of voice and related products
By acquiring voice data through the terminal and using the VAD engine to monitor in-vehicle dialogue, the problem of discontinuous in-vehicle voice interaction has been solved, resulting in a better user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI PATEO ELECTRONIC EQUIPMENT MANUFACTURING CO LTD
- Filing Date
- 2021-12-02
- Publication Date
- 2026-05-12
AI Technical Summary
Existing in-vehicle voice interaction systems cannot achieve continuous interaction, which affects the user experience.
By acquiring voice data through the terminal and continuously monitoring the dialogue between the target object and the vehicle system using the VAD engine, continuous voice interaction can be achieved.
It improves the continuity of voice interaction and enhances the user experience.
Smart Images

Figure CN116259304B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, and in particular to a continuous speech interaction method and related products. Background Technology
[0002] In-vehicle voice interaction is crucial in connected car products. It is becoming increasingly common to use voice commands to operate navigation and listen to multimedia while driving. However, existing in-vehicle voice interaction systems cannot achieve continuous interaction, which affects the user experience. Summary of the Invention
[0003] This application discloses a continuous voice interaction method and related products, which can realize continuous in-vehicle voice interaction and improve the user experience.
[0004] Firstly, a continuous voice interaction method is provided, the method comprising the following steps:
[0005] The terminal acquires the first voice data input by the target object, and after recognizing the first voice data as a wake-up voice, it sends the preset data in the cache to the terminal's voice activity detection VAD engine.
[0006] The terminal's VAD engine continuously monitors the target object and the vehicle-to-everything (V2X) dialogue and processes the dialogue.
[0007] After the target has finished speaking, the terminal's VAD engine sends the preset data in the cache back to the terminal's VAD engine to achieve continuous voice interaction.
[0008] Optionally, the method further includes:
[0009] If the target object finishes speaking, check if the continuous voice end condition is met. If the continuous voice end condition is met, stop sending the preset data in the cache to the terminal's VAD engine to stop the continuous voice interaction.
[0010] Optionally, the detection of whether the continuous speech termination condition is met specifically includes:
[0011] After the target finishes speaking, a timer is started. The timer stops when the target's voice data is received again. The first duration of the timer is obtained. If the first duration is greater than the time threshold, it is determined that the continuous voice termination condition is met.
[0012] Optionally, the detection of whether the continuous speech termination condition is met specifically includes:
[0013] After the target finishes speaking, the system receives a second voice data input from the target. When the second voice data is identified as belonging to a specific voice that ends continuous voice interaction, the continuous voice interaction end condition is met.
[0014] Optionally, the preset data in the cache specifically includes:
[0015] The terminal obtains preset recording data, performs noise reduction and echo cancellation on the recording data to obtain processed data, and stores the processed data in the cache as preset data.
[0016] Optionally, the step of identifying and determining the specific voice belonging to the end of continuous voice interaction from the second voice data specifically includes:
[0017] For the second speech data, an RNN recognition algorithm is used to determine x confidence rates for x words corresponding to each pronunciation group in the speech data; an LSTM recognition algorithm is used to determine y confidence rates for y words corresponding to each pronunciation group in the speech data.
[0018] The terminal device adds the two confidence rates of the same first word among x and y words in the first pronunciation group to obtain the confidence rate sum of the first word. It then iterates through the same words among x and y words to obtain the confidence rate sum of each same word. The word corresponding to the maximum confidence rate sum is determined as the word corresponding to the first pronunciation group. Iterates through all pronunciation groups to obtain the words corresponding to all pronunciation groups. All words are arranged in chronological order to form the text data of the second speech data. If the text data belongs to a specific text, the second speech data is determined to belong to the specific speech that ends continuous speech interaction.
[0019] Optionally, the step of using an RNN recognition algorithm to determine the x confidence rates for each pronunciation group corresponding to x words in the second speech data specifically includes:
[0020] S t =X t ×W+S t-1 ×W
[0021] O t =f(S) t )
[0022] Where W represents the weight, X t-1 X represents the input data of the input layer at time t-1 of the second speech data. t S represents the input data of the input layer at time t of the second speech data. t-1 This represents the output of the hidden layer at time t-1, where f represents the activation function, and O is the time limit. t-1 This represents the output result of the output layer at time t-1; based on the output result, determine the x confidence rates for each pronunciation group corresponding to x words.
[0023] Optionally, the activation function specifically includes:
[0024] sigmoid function or tanh function
[0025]
[0026] Secondly, a continuous voice interaction system is provided, the system comprising:
[0027] The acquisition unit is used to acquire the first voice data input by the target object;
[0028] The processing unit is used to identify the first voice data as a wake-up voice and then send the preset data in the cache to the terminal's voice activity detection VAD engine.
[0029] The VAD engine is used to continuously monitor the dialogue between the target object and the vehicle system, and process the dialogue; after detecting that the target object has finished speaking;
[0030] The processing unit is also used to send the preset data in the cache back to the terminal's VAD engine to achieve continuous voice interaction.
[0031] To achieve the above objectives, a third aspect provides an electronic device, the electronic device comprising:
[0032] At least one processor, the at least one processor communicating with the positioning module; and
[0033] A memory configured to store instructions that, when executed by the at least one processor, cause the at least one processor to perform steps, the steps including:
[0034] The system acquires the first voice data input by the target object, identifies the first voice data as a wake-up voice, and sends the preset data in the cache to the terminal's voice activity detection VAD engine.
[0035] The VAD engine continuously monitors the target object and the vehicle-to-everything (V2X) dialogue and processes the dialogue.
[0036] After detecting that the target has finished speaking, the VAD engine sends the preset data in the cache back to the terminal's VAD engine to achieve continuous voice interaction.
[0037] To achieve the above objectives, a fourth aspect provides a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform the method provided in the first aspect.
[0038] To achieve the above objectives, in a fifth aspect, a computer program product is provided, comprising a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in the first aspect of the embodiments of this application. The computer program product may be a software installation package.
[0039] By implementing the embodiments of this application, the technical solution provided by this application obtains the location information of vehicles to be put into storage. When the location information is within a preset range, the system automatically performs an inventory check to determine whether the vehicle to be put into storage belongs to a vehicle eligible for financial financing. When it is determined that the vehicle to be put into storage belongs to a vehicle eligible for financial financing, the electronic seal of the vehicle to be put into storage is opened. After obtaining the abnormal data of the vehicle eligible for financial financing, the system queries the status of the electronic seal of the vehicle eligible for financial financing. If the electronic seal is in an opened state, the electronic seal is pushed to the vehicle eligible for financial financing to execute an alarm process. This can conveniently obtain information about the vehicle eligible for financial financing, reduce labor costs, and the alarm process can reduce the financial risk of the vehicle eligible for financial financing. Attached Figure Description
[0040] The accompanying drawings used in the embodiments of this application are described below.
[0041] Figure 1 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;
[0042] Figure 2 This is a flowchart illustrating a risk management method for warehouse financing vehicles provided in an embodiment of this application;
[0043] Figure 3 This is a schematic diagram of the structure of a risk management system for warehouse financing vehicles provided in an embodiment of this application;
[0044] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0045] The embodiments of this application are described below with reference to the accompanying drawings.
[0046] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document indicates that the preceding and following related objects have an "or" relationship.
[0047] In this application's embodiments, "multiple" refers to two or more. The use of terms like "first," "second," etc., in this application's embodiments is merely illustrative and for distinguishing the described objects; it has no order and does not indicate a specific limitation on the number of devices in this application's embodiments, nor does it constitute any limitation on the embodiments of this application. The term "connection" in this application's embodiments refers to various connection methods, such as direct or indirect connections, to achieve communication between devices; this application's embodiments do not impose any limitations on this.
[0048] See Figure 1 , Figure 1A structural block diagram of a terminal device is provided, such as Figure 1 As shown, the terminal device may include a processor, memory, communication unit, and bus. Depending on the function, the processor may be equipped with hardware structures such as a microphone and mobile devices. In practical applications, it can also be implemented based on different hardware. In practical applications, the terminal device may also be integrated into other hardware devices, such as smart vehicle devices, smartphones, servers, computer devices, etc.
[0049] In-vehicle voice interaction is crucial in connected car products. Using voice commands for navigation and multimedia while driving is becoming increasingly common. However, existing voice interaction methods are inconvenient for continuous conversations with the vehicle's infotainment system; each interaction requires waking it up. Similar to human conversation, each interaction requires calling out the system's name, such as "Xiao Ai" or "Siri," making continuous dialogue cumbersome. Therefore, a method is needed that allows users to call out the other party's name only once, enabling continuous conversation without needing to say the name before each sentence. This technology makes in-vehicle systems more intelligent and voice interaction more user-friendly.
[0050] See Figure 2 , Figure 2 A continuous voice interaction method is provided, which can be executed on a terminal device, specifically a smart vehicle terminal. (See reference...) Figure 2 The method includes the following steps:
[0051] Step S201: The terminal obtains the first voice data input by the target object. After recognizing the first voice data as a wake-up voice, the terminal sends the preset data in the cache to the terminal's VAD (Voice Activity Detection) engine.
[0052] The terminal can acquire the first voice data input from the target object through its microphone. In practical applications, it can also receive voice data from other terminal devices via its communication unit. This voice data is generally the raw, unprocessed voice data.
[0053] Step S202: The terminal's VAD engine continuously monitors the dialogue between the target object and the vehicle system, and processes the dialogue.
[0054] Voice activity detection (VAD) is a technique used in speech processing to detect the presence of speech signals. VAD is primarily used for speech coding and speech recognition. It can simplify speech processing and can also be used to remove non-speech segments during audio sessions: in IP telephony applications, it can avoid encoding and transmitting silence packets, saving computation time and bandwidth.
[0055] VAD technology enables a range of speech-based applications. Therefore, there are various VAD algorithms with different characteristics and latency, sensitivity, accuracy, and computational cost. Some VAD algorithms also provide further analysis, such as whether the speech is voiced, unvoiced, or sustained. Speech activity detection is generally language-independent.
[0056] Step S203: After the target object finishes speaking, the terminal's VAD engine sends the preset data in the cache back to the terminal's VAD engine to achieve continuous voice interaction.
[0057] The aforementioned pre-stored data can be wake-up voice data for in-vehicle terminal devices, such as "Xiao Ai Tongxue" or "Siri," or other voice data or data that the VAD engine can recognize. This application does not limit the specific types mentioned above.
[0058] The technical solution provided in this application allows the terminal to acquire first voice data input by the target object. After recognizing the first voice data as a wake-up voice, the terminal sends the preset data in the cache to the terminal's Voice Activity Detection (VAD) engine. The terminal's VAD engine continuously monitors the dialogue between the target object and the vehicle system and processes the dialogue. After detecting that the target object has finished speaking, the terminal's VAD engine sends the preset data in the cache back to the terminal's VAD engine to achieve continuous voice interaction. This enables continuous voice interaction without any interruptions, improving the continuity of voice interaction and enhancing the user experience.
[0059] Optionally, the method further includes:
[0060] If the target object finishes speaking, check if the continuous voice end condition is met. If the continuous voice end condition is met, stop sending the preset data in the cache to the terminal's VAD engine to stop the continuous voice interaction.
[0061] For example, the above detection of whether the continuous speech termination condition is met may specifically include:
[0062] After the target finishes speaking, a timer is started. The timer stops when the target's voice data is received again. The first duration of the timer is obtained. If the first duration is greater than the time threshold, it is determined that the continuous voice termination condition is met.
[0063] The aforementioned timer can be a software-generated timer, or it can be other types of timers, such as a timing circuit. The aforementioned time threshold can be set by the user at any time, or it can be configured by the manufacturer, such as 5 seconds, 10 seconds, etc.
[0064] For example, the above detection of whether the continuous speech termination condition is met specifically includes:
[0065] After the target finishes speaking, the system receives a second voice data input from the target. When the second voice data is identified as belonging to a specific voice that ends continuous voice interaction, the continuous voice interaction end condition is met.
[0066] For example, the preset data in the cache specifically includes:
[0067] The terminal obtains preset recording data, performs noise reduction and echo cancellation on the recording data to obtain processed data, and stores the processed data in the cache as preset data.
[0068] The aforementioned recording data can be data previously recorded by the user, such as data collected during historical wake-ups, or other types of recording data. For example, the recording data mentioned above is the user's own data.
[0069] For example, the above-mentioned identification of specific speech belonging to the end of continuous speech interaction based on the second speech data specifically includes:
[0070] For the second speech data, an RNN recognition algorithm is used to determine x confidence rates for x words corresponding to each pronunciation group in the speech data; an LSTM recognition algorithm is used to determine y confidence rates for y words corresponding to each pronunciation group in the speech data.
[0071] The terminal device adds the two confidence rates of the same first word among x and y words in the first pronunciation group to obtain the confidence rate sum of the first word. It then iterates through the same words among x and y words to obtain the confidence rate sum of each same word. The word corresponding to the maximum confidence rate sum is determined as the word corresponding to the first pronunciation group. Iterates through all pronunciation groups to obtain the words corresponding to all pronunciation groups. All words are arranged in chronological order to form the text data of the second speech data. If the text data belongs to a specific text, the second speech data is determined to belong to the specific speech that ends continuous speech interaction.
[0072] For example, the step of using an RNN recognition algorithm to determine the x confidence rates for each pronunciation group corresponding to x words in the second speech data specifically includes:
[0073] S t =X t ×W+S t-1 ×W
[0074] O t =f(S) t )
[0075] Where W represents the weight, X t-1 X represents the input data of the input layer at time t-1 of the second speech data.t S represents the input data of the input layer at time t of the second speech data. t-1 This represents the output of the hidden layer at time t-1, where f represents the activation function, and O is the time limit. t-1 This represents the output result of the output layer at time t-1; based on the output result, determine the x confidence rates for each pronunciation group corresponding to x words.
[0076] For example, the activation function specifically includes:
[0077] sigmoid function or tanh function
[0078]
[0079] See Figure 3 , Figure 3 A continuous voice interaction system is provided, the system comprising:
[0080] Acquisition unit 301 is used to acquire the first voice data input by the target object;
[0081] Processing unit 302 is used to identify the first voice data as wake-up voice and then send the preset data in the cache to the terminal's voice activity detection VAD engine.
[0082] The VAD engine 303 is used to continuously monitor the dialogue between the target object and the vehicle system, and process the dialogue; after detecting that the target object has finished speaking;
[0083] The processing unit 302 is also used to send the preset data in the cache back to the terminal's VAD engine to realize continuous voice interaction.
[0084] Optionally, the processing unit 302 is further configured to detect whether the continuous voice end condition is met after the target object finishes speaking, and if the continuous voice end condition is met, stop sending the preset data in the cache to the terminal's VAD engine to stop the continuous voice interaction.
[0085] Optionally, the processing unit 302 is specifically used to start a timer after the target object finishes speaking. The timer is used to stop when the target object's voice data is received again. The first duration of the timer is obtained. If the first duration is greater than the time threshold, it is determined that the continuous voice termination condition is met.
[0086] For example, the processing unit 302 is also configured to receive the second voice data input by the target object again after the target object has finished speaking, and determine that the continuous voice end condition is met when the second voice data is identified as belonging to a specific voice that ends continuous voice interaction.
[0087] For example, processing unit 302 acquires preset recording data, performs noise reduction and echo cancellation on the recording data to obtain processed data, and stores the processed data in a buffer as preset data.
[0088] Optionally, the processing unit 302 is further configured to use an RNN recognition algorithm to determine x confidence rates for x words corresponding to each pronunciation group in the second speech data; use an LSTM recognition algorithm to determine y confidence rates for y words corresponding to each pronunciation group in the speech data; add the two confidence rates of the same first word among the x words and y words in the first pronunciation group to obtain the confidence rate sum of the first word; traverse the same words among the x words and y words to obtain the confidence rate sum of each same word; determine the word corresponding to the maximum confidence rate sum as the word corresponding to the first pronunciation group; traverse all pronunciation groups to obtain the words corresponding to all pronunciation groups; assemble all the words into text data of the second speech data in chronological order; if the text data belongs to a specific text, determine that the second speech data belongs to a specific speech that ends continuous speech interaction.
[0089] For example, the step of using an RNN recognition algorithm to determine the x confidence rates for each pronunciation group corresponding to x words in the second speech data specifically includes:
[0090] S t =X t ×W+S t-1 ×W
[0091] O t =f(S) t )
[0092] Where W represents the weight, X t-1 X represents the input data of the input layer at time t-1 of the second speech data. t S represents the input data of the input layer at time t of the second speech data. t-1 This represents the output of the hidden layer at time t-1, where f represents the activation function, and O is the time limit. t-1 This represents the output result of the output layer at time t-1; based on the output result, determine the x confidence rates for each pronunciation group corresponding to x words.
[0093] Please see Figure 4 , Figure 4 This application provides an electronic device 40, which includes a processor 401, a memory 402, and a communication interface 403. The processor 401, the memory 402, and the communication interface 403 are interconnected via a bus.
[0094] The memory 402 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), and is used for related computer programs and data. The communication interface 403 is used for receiving and sending data.
[0095] Processor 401 can be one or more central processing units (CPUs). If processor 401 is a CPU, the CPU can be a single-core CPU or a multi-core CPU.
[0096] Processor 401 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). Different processing units may be independent components or integrated into one or more processors. In some embodiments, the user equipment may also include one or more processing units. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. In other embodiments, the processing unit may also include a memory for storing instructions and data. For example, the memory in the processing unit may be a cache memory. This memory can store instructions or data that the processing unit has just used or is repeatedly used. If the processing unit needs to reuse the instruction or data, it can directly retrieve it from the memory. This avoids repeated access, reduces the waiting time of the processing unit, and thus improves the efficiency of the user equipment in processing data or executing instructions.
[0097] In some embodiments, the processor 401 may include one or more interfaces. These interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface. The USB interface is a USB-compliant interface, specifically a Mini USB interface, a Micro USB interface, a USB Type-C interface, etc. The USB interface can be used to connect a charger to charge the user device, and can also be used for data transfer between the user device and peripheral devices. The USB interface can also be used to connect headphones for audio playback.
[0098] If the electronic device 40 is a user device, such as a smartphone, the processor 401 in the electronic device 40 is used to read the computer program code stored in the memory 402 and execute it, such as... Figure 2 The method shown.
[0099] All relevant content in each scenario involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0100] This application also provides a computer-readable storage medium storing a computer program that, when run on a network device,... Figure 2 The method flow shown is thus implemented.
[0101] This application also provides a computer program product, which, when run on a terminal, provides a method for... Figure 2 The method flow shown is thus implemented.
[0102] This application also provides an electronic device, including a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, the programs including functions for executing... Figure 2 Instructions for the steps in the method of the illustrated embodiment.
[0103] The above primarily describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the electronic device includes the corresponding hardware structure and / or software template for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0104] This application embodiment can divide the electronic device into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0105] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and templates involved are not necessarily essential to this application.
[0106] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0107] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0108] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0109] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0110] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0111] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
Claims
1. A continuous voice interaction method, characterized in that, The method includes the following steps: The terminal acquires the first voice data input by the target object, and after recognizing the first voice data as a wake-up voice, it sends the preset data in the cache to the terminal's voice activity detection VAD engine. The terminal's VAD engine continuously monitors the target object and the vehicle-to-everything (V2X) dialogue and processes the dialogue. After the target has finished speaking, the terminal's VAD engine will send the preset data in the cache back to the terminal's VAD engine to achieve continuous voice interaction. If the target object finishes speaking, check whether the continuous voice end condition is met. If the continuous voice end condition is met, stop sending the preset data in the buffer to the terminal's VAD engine to stop the continuous voice interaction. The detection of whether the continuous speech termination condition is met specifically includes: After the target finishes speaking, the system receives the second voice data input by the target again. When the second voice data is identified as belonging to a specific voice that ends continuous voice interaction, the continuous voice end condition is determined to be met. The specific steps of identifying and determining the second voice data as belonging to the end of continuous voice interaction include: For the second speech data, an RNN recognition algorithm is used to determine x confidence rates for x words corresponding to each pronunciation group in the speech data; an LSTM recognition algorithm is used to determine y confidence rates for y words corresponding to each pronunciation group in the speech data. The terminal device adds the two confidence rates of the same first word among x and y words in the first pronunciation group to obtain the confidence rate sum of the first word. It then iterates through the same words among x and y words to obtain the confidence rate sum of each same word. The word corresponding to the maximum confidence rate sum is determined as the word corresponding to the first pronunciation group. Iterates through all pronunciation groups to obtain the words corresponding to all pronunciation groups. All words are arranged in chronological order to form the text data of the second speech data. If the text data belongs to a specific text, the second speech data is determined to belong to the specific speech that ends continuous speech interaction.
2. The method according to claim 1, characterized in that, The detection of whether the continuous speech termination condition is met specifically includes: After the target finishes speaking, a timer is started. The timer stops when the target's voice data is received again. The first duration of the timer is obtained. If the first duration is greater than the time threshold, it is determined that the continuous voice termination condition is met.
3. The method according to claim 1, characterized in that, The preset data in the cache specifically includes: The terminal obtains preset recording data, performs noise reduction and echo cancellation on the recording data to obtain processed data, and stores the processed data in the cache as preset data.
4. The method according to claim 1, characterized in that, The step of using an RNN recognition algorithm to determine the x confidence rates for each pronunciation group corresponding to x words in the second speech data specifically includes: Where W represents the weight, X t-1 X represents the input data of the input layer at time t-1 of the second speech data. t S represents the input data of the input layer at time t of the second speech data. t-1 This represents the output of the hidden layer at time t-1, where f represents the activation function, and O is the time limit. t-1 This represents the output result of the output layer at time t-1; based on the output result, determine the x confidence rates for each pronunciation group corresponding to x words.
5. The method according to claim 4, characterized in that, The activation function specifically includes: sigmoid function or tanh function 。 6. A continuous voice interaction system, characterized in that, The system includes: The acquisition unit is used to acquire the first voice data input by the target object; The processing unit is used to identify the first voice data as a wake-up voice and then send the preset data in the cache to the terminal's voice activity detection VAD engine. The VAD engine is used to continuously monitor the dialogue between the target object and the vehicle system, and process the dialogue; after detecting that the target object has finished speaking; The processing unit is also used to send the preset data in the cache back to the terminal's VAD engine to achieve continuous voice interaction; The processing unit is specifically used to start a timer after the target object finishes speaking. The timer is used to stop when the target object's voice data is received again. The first duration of the timer is obtained. If the first duration is greater than the time threshold, it is determined that the continuous voice termination condition is met. The processing unit is also used to receive the second voice data input by the target object again after the target object has finished speaking, and to determine that the continuous voice end condition is met when the second voice data is identified as belonging to a specific voice that ends continuous voice interaction. The processing unit is further configured to use an RNN recognition algorithm to determine x confidence rates for x words corresponding to each pronunciation group in the second speech data; use an LSTM recognition algorithm to determine y confidence rates for y words corresponding to each pronunciation group in the speech data; add the two confidence rates of the same first word among the x words and y words in the first pronunciation group to obtain the confidence rate sum of the first word; traverse the same words among the x words and y words to obtain the confidence rate sum of each same word; determine the word corresponding to the maximum confidence rate sum as the word corresponding to the first pronunciation group; traverse all pronunciation groups to obtain the words corresponding to all pronunciation groups; assemble all the words into text data of the second speech data in chronological order; if the text data belongs to a specific text, determine that the second speech data belongs to a specific speech that ends continuous speech interaction.
7. A computer-readable storage medium, characterized in that, A computer program for storing electronic data interchange is provided, wherein the computer program causes a computer to perform the method as described in any one of claims 1-5.