Voice interaction method, device, electronic device, storage medium and program product
By acquiring real-time data acquisition during voice front-end point detection and using large language models, the problem of low speech interaction accuracy is solved, and multimodal interaction with high accuracy and fast response is achieved.
Patent Information
- Application Number
- CN202510637530.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The accuracy of voice interaction in the prior art leads to poor user interaction experience, especially when there is lag in obtaining other modal data, it is impossible to accurately answer user questions.
When the voice front endpoint of the audio data is detected, the collected data consistent with the timestamp is obtained, including real-time acquisition and control of the acquisition device to collect data, ensuring the real-time and accuracy of the data, and using a large language model for comprehensive understanding and analysis.
It improves the accuracy and response speed of voice interaction, improves the user interaction experience, and achieves high-accuracy multimodal interaction.
Smart Images

Figure CN120164465B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction technology, and in particular to a voice interaction method, device, electronic device, storage medium and program product. Background Art
[0002] Voice interaction is a convenient method of human-computer interaction in smart cars. Currently, user questions are becoming increasingly diverse, extending beyond simple vehicle control and navigation functions. For example, users may ask what brand or model of a car is outside the vehicle, or what a building is. Therefore, in addition to analyzing and understanding user voice data, other modal data, such as images outside the vehicle, is also needed to understand and answer the question. Therefore, multimodal interaction is necessary.
[0003] Currently, multimodal interaction is achieved by first acquiring user voice data, then determining the other modal data to be collected based on the user voice data, and then determining the answer based on the user voice data and the other modal data. However, the other modal data used to determine the answer in the existing technology has a lag, that is, the other modal data is not accurate, resulting in a decrease in the accuracy of voice interaction, which in turn affects the user interaction experience. For example, if the transcribed text of the user voice data is "What scenic spot did we just pass by?", if the other modal data to be acquired is an image outside the vehicle, the image outside the vehicle is collected after the user has said the complete sentence "What scenic spot did we just pass by?" and after voice recognition and voice understanding are performed on the user voice data. Therefore, the collected image outside the vehicle may no longer cover the scenic spot, resulting in an inability to accurately obtain the answer. Summary of the Invention
[0004] The present invention provides a voice interaction method, device, electronic device, storage medium and program product, which are used to solve the defect of low accuracy of voice interaction in the prior art and realize a high-accuracy multimodal interaction solution.
[0005] The present invention provides a voice interaction method, comprising:
[0006] When a voice front-end point of the audio data is detected, acquiring collected data consistent with a timestamp of the voice front-end point; the voice front-end point represents the starting time of input of the user's voice data, and the collected data is used to assist in understanding the problem represented by the user's voice data;
[0007] Based on the user voice data and the collected data, an answer result corresponding to the user voice data is determined.
[0008] According to a voice interaction method provided by the present invention, when a voice front-end point of audio data is detected, acquiring collected data consistent with a timestamp of the voice front-end point includes:
[0009] Upon detecting a voice front-end point of audio data input in real time, acquiring first collected data consistent with a timestamp of the voice front-end point from the collected data, and / or controlling a collection device to collect second collected data consistent with a timestamp of the voice front-end point, and acquiring the second collected data;
[0010] Wherein, when the first collected data is obtained, the collected data includes the first collected data; when the second collected data is obtained, the collected data includes the second collected data.
[0011] According to a voice interaction method provided by the present invention, determining an answer result corresponding to the user voice data based on the user voice data and the collected data includes:
[0012] If the intention recognition result of the user voice data is a preset intention recognition result, comprehensively determining an answer result corresponding to the user voice data based on the user voice data and the collected data;
[0013] After acquiring the collected data consistent with the timestamp of the voice front-end point, the method further includes:
[0014] In a case where the intention recognition result of the user voice data is not a preset intention recognition result, an answer result corresponding to the user voice data is determined based on the user voice data.
[0015] According to a voice interaction method provided by the present invention, after acquiring the collected data consistent with the timestamp of the voice front-end point, the method further includes:
[0016] When it is determined based on the intention recognition result of the user voice data that the collected data is not needed, performing data post-processing on the collected data;
[0017] The data post-processing includes at least one of the following:
[0018] In the case that there is data that needs to be pre-processed in the collected data, canceling the data pre-processing process of the collected data;
[0019] Delete the deletable data in the collected data.
[0020] According to a voice interaction method provided by the present invention, the voice interaction method is applied to a processor in a car;
[0021] In the case where a voice front-end point of the audio data is detected, before acquiring the collected data consistent with the timestamp of the voice front-end point, the method further includes:
[0022] When the vehicle is in a voice wake-up state and the vehicle is in a driving state, collecting audio data in real time;
[0023] Perform voice front-end detection on the audio data.
[0024] According to a voice interaction method provided by the present invention, the user voice data is determined based on the following method:
[0025] When a voice front-end point of the audio data is detected, the voice front-end point is used as the starting acquisition moment to acquire input voice data frames until the endpoint stops acquiring after the voice is detected;
[0026] Determining the user voice data based on the collected voice data frames;
[0027] The voice end point represents the termination time of the user voice data input.
[0028] The present invention also provides a voice interaction device, comprising:
[0029] A data acquisition module, configured to, upon detecting a voice front-end point in audio data, acquire collected data consistent with a timestamp of the voice front-end point; the voice front-end point represents the starting time of input of the user's voice data, and the collected data is used to assist in understanding the problem represented by the user's voice data;
[0030] The answer determination module is used to determine the answer result corresponding to the user voice data based on the user voice data and the collected data.
[0031] The present invention also provides a movable device, comprising:
[0032] An audio acquisition device, wherein the audio acquisition device is used to acquire audio data;
[0033] A processor, wherein the processor is used to execute any of the voice interaction methods described above.
[0034] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the voice interaction method described above is implemented.
[0035] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described voice interaction methods.
[0036] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned voice interaction methods.
[0037] The voice interaction method, device, electronic device, storage medium and program product provided by the present invention, when the voice front-end point of the audio data is detected, obtains the collected data consistent with the timestamp of the voice front-end point, and the voice front-end point represents the starting input time of the user voice data, thereby ensuring that the collected data consistent with the timestamp of the voice front-end point is immediately obtained when the voice front-end point of the audio data is detected, thereby avoiding the lag of the collected data, ensuring that the answer result corresponding to the user voice data is accurately determined based on the user voice data and the real-time collected data, thereby improving the accuracy of the voice interaction and further improving the user interaction experience; and when the voice front-end point of the audio data is detected, the collected data consistent with the timestamp of the voice front-end point is obtained, thereby timely preparing the collected data for subsequent timely determination of the answer result based on the collected data, thereby improving the response speed of the voice interaction; at the same time, the collected data is used to assist in understanding the problem represented by the user voice data, so based on the user voice data and the collected data, the answer result corresponding to the user voice data can be more accurately determined. In summary, the present invention can output the answer result corresponding to the user voice data in a timely and accurate manner, thereby realizing a highly accurate multimodal interaction solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 This is one of the flow charts of the voice interaction method provided by the present invention.
[0040] Figure 2 This is the second flow chart of the voice interaction method provided by the present invention.
[0041] Figure 3 This is the third flow chart of the voice interaction method provided by the present invention.
[0042] Figure 4 This is the fourth flow chart of the voice interaction method provided by the present invention.
[0043] Figure 5 This is the fifth flow chart of the voice interaction method provided by the present invention.
[0044] Figure 6 This is the sixth flow chart of the voice interaction method provided by the present invention.
[0045] Figure 7It is a structural diagram of the voice interaction device provided by the present invention.
[0046] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0047] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0048] Current voice interaction solutions often rely on complete user voice data to determine the required data for other modalities before collecting them. However, this data is often delayed, resulting in reduced voice interaction accuracy and a poor user experience.
[0049] Given the poor accuracy of current voice interaction solutions and the poor user interaction experience, the present invention conducted research. The initial idea was to immediately collect other modal data after obtaining complete user voice data. However, the present invention's research on this idea found that although this idea can reduce the time required to determine the other modal data to be collected based on complete user voice data, it still requires waiting until complete user voice data is obtained or a voice question and answer instruction is received before collecting other modal data. This results in a lag in the other modal data, resulting in poor accuracy in voice interaction and an inability to meet the real-time problem of "what intersection just passed by".
[0050] In response to the problems existing in the above-mentioned ideas, the present invention has continued to study and finally proposed a voice interaction method. When the voice front-end point of the audio data is detected, the voice interaction method obtains the collected data (other modal data) consistent with the timestamp of the voice front-end point, so that there will be no lag in the other modal data, ensuring that the answer result corresponding to the user voice data is accurately determined based on the user voice data and the collected data, thereby improving the accuracy of voice interaction and further enhancing the user interaction experience.
[0051] Next, the voice interaction method provided by the present invention is introduced through the following embodiments. Figures 1-6 The voice interaction method of the present invention is described.
[0052] The voice interaction method provided by the present invention can be applied to electronic devices, which can be configured according to actual needs. For example, the electronic device can be a mobile vehicle or a mobile robot. In a specific embodiment, the voice interaction method can be applied to a processor of a mobile device, such as a car, to implement the voice interaction method for a smart car.
[0053] Figure 1 This is one of the flow charts of the voice interaction method provided by the present invention, such as Figure 1 As shown, the voice interaction method may include: step 110 and step 120.
[0054] Step 110 : When a voice front-end point of the audio data is detected, acquisition data consistent with a timestamp of the voice front-end point is obtained.
[0055] Here, audio data refers to data containing a voice front-end to be detected. This audio data includes multiple audio data frames, the number of which is not specifically limited. If a voice front-end can be detected in the audio data, the audio data includes both non-voice data frames and voice data frames. If no voice front-end can be detected in the audio data, the audio data includes only non-voice data frames. Non-voice data frames can be either silent data frames (no sound) or noise data frames (background sound), and the last frame of the audio data is typically a voice data frame.
[0056] The audio data is collected. In one embodiment, the audio data can be collected by an audio collection device (e.g., a microphone) and forwarded to the execution entity of the voice interaction method provided by the present invention. It should be understood that if the voice front-end has been detected but the voice back-end has not been detected, audio data may not be collected for voice front-end detection.
[0057] In one specific embodiment, the audio data is input in real time. In other words, the audio data is collected in real time, that is, the voice front-end detection is performed on the audio data at the same time as the audio data is collected, thereby ensuring in a more real-time manner that when the voice front-end of the audio data is detected, the collected data consistent with the timestamp of the voice front-end is immediately obtained, thereby avoiding delays and ensuring that the answer corresponding to the user voice data is accurately determined based on the user voice data and the collected data, thereby improving the accuracy of voice interaction and further enhancing the user interaction experience. For example, the duration of the audio data is 3 seconds, and the user does not ask any questions in the first nearly 3 seconds. It is not until the user begins to ask a question near the third second that the collected data consistent with the timestamp of the voice front-end is immediately obtained. For example, when the user says "just now", the collected data consistent with the timestamp of the voice front-end is immediately obtained.
[0058] In one embodiment, while in the voice-activated state, audio data is collected in real time and voice front-end detection is performed on the audio data. This prevents constant voice front-end detection and conserves computing resources. In other words, voice front-end detection is performed on the input audio data only when the user activates the voice interaction system.
[0059] In one embodiment, if the voice interaction method is applied to a processor in a car, and when the car is in motion (driving state), audio data is collected in real time, and voice front-end detection is performed on the audio data, so that voice front-end detection is performed only when the car is in motion, avoiding continuous detection of the voice front-end, thereby saving computing resources; and considering that the vehicle is stationary and the real-time requirements are not high, the voice interaction method can be executed only when the car is in motion.
[0060] It should be understood that the last frame included in the audio data is the currently input or currently collected audio data frame. As for how long the audio data needs to retain the previous audio data frames, there is no limitation here. It should be noted that the previous audio data frames are retained to facilitate the detection of the voice front-end point.
[0061] Among them, the voice front-end point represents the starting input moment of the user's voice data. In other words, the voice front-end point is the moment of transition from the non-voice data frame to the voice data frame in the audio data. Therefore, the voice front-end point is usually the input moment or acquisition moment of the last frame in the audio data, thereby ensuring that the voice front-end point is the current moment, thereby improving the real-time acquisition of the collected data, thereby improving the accuracy of voice interaction; of course, the voice front-end point may not be the input moment or acquisition moment of the last frame in the audio data. For example, if the transcribed text of the user's voice data is "just passed by some scenic spot", the voice front-end point is the first input moment of "just". In other words, when the user starts speaking, the collected data consistent with the timestamp of the voice front-end point is obtained.
[0062] The user voice data is the voice data spoken by the user and may include multiple voice data frames. For example, the transcribed text of the user voice data may include "What is the name of the intersection just now?", "What is the brand of the car in front?", "What is the model of the car in front?", "What is the building we just passed by?", or "What scenic spot we just passed by?", etc. It should be noted that the user questions addressed by the embodiments of the present invention are mostly real-time questions, that is, questions asked in real time. For such real-time questions, the present invention can accurately perform voice interaction and provide timely responses.
[0063] The starting input time of the user voice data is the input time (collection time) corresponding to the first word in the transcribed text of the user voice data. For example, if the transcribed text of the user voice data is "What is the name of the intersection just now?", the starting input time of the user voice data is the input time of the first word "just now". Of course, the utterance time corresponding to the first word in the transcribed text of the user voice data can also be determined based on the starting input time, and the voice front-end point can be determined based on the utterance time. The voice front-end point represents the starting utterance time of the user voice data, thereby improving the real-time acquisition of collected data and the accuracy of voice interaction.
[0064] The voice front-end point can be detected by a voice endpoint detection method. The voice endpoint detection method may include but is not limited to the following two embodiments.
[0065] In one embodiment, the speech endpoint detection method extracts spectral features from audio data frame by frame. Next, the posterior probabilities of speech and non-speech in each frame are determined based on the extracted spectral features and a pre-built endpoint detection model. Finally, a speech front-end detection result is output based on the posterior probabilities of speech in each frame. The endpoint detection model is trained based on sample audio data and its corresponding speech front-end detection result labels.
[0066] In another embodiment of the speech endpoint detection method, the energy value of each frame in the audio data is determined, and then the speech front-end detection result is output based on the energy value of each frame. For example, the time point corresponding to the audio data frame with an energy value greater than a preset threshold is the speech front-end.
[0067] Here, the collected data refers to data collected by a collection device. For example, the collected data may be images collected by an image collection device, audio collected by an audio collection device, speed collected by a speed sensor, or heart rate collected by a heart rate sensor. The collected data may be one or more, for example, multiple collected data may include images and sounds outside the vehicle.
[0068] The collection time of the collected data that coincides with the timestamp of the voice front-end point is the voice front-end point. Furthermore, the collected data can include data collected at the voice front-end point and data collected within a preset time period before the voice front-end point. For example, if the voice front-end point is 3 seconds, data from 2 to 3 seconds can be collected. This provides more collected data to assist in understanding the question represented by the user's voice data, thereby improving the accuracy of voice interaction. This can also avoid missing useful collected data. For example, if the user's voice data asks "What's the intersection just now?", because the mobile device (e.g., a car) is moving and the intersection just now is quickly missed, it is necessary to collect data from 2 to 3 seconds, thereby improving the accuracy of voice interaction. Of course, the collected data can also include data collected at the voice front-end point, data collected within a first preset time period before the voice front-end point, and data collected within a second preset time period after the voice front-end point, which is not specifically limited in this embodiment of the present invention. The first and second preset time periods can be set according to actual circumstances, for example, both can be 1 second.
[0069] The collected data may include first collected data and / or second collected data.
[0070] The first collected data is obtained from the already collected data; specifically, when a voice front-end point is detected in the audio data, the first collected data is obtained from the already collected data, which is consistent with the timestamp of the voice front-end point. The already collected data is data that can be collected automatically without the need for controlled collection. For example, if it is necessary to collect images outside the vehicle, if the vehicle has a dashcam, onboard camera, or onboard 360-degree panoramic imaging system turned on, since they have already been collected, there is no need to control the collection of images outside the vehicle.
[0071] The second collected data is collected by controlling the collection device. Specifically, when a voice front-end point is detected in the audio data, the collection device is controlled to collect the second collected data that is consistent with the timestamp of the voice front-end point, and the second collected data is obtained. For example, when it is necessary to collect images outside the vehicle, the vehicle camera is controlled to take real-time photos of the environment outside the vehicle to collect images outside the vehicle.
[0072] Furthermore, when a voice front-end point in the audio data is detected, a collected data set is obtained that matches the timestamp of the voice front-end point. Based on the intent recognition result of the user voice data, a number of collected data matching the intent recognition result are determined from the collected data set. This is then used to determine the answer corresponding to the user voice data based on the user voice data and the number of collected data. For example, if the user voice data is "I just passed by a certain scenic spot," and the collected data set includes both exterior vehicle images and exterior vehicle sounds, then only the exterior vehicle images need to be determined from these two data sets to determine the answer. Thus, the answer is determined based solely on the collected data matching the user intent, eliminating the need to determine the answer based on other irrelevant data, thereby improving the efficiency and accuracy of voice interaction. Of course, embodiments of the present invention can be applied to only one scenario, where the required data is known. Instead of determining the number of collected data from the collected data set, the answer corresponding to the user voice data is determined based on the user voice data and the collected data set. For example, if embodiments of the present invention are applied only to scenarios where the user asks questions about the exterior vehicle environment, and all questions are related to exterior vehicle images, then it is known that exterior vehicle images are being collected, eliminating the need for filtering operations.
[0073] The collected data is used to assist in understanding the question represented by the user voice data, thereby facilitating more accurate determination of the answer corresponding to the user voice data, thereby improving the accuracy of voice interaction and further enhancing the user interaction experience.
[0074] For example, if the transcribed text of the user's voice data is "What is the scenic spot we just passed by", the collected data is the image outside the car; if the transcribed text of the user's voice data is "What sound was outside the car just now", the collected data is the sound outside the car; if the transcribed text of the user's voice data is "What expression did I have just now", the collected data is the image inside the car; if the transcribed text of the user's voice data is "What is the current speed of the car", the collected data is the car speed; if the transcribed text of the user's voice data is "What was my heart rate just now", the collected data is the user's heart rate.
[0075] In one specific embodiment, upon detecting the voice front-end of the audio data, collected data consistent with the timestamp of the voice front-end is obtained, thereby ensuring that the collected data consistent with the timestamp of the voice front-end is obtained immediately upon detecting the voice front-end of the audio data in a more real-time manner, thereby avoiding delays and ensuring that the answer corresponding to the user voice data is accurately determined based on the user voice data and the collected data, thereby improving the accuracy of voice interaction and, in turn, enhancing the user interaction experience; and upon detecting the voice front-end of the audio data, collected data consistent with the timestamp of the voice front-end is obtained, thereby promptly preparing the collected data for subsequent timely determination of the answer based on the collected data, thereby improving the responsiveness of the voice interaction. For example, if the audio data is 3 seconds long, and the user does not ask any questions in the first nearly 3 seconds, until the user begins to ask a question near the third second, such as when the user says "just now", collected data consistent with the timestamp of the voice front-end is immediately obtained.
[0076] Furthermore, the collected data is timestamped to ensure the synchronization of the collected data with the user's voice data, to ensure that the answer results corresponding to the user's voice data are obtained more accurately, and to improve the accuracy of voice interaction.
[0077] Step 120: Determine an answer result corresponding to the user voice data based on the user voice data and the collected data.
[0078] Here, the answer result is the answer to the question represented by the user voice data. For example, if the transcribed text of the user voice data is "What is the scenic spot I just passed by", the answer result can be "a certain scenic spot".
[0079] In one specific embodiment, user voice data and collected data are input into a voice interaction model, and the voice interaction model outputs a response. Furthermore, the voice interaction model can be constructed based on a large language model. Furthermore, a transcribed text of the user voice data and collected data are input into the voice interaction model, and the voice interaction model outputs a response.
[0080] The reason for using a large language model (LLM) to build a voice interaction model is that it has richer prior knowledge and reasoning capabilities than traditional small models or pre-trained models. At the same time, compared with implicit information extraction and modeling, the large language model can explicitly output thinking, analysis, and reasoning steps, thereby enhancing the correctness of the final reasoning results at the semantic level. This can improve the robustness of the voice interaction model and thus improve the accuracy of voice interaction.
[0081] Furthermore, the user's voice data is converted into a transcript. Based on the transcript and the collected data, the corresponding answer to the user's voice data is determined. The transcript can better assist in understanding the question represented by the user's voice data, thereby further improving the accuracy of voice interaction.
[0082] It should be understood that comprehensive understanding and analysis based on user voice data and collected data can more accurately determine the answer results corresponding to the user voice data, that is, multimodal understanding and analysis can improve the accuracy of voice interaction.
[0083] It is understandable that, considering that other modal data are collected only after the complete user voice data is obtained, it is mostly used when the vehicle is stationary. In the driving state, the movement of the vehicle will cause inconsistencies and mismatches between the user voice data and the actual scene. Therefore, when the voice front-end point of the audio data is detected, the embodiment of the present invention immediately obtains the collected data consistent with the timestamp of the voice front-end point, avoiding the lag of the collected data, so that the user voice data and the collected data are consistent and matched, thereby improving the accuracy of voice interaction in the driving state. That is, the embodiment of the present invention can provide a voice interaction method that can accurately respond to user needs in the driving state, and improve the real-time and accuracy of multimodal interaction; and the embodiment of the present invention can avoid limiting the use scenarios of voice interaction, that is, it can be used regardless of driving state or stationary state.
[0084] The voice interaction method provided by the embodiment of the present invention obtains collected data consistent with the timestamp of the voice front-end point when the voice front-end point of the audio data is detected, and the voice front-end point represents the starting input time of the user voice data, thereby ensuring that the collected data consistent with the timestamp of the voice front-end point is immediately obtained when the voice front-end point of the audio data is detected, thereby avoiding the lag of the collected data, ensuring that the answer result corresponding to the user voice data is accurately determined based on the user voice data and the real-time collected data, thereby improving the accuracy of the voice interaction and further improving the user interaction experience; and when the voice front-end point of the audio data is detected, the collected data consistent with the timestamp of the voice front-end point is obtained, thereby timely preparing the collected data for subsequent timely determination of the answer result based on the collected data, thereby improving the response speed of the voice interaction; at the same time, the collected data is used to assist in understanding the problem represented by the user voice data, so based on the user voice data and the collected data, the answer result corresponding to the user voice data can be more accurately determined. In summary, the present invention can output the answer result corresponding to the user voice data in a timely and accurate manner, thereby realizing a highly accurate multimodal interaction solution.
[0085] Based on any of the above embodiments, a specific embodiment of the voice interaction method is given below. Figure 2This is the second flow chart of the voice interaction method provided by the present invention, such as Figure 2 As shown, the voice interaction method includes: step 111 and step 120.
[0086] Step 111, while detecting the voice front-end point of the real-time input audio data, obtain the first collected data consistent with the timestamp of the voice front-end point from the collected data, and / or control the collection device to collect the second collected data consistent with the timestamp of the voice front-end point, and obtain the second collected data.
[0087] Here, the audio data is input in real time, so the voice front-end detection is performed at the same time as the audio data is input in real time, and the voice front-end detection is performed as long as there is audio data input, instead of first obtaining a complete audio data and then post-processing it for voice front-end detection, thereby avoiding delays in data collection and improving the accuracy of voice interaction.
[0088] In other words, the last frame of audio data is the current real-time input, meaning the last frame included in the audio data is the currently input or currently collected audio data frame. There's no limit on how long the audio data needs to retain the previous audio data frames. It should be noted that retaining the previous audio data frames facilitates the detection of the speech front-end. For example, if the user's voice data is transcribed as "I just passed by some scenic spot," the speech front-end is the input moment of the first "just," and the corresponding last frame of audio data is the speech front-end. This allows the user to obtain collected data that matches the speech front-end timestamp when they begin speaking.
[0089] In other words, the audio data is collected in real time, that is, the voice front-end point detection is performed on the audio data while the audio data is collected, so as to ensure in a more real-time manner that when the voice front-end point of the audio data is detected, the collected data consistent with the voice front-end point timestamp is immediately obtained, thereby avoiding delays.
[0090] It should be noted that, upon detecting the voice front-end of the real-time input audio data, collected data consistent with the timestamp of the voice front-end is immediately obtained, thereby avoiding lags in the collected data and ensuring that the answer corresponding to the user's voice data is accurately determined based on the user's voice data and the collected data, thereby improving the accuracy of voice interaction and, in turn, the user interaction experience. Furthermore, upon detecting the voice front-end of the audio data, collected data consistent with the timestamp of the voice front-end is obtained, thereby promptly preparing the collected data for subsequent timely determination of the answer based on the collected data, thereby improving the responsiveness of the voice interaction. For example, if the audio data is 3 seconds long, and the user does not ask any questions in the first nearly 3 seconds, until the user begins to ask a question near the third second, such as when the user says "just now," collected data consistent with the timestamp of the voice front-end is immediately obtained.
[0091] Here, the collected data refers to data collected by a collection device. For example, the collected data may be images collected by an image collection device, audio collected by an audio collection device, speed collected by a speed sensor, heart rate collected by a heart rate sensor, and so on. The number of collected data may be one or more. For example, the multiple collected data may include images outside the vehicle and sounds outside the vehicle. The collected data refers to data that can be collected automatically without the need for controlled collection. For example, if the vehicle has already turned on a dashcam, on-board camera, or on-board 360-degree panoramic imaging system, then since these images have already been collected, there is no need to control the collection of images outside the vehicle.
[0092] The first collected data, whose timestamp matches the voice front-end point, is collected at the time of collection. Furthermore, the first collected data may include data collected at the voice front-end point and data collected within a preset time period before the voice front-end point. For example, if the voice front-end point is 3 seconds, data from 2 to 3 seconds may be collected. This provides more collected data to assist in understanding the question represented by the user's voice data, thereby improving the accuracy of voice interaction. This also avoids missing useful collected data. For example, if the user's voice data asks "What was the intersection just now?", because the mobile device (e.g., a car) is moving and the intersection just now was quickly missed, data from 2 to 3 seconds may be collected, thereby improving the accuracy of voice interaction. Of course, the first collected data may also include data collected at the voice front-end point, data collected within a first preset time period before the voice front-end point, and data collected within a second preset time period after the voice front-end point, which is not specifically limited in this embodiment of the present invention. Accordingly, the collected data needs to include data of the same type as the first collected data.
[0093] Here, the acquisition device is used to acquire the second acquired data. The acquisition device can be configured based on the type of the second acquired data and is not specifically limited herein. For example, if the second acquired data is an image, the acquisition device is an image acquisition device; if the second acquired data is audio, the acquisition device is an audio acquisition device; if the second acquired data is speed, the acquisition device is a speed sensor; and if the second acquired data is heart rate, the acquisition device is a heart rate sensor.
[0094] Here, the second collected data is collected by controlling the collection device. For example, when it is necessary to collect images outside the vehicle, the vehicle-mounted camera is controlled to take real-time photos of the environment outside the vehicle to collect images outside the vehicle.
[0095] The collection time of the second collected data that is consistent with the timestamp of the voice front-end point is the voice front-end point. Furthermore, the second collected data can include data collected by the voice front-end point and data collected within a preset time length before the voice front-end point. For example, if the voice front-end point is the 3rd second, data from the 2nd to 3rd seconds can be collected, thereby providing more collected data to assist in understanding the problem represented by the user voice data, thereby improving the accuracy of voice interaction; and it can avoid missing useful collected data. For example, the user voice data is "What is the intersection just now?" Because the car is moving, the intersection just now is quickly missed, so it is necessary to collect data from the 2nd to 3rd seconds, thereby improving the accuracy of voice interaction. Of course, the second collected data can also include data collected by the voice front-end point, data collected within the first preset time length before the voice front-end point, and data collected within the second preset time length after the voice front-end point. This embodiment of the present invention does not specifically limit this.
[0096] It should be noted that, since the second collected data is collected by the collection device, the execution subject of the voice interaction method provided by the present invention needs to obtain the second collected data collected by the collection device, that is, the collection device forwards the second collected data to the execution subject.
[0097] Wherein, when the first collected data is obtained, the collected data includes the first collected data, and when the second collected data is obtained, the collected data includes the second collected data. It should be understood that the collected data may include the first collected data and / or the second collected data.
[0098] For example, if the transcribed text of the user voice data is “What is the scenic spot we just passed by?”, the collected data is the image outside the car, and it can be the first collected data or the second collected data; if the transcribed text of the user voice data is “What sound was outside the car just now?”, the collected data is the sound outside the car, and it can be the first collected data or the second collected data; if the transcribed text of the user voice data is “What expression did I have just now?”, the collected data is the image inside the car, and it can be the first collected data or the second collected data; if the transcribed text of the user voice data is “What is the current speed of the car?”, the collected data is the car speed, and it can be the first collected data or the second collected data; if the transcribed text of the user voice data is “What was my heart rate just now?”, the collected data is the user's heart rate, and it can be the first collected data or the second collected data.
[0099] Step 120: Determine an answer result corresponding to the user voice data based on the user voice data and the collected data.
[0100] Specifically, the answer result is determined based on the user voice data, and the first collected data and / or the second collected data.
[0101] The voice interaction method provided by the embodiment of the present invention, when detecting the voice front-end point of the real-time input audio data, obtains the first collected data consistent with the timestamp of the voice front-end point from the collected data, and / or controls the collection device to collect the second collected data consistent with the timestamp of the voice front-end point, and the voice front-end point represents the starting input time of the user voice data, thereby ensuring that when the voice front-end point of the real-time input audio data is detected, the collected data consistent with the timestamp of the voice front-end point is immediately obtained, thereby avoiding the delay of the collected data, ensuring that the answer result corresponding to the user voice data is accurately determined based on the user voice data and the real-time collected data, thereby improving the accuracy of the voice interaction and further improving the user interaction experience; and when the voice front-end point of the real-time input audio data is detected, the collected data consistent with the timestamp of the voice front-end point is obtained, thereby preparing the collected data in time for the subsequent timely determination of the answer result based on the collected data, thereby improving the response speed of the voice interaction.
[0102] Based on any of the above embodiments, a specific embodiment of the voice interaction method is given below. Figure 3 This is the third flow chart of the voice interaction method provided by the present invention, such as Figure 3 As shown, the voice interaction method includes: step 110, step 121 and step 130. Step 130 is after step 110.
[0103] Step 110 : When a voice front-end point of the audio data is detected, acquisition data consistent with a timestamp of the voice front-end point is obtained.
[0104] Specifically, when a voice front-end is detected in the audio data, user voice data is collected. This user voice data is the voice data spoken by the user, and the user voice data may include multiple voice data frames. Based on the user voice data, the user's complete question can be determined, ensuring that an accurate answer is subsequently obtained.
[0105] Step 121, when the intention recognition result of the user voice data is a preset intention recognition result, based on the user voice data and the collected data, comprehensively determine the answer result corresponding to the user voice data.
[0106] Here, the intent recognition result is obtained by performing intent recognition on the user voice data. Further, the intent recognition result is obtained by performing intent recognition on the transcribed text of the user voice data.
[0107] In one embodiment, speech recognition is performed on user speech data to obtain a transcribed text of the user speech data, and semantic understanding is performed on the transcribed text to obtain an intent recognition result. Both speech recognition and semantic understanding can be achieved through artificial intelligence models.
[0108] In another embodiment, user voice data is input into an intent recognition model to obtain an intent recognition result output by the intent recognition model. The intent recognition model can be constructed based on a large language model, or trained on a self-constructed initial model based on sample user voice data and its corresponding intent recognition result labels.
[0109] Here, the preset intent recognition results are pre-set, for example, the intent to inquire about information about buildings outside the vehicle, the intent to inquire about information about scenic spots outside the vehicle, the intent to inquire about information about cars outside the vehicle, and so on. These preset intent recognition results are all corresponding to intents that require multimodal interaction, that is, intents that require collected data (other modal data) to determine the answer. The number of preset intent recognition results can be one or more. For example, if there are multiple intents that all require multimodal interaction, the number of preset intent recognition results will be multiple.
[0110] Furthermore, considering that not all collected data is relevant to the user's question, that is, not all collected data is relevant to the intent recognition results, embodiments of the present invention, based on the intent recognition results of the user's voice data, filter out target data from the collected data that matches the question represented by the user's voice data. Based on the user's voice data and the target data, a corresponding answer is comprehensively determined. This avoids determining an answer based on irrelevant collected data, thereby improving the efficiency and accuracy of voice interaction.
[0111] It should be understood that based on the user voice data and collected data, comprehensively determining the answer results corresponding to the user voice data, that is, conducting comprehensive understanding and analysis, can more accurately determine the answer results corresponding to the user voice data, that is, conducting multimodal understanding and analysis, which can improve the accuracy of voice interaction.
[0112] Step 130: When the intention recognition result of the user voice data is not a preset intention recognition result, determine an answer result corresponding to the user voice data based on the user voice data.
[0113] The preset intent recognition results are all intentions corresponding to multimodal interaction. If the intent recognition result of the user's voice data is not the preset intent recognition result, it means that no other modal data (collected data) is required. Therefore, it is only necessary to determine the answer result corresponding to the user's voice data based on the user's voice data, thereby improving the efficiency of voice interaction and improving the accuracy of voice interaction.
[0114] For example, if the transcribed text of the user voice data is "What time is it now?", then its intent recognition result is not a preset intent recognition result, so it is only necessary to determine the answer result based on the user voice data.
[0115] In one specific embodiment, user voice data is input into a voice interaction model, and the voice interaction model outputs a response. Furthermore, the voice interaction model can be constructed based on a large language model. Furthermore, a transcribed text of the user voice data is input into the voice interaction model, and the voice interaction model outputs a response.
[0116] Large language models are used to build voice interaction models because they possess richer prior knowledge and reasoning capabilities than traditional small or pre-trained models. Furthermore, compared to implicit information extraction and modeling, large language models can explicitly output thinking, analysis, and reasoning steps, thereby enhancing the correctness of the final reasoning results at the semantic level. This improves the robustness of the voice interaction model and, consequently, the accuracy of voice interaction.
[0117] Furthermore, the user's voice data is converted into a transcript, and based on the transcript, the corresponding answer to the user's voice data is determined. The transcript can better assist in understanding the question represented by the user's voice data, thereby further improving the accuracy of voice interaction.
[0118] The voice interaction method provided by an embodiment of the present invention, when the intention recognition result of the user voice data is a preset intention recognition result, comprehensively determines the answer result corresponding to the user voice data based on the user voice data and the collected data, and can more accurately determine the answer result corresponding to the user voice data, that is, perform multimodal understanding and analysis, which can improve the accuracy of voice interaction; and when the intention recognition result of the user voice data is not a preset intention recognition result, the answer result corresponding to the user voice data is determined based on the user voice data, so that the answer result corresponding to the user voice data can be determined based on the user voice data without the need for other modal data (collected data), thereby improving the efficiency of voice interaction and improving the accuracy of voice interaction.
[0119] Based on any of the above embodiments, another embodiment of the voice interaction method is given below. Figure 4 This is the fourth flow chart of the voice interaction method provided by the present invention, such as Figure 4 As shown, the voice interaction method includes: step 110, step 120 and step 140. Step 140 is after step 110.
[0120] Step 110 : When a voice front-end point of the audio data is detected, acquisition data consistent with a timestamp of the voice front-end point is obtained.
[0121] Here, the collected data refers to data collected by a collection device. For example, the collected data may be images collected by an image collection device, which typically requires data preprocessing; or audio collected by an audio collection device, which typically requires data preprocessing; or speed collected by a speed sensor, or heart rate collected by a heart rate sensor, etc. Furthermore, to improve the real-time nature of voice interaction, after obtaining the collected data that matches the timestamp of the voice front-end, data preprocessing is immediately performed, so that the answer can be quickly and directly determined based on the preprocessed collected data.
[0122] It should be understood that after the collected data is acquired, it usually needs to be cached or stored in a storage space for use in subsequent steps.
[0123] Step 120: Determine an answer result corresponding to the user voice data based on the user voice data and the collected data.
[0124] In a specific embodiment, if the collected data requires data preprocessing, the answer result corresponding to the user voice data is determined based on the user voice data and the collected data after data preprocessing.
[0125] Step 140 : When it is determined that the collected data is not needed based on the intention recognition result of the user voice data, post-process the collected data.
[0126] Here, the intent recognition result is obtained by performing intent recognition on the user voice data. Further, the intent recognition result is obtained by performing intent recognition on the transcribed text of the user voice data.
[0127] Considering that not all collected data is relevant to user questions, that is, not all collected data is relevant to intent recognition results, some intent recognition results do not require data collection, and therefore require data post-processing. This is necessary because even if the voice front-end of the audio data is detected, the complete user voice data has not yet been obtained, making it impossible to determine whether data collection is necessary. In other words, regardless of whether it is necessary, collected data that matches the timestamp of the voice front-end is obtained first.
[0128] In a specific embodiment, when the intention recognition result of the user voice data is not a preset intention recognition result, it is determined that data collection is not necessary.
[0129] The preset intent recognition results are pre-set, for example, the intent to inquire about information about buildings outside the vehicle, the intent to inquire about information about scenic spots outside the vehicle, the intent to inquire about information about cars outside the vehicle, and so on. These preset intent recognition results are all corresponding to intents that require multimodal interaction, that is, they require collected data (other modal data) to determine the answer. Because these preset intent recognition results are all corresponding to intents that require multimodal interaction, if the intent recognition result of the user's voice data is not a preset intent recognition result, it indicates that no other modal data (collected data) is required.
[0130] For example, the preset intention recognition result is the intention to inquire about the information of buildings outside the vehicle, and the collected data includes images outside the vehicle. If the intention recognition result of the user voice data is not the intention to inquire about the information of buildings outside the vehicle, the collected data is post-processed (such as deleting the images outside the vehicle and pausing the image preprocessing process).
[0131] The data post-processing includes at least one of the following: canceling the data pre-processing process of the collected data when there is data requiring data pre-processing in the collected data; deleting deletable data in the collected data.
[0132] Taking into account that some collected data need to be preprocessed before the subsequent answer results can be determined, while some collected data do not need data preprocessing, therefore, when there is data that needs data preprocessing in the collected data, the data preprocessing process of the collected data is canceled to avoid wasting computing resources, thereby improving the efficiency of voice interaction.
[0133] For example, if the collected data is image data, it may need to be preprocessed; if the collected data is speed data, it may not need to be preprocessed.
[0134] The specific method of data preprocessing is related to the data type of the collected data and is not specifically limited here. For example, if the collected data is image data, the corresponding data preprocessing method may include but is not limited to: image cropping, denoising, etc.
[0135] Canceling the data preprocessing process of the collected data can be understood as canceling the data preprocessing process when the collected data is about to start data preprocessing, that is, not performing the data preprocessing process; or canceling the data preprocessing process when the collected data is already in the data preprocessing process, that is, interrupting the data preprocessing process, and not performing the data preprocessing process subsequently.
[0136] Considering that not all collected data can be deleted, only deletable data is deleted. For example, dashcams and in-vehicle 360-degree panoramic imaging systems do not require deletion due to their functional requirements. Therefore, deleting deletable data from the collected data can reduce storage space usage and avoid wasting storage resources. It is understood that the deletable data is only the data collected and stored for the voice interaction method.
[0137] In addition, users can also perform post-processing on the collected data through voice. For example, although a user asks a question, he or she may cancel it immediately.
[0138] The voice interaction method provided by an embodiment of the present invention cancels the data preprocessing process of the collected data when it is determined based on the intention recognition result of the user voice data that data collection is not necessary, and when there is data that needs to be preprocessed in the collected data, thereby avoiding wasting computing resources and improving the efficiency of voice interaction; at the same time, when it is determined based on the intention recognition result of the user voice data that data collection is not necessary, deleting the deletable data in the collected data can reduce the storage space occupied, thereby avoiding wasting storage resources and providing more resources to where the voice interaction method really needs them, thereby ensuring the stability of the voice interaction.
[0139] Based on any of the above embodiments, another embodiment of the voice interaction method is given below. Figure 5 This is the fifth flow chart of the voice interaction method provided by the present invention, such as Figure 5 As shown, the voice interaction method includes: step 510, step 520, step 110 and step 120.
[0140] This voice interaction method is applied to the processor in the car. Based on this, voice interaction capabilities and multimodal interaction capabilities can be realized in smart cars.
[0141] Step 510 : When the vehicle is in a voice awakening state and the vehicle is in a driving state, audio data is collected in real time.
[0142] Here, the voice wake-up state refers to the user waking up the voice interaction system. For example, the user speaks a preset wake-up word to wake up the voice interaction system. The preset wake-up word is a predefined wake-up word, such as "Xiao Fei, Xiao Fei." It should be understood that after waking up the voice interaction, it is in the voice wake-up state, during which it can receive user voice data for voice interaction.
[0143] It can be understood that real-time audio data collection, i.e., the execution of the voice interaction method, is performed only when the system is in the voice wake-up state, thus avoiding constant audio data collection and voice front-end detection, thereby avoiding wasting resources and reducing device power consumption. In other words, voice front-end detection is performed only when the system is in the voice wake-up state, avoiding constant voice front-end detection and thus saving computing resources; that is, voice front-end detection is performed on the input audio data only when the user wakes up the voice interaction system.
[0144] Here, "driving state" refers to the vehicle's driving (mobile) state. Therefore, real-time audio data is collected, and the voice interaction method is executed, only when the vehicle is driving. This allows responses to be determined directly based on user voice data when the vehicle is stationary, thereby improving the efficiency of voice interaction in stationary states. Furthermore, performing voice front-end detection only when the vehicle is driving avoids constant voice front-end detection, thereby conserving computing resources. Furthermore, since real-time performance is less critical when the vehicle is stationary, the voice interaction method can be executed only when the vehicle is driving.
[0145] Here, audio data is collected in real time, so voice front-end detection is performed at the same time as the audio data is collected in real time. As long as there is audio data collection, voice front-end detection is performed, instead of first obtaining a complete audio data and then post-processing it for voice front-end detection, thereby avoiding delays in collecting data and improving the accuracy of voice interaction.
[0146] In other words, the last frame of audio data is currently being collected in real time. This means the last frame included in the audio data is the currently collected audio data frame. There's no limit on how long the audio data needs to retain the previous audio data frames. It should be noted that retaining the previous audio data frames facilitates the detection of the speech front-end. For example, if the user's voice data is transcribed as "I just passed by some scenic spot," the speech front-end is the moment the first "just" is input. The corresponding last frame of audio data is collected at the speech front-end, allowing the user to obtain data that matches the timestamp of the speech front-end when they begin speaking.
[0147] In other words, the audio data is input in real time, so that the voice front-end point of the audio data is detected at the same time as the audio data is collected, thereby ensuring in more real time that when the voice front-end point of the audio data is detected, the collected data consistent with the voice front-end point timestamp is immediately obtained, thereby avoiding delays in the collected data and improving the accuracy of voice interaction.
[0148] Step 520: Perform voice front-end detection on the audio data.
[0149] The voice front-end point can be detected by a voice endpoint detection method, which can refer to the above two embodiments.
[0150] Step 110 : When a voice front-end point of the audio data is detected, acquisition data consistent with a timestamp of the voice front-end point is obtained.
[0151] After the above-mentioned voice front-end detection of the audio data, when the voice front-end of the audio data is detected, the collected data consistent with the timestamp of the voice front-end is obtained, thereby ensuring in real time that when the voice front-end of the audio data is detected, the collected data consistent with the timestamp of the voice front-end is immediately obtained, thereby avoiding delays in the collected data, ensuring that the answer result corresponding to the user voice data is accurately determined based on the user voice data and the collected data, thereby improving the accuracy of voice interaction and further enhancing the user interaction experience; and when the voice front-end of the audio data is detected, the collected data consistent with the timestamp of the voice front-end is obtained, thereby preparing the collected data in time for subsequent timely determination of the answer result based on the collected data, that is, improving the response speed of voice interaction.
[0152] Step 120: Determine an answer result corresponding to the user voice data based on the user voice data and the collected data.
[0153] Specifically, based on the user voice data and the collected data, the answer result when the car is in the voice awakening state and the car is in the driving state is determined.
[0154] The voice interaction method provided in an embodiment of the present invention collects audio data and performs voice front-end detection only when the vehicle is in a voice wake-up state, avoiding continuous detection of the voice front-end and thus saving computing resources; and collects audio data and performs voice front-end detection only when the vehicle is in a driving state, avoiding continuous detection of the voice front-end and thus saving computing resources.
[0155] Based on any of the above embodiments, another embodiment of the voice interaction method is given below. Figure 6 This is the sixth flow chart of the voice interaction method provided by the present invention, such as Figure 6As shown, the user voice data is determined based on the following steps.
[0156] Step 610: When a voice front-end point of the audio data is detected, the voice front-end point is used as the starting acquisition moment to acquire input voice data frames until the voice back-end point is detected and acquisition stops.
[0157] Here, the voice front end refers to the moment when the user's voice data starts entering the input. In other words, the voice front end is the moment when the audio data transitions from a non-voice data frame to a voice data frame. Therefore, the voice front end is used as the starting point for collection, and the collected voice data frames are all collected until the voice back end is detected.
[0158] The starting collection time is the time at which the user's voice data is collected, i.e., the voice front-end node indicates the time at which the user's voice data is collected. It is understood that the starting collection time of the user's voice data is the collection time corresponding to the first character in the transcribed text of the user's voice data. For example, if the transcribed text of the user's voice data is "What is the name of the intersection just now?", the starting collection time of the user's voice data is the collection time of the first character "just now".
[0159] The post-speech endpoint represents the moment when the user's voice data input terminates. In other words, the post-speech endpoint is the moment when the voice data frame transitions to the non-voice data frame. Therefore, the post-speech endpoint is typically the input or acquisition moment of the last frame in the user's voice data, thereby ensuring the accuracy of the user's voice data and, in turn, improving the accuracy of voice interaction. For example, if the transcribed text of the user's voice data is "Just passed by what scenic spot," the post-speech endpoint is the input moment of "region." In other words, the moment the user finishes speaking is the post-speech endpoint.
[0160] The speech endpoint can be detected by a speech endpoint detection method, which can refer to the above two embodiments.
[0161] Step 620: Determine the user voice data based on the collected voice data frames.
[0162] Here, user voice data refers to the voice data spoken by the user, and includes individual voice data frames. For example, the transcribed text of the user voice data may include "What is the name of the intersection just now?", "What is the brand of the car in front?", "What is the model of the car in front?", "What is the name of the building we just passed?", "What is the name of the scenic spot we just passed?", etc.
[0163] The voice interaction method provided by the embodiment of the present invention, when the voice front-end point of the audio data is detected, immediately uses the voice front-end point as the starting collection moment to immediately collect the input voice data frame until the endpoint after the voice is detected stops collecting, and then the user voice data can be obtained when the last frame of the user voice data is collected (for example, the user voice data can be obtained immediately after the user finishes asking the question), so that the answer result can be quickly determined based on the user voice data. Compared with the existing method of first obtaining the complete audio data and then performing endpoint detection on the complete audio data to obtain the user voice data, the present invention will not have lag, can improve the real-time performance of voice interaction, and improve the response speed of voice interaction.
[0164] The voice interaction device provided by the present invention is described below. The voice interaction device described below and the voice interaction method described above can be referenced to each other.
[0165] Figure 7 This is a schematic diagram of the structure of the voice interaction device provided by the present invention. Figure 7 As shown, the voice interaction device includes a data acquisition module 710 and an answer determination module 720.
[0166] The data acquisition module 710 is used to obtain collected data consistent with the timestamp of the voice front-end point when a voice front-end point of the audio data is detected; the voice front-end point represents the starting input moment of the user voice data, and the collected data is used to assist in understanding the problem represented by the user voice data.
[0167] The answer determination module 720 is used to determine the answer result corresponding to the user voice data based on the user voice data and the collected data.
[0168] The voice interaction device provided by the embodiment of the present invention obtains collected data consistent with the timestamp of the voice front-end point when the voice front-end point of the audio data is detected, and the voice front-end point indicates the starting input time of the user voice data, thereby ensuring that the collected data consistent with the timestamp of the voice front-end point is immediately obtained when the voice front-end point of the audio data is detected, thereby avoiding the delay of the collected data, ensuring that the answer result corresponding to the user voice data is accurately determined based on the user voice data and the real-time collected data, thereby improving the accuracy of the voice interaction and further enhancing the user interaction experience; and when the voice front-end point of the audio data is detected, the collected data consistent with the timestamp of the voice front-end point is obtained, thereby timely preparing the collected data for subsequent timely determination of the answer result based on the collected data, thereby improving the response speed of the voice interaction; at the same time, the collected data is used to assist in understanding the problem represented by the user voice data, so based on the user voice data and the collected data, the answer result corresponding to the user voice data can be more accurately determined. In summary, the present invention can output the answer result corresponding to the user voice data in a timely and accurate manner.
[0169] Based on any of the above embodiments, the data acquisition module 710 is specifically configured to:
[0170] Upon detecting a voice front-end point of audio data input in real time, acquiring first collected data consistent with a timestamp of the voice front-end point from the collected data, and / or controlling a collection device to collect second collected data consistent with a timestamp of the voice front-end point, and acquiring the second collected data;
[0171] Wherein, when the first collected data is obtained, the collected data includes the first collected data; when the second collected data is obtained, the collected data includes the second collected data.
[0172] Based on any of the above embodiments, the answer determination module 720 is specifically used to: when the intention recognition result of the user voice data is a preset intention recognition result, comprehensively determine the answer result corresponding to the user voice data based on the user voice data and the collected data.
[0173] The answer determination module 720 is also used to: when the intention recognition result of the user voice data is not a preset intention recognition result, determine the answer result corresponding to the user voice data based on the user voice data.
[0174] Based on any of the above embodiments, the device further includes a data post-processing module, which is configured to:
[0175] When it is determined based on the intention recognition result of the user voice data that the collected data is not needed, performing data post-processing on the collected data;
[0176] The data post-processing includes at least one of the following:
[0177] In the case that there is data that needs to be pre-processed in the collected data, canceling the data pre-processing process of the collected data;
[0178] Delete the deletable data in the collected data.
[0179] Based on any of the above embodiments, the voice interaction method is applied to a processor in a car; the device further includes:
[0180] A data acquisition module, configured to collect audio data in real time when the vehicle is in a voice awakening state and the vehicle is in a driving state;
[0181] The endpoint detection module is used to perform voice front-end detection on the audio data.
[0182] Based on any of the above embodiments, the user voice data is determined based on the following data determination module, and the data determination module is used to:
[0183] When a voice front-end point of the audio data is detected, the voice front-end point is used as the starting acquisition moment to acquire input voice data frames until the endpoint stops acquiring after the voice is detected;
[0184] Determining the user voice data based on the collected voice data frames;
[0185] The voice end point represents the termination time of the user voice data input.
[0186] The mobile device provided by the present invention is described below. The mobile device described below and the voice interaction method described above can be referenced to each other.
[0187] The present invention provides a mobile device comprising an audio acquisition device and a processor. The audio acquisition device is configured to acquire audio data; the processor is configured to execute the voice interaction method described in any of the above embodiments. For example, the mobile device may be a mobile device such as a car or a mobile robot.
[0188] Figure 8 An example of a physical structure diagram of an electronic device is shown below. Figure 8As shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call logic instructions in the memory 830 to execute a voice interaction method, which includes: upon detecting a voice front-end point of audio data, obtaining collected data consistent with a timestamp of the voice front-end point; the voice front-end point indicates the starting input moment of the user voice data, and the collected data is used to assist in understanding the question represented by the user voice data; and determining an answer result corresponding to the user voice data based on the user voice data and the collected data.
[0189] Furthermore, the logic instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0190] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the voice interaction method provided by the above methods, the method including: when a voice front-end point of audio data is detected, obtaining collected data consistent with the timestamp of the voice front-end point; the voice front-end point represents the starting input moment of the user voice data, and the collected data is used to assist in understanding the problem represented by the user voice data; based on the user voice data and the collected data, determining the answer result corresponding to the user voice data.
[0191] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the voice interaction method provided by the above-mentioned methods, the method comprising: when a voice front-end point of audio data is detected, obtaining collected data consistent with the timestamp of the voice front-end point; the voice front-end point represents the starting input moment of the user voice data, and the collected data is used to assist in understanding the problem represented by the user voice data; based on the user voice data and the collected data, determining the answer result corresponding to the user voice data.
[0192] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0193] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A voice interaction method, characterized in that: include: When a voice front-end point of real-time input audio data is detected, immediately acquiring collected data consistent with the timestamp of the voice front-end point, and immediately performing data preprocessing on the collected data if data preprocessing is required; The voice front-end point represents the starting input moment of the user's voice data, and the collected data is used to assist in understanding the problem represented by the user's voice data; the audio data includes non-voice data frames and voice data frames; Based on the user voice data and the collected data, an answer result corresponding to the user voice data is determined.
2. The voice interaction method according to claim 1, characterized in that: The method of immediately acquiring collected data consistent with a timestamp of the voice front-end point upon detecting the voice front-end point of the real-time input audio data includes: Upon detecting a voice front-end point of audio data input in real time, immediately acquiring first collected data consistent with a timestamp of the voice front-end point from the collected data, and / or immediately controlling a collection device to acquire second collected data consistent with a timestamp of the voice front-end point, and acquiring the second collected data; Wherein, when the first collected data is obtained, the collected data includes the first collected data; when the second collected data is obtained, the collected data includes the second collected data.
3. The voice interaction method according to claim 1, wherein: The determining, based on the user voice data and the collected data, an answer result corresponding to the user voice data includes: If the intention recognition result of the user voice data is a preset intention recognition result, comprehensively determining an answer result corresponding to the user voice data based on the user voice data and the collected data; After immediately acquiring the collected data consistent with the timestamp of the voice front-end point, the method further includes: In a case where the intention recognition result of the user voice data is not a preset intention recognition result, an answer result corresponding to the user voice data is determined based on the user voice data.
4. The voice interaction method according to any one of claims 1 to 3, characterized in that: After immediately acquiring the collected data consistent with the timestamp of the voice front-end point, the method further includes: When it is determined based on the intention recognition result of the user voice data that the collected data is not needed, performing data post-processing on the collected data; The data post-processing includes at least one of the following: In the case that there is data that needs to be pre-processed in the collected data, canceling the data pre-processing process of the collected data; Delete the deletable data in the collected data.
5. The voice interaction method according to any one of claims 1 to 3, characterized in that: The voice interaction method is applied to a processor in a car; Before immediately acquiring collected data consistent with a timestamp of the voice front-end point upon detecting the voice front-end point of the real-time input audio data, the method further includes: When the vehicle is in a voice wake-up state and the vehicle is in a driving state, collecting audio data in real time; Perform voice front-end detection on the audio data.
6. The voice interaction method according to any one of claims 1 to 3, characterized in that: The user voice data is determined based on the following method: When a voice front-end point of the audio data is detected, the voice front-end point is used as the starting acquisition moment to acquire input voice data frames until the endpoint stops acquiring after the voice is detected; Determining the user voice data based on the collected voice data frames; The voice end point represents the termination time of the user voice data input.
7. A voice interaction device, characterized in that: include: A data acquisition module is used to immediately acquire collected data consistent with the timestamp of the voice front-end point when a voice front-end point of the real-time input audio data is detected, and to immediately perform data preprocessing on the collected data if data preprocessing is required; The voice front-end point represents the starting input moment of the user's voice data, and the collected data is used to assist in understanding the problem represented by the user's voice data; the audio data includes non-voice data frames and voice data frames; The answer determination module is used to determine the answer result corresponding to the user voice data based on the user voice data and the collected data.
8. A movable device, characterized in that: include: An audio acquisition device, wherein the audio acquisition device is used to acquire audio data; A processor, wherein the processor is used to execute the voice interaction method according to any one of claims 1 to 6.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the voice interaction method according to any one of claims 1 to 6 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the voice interaction method according to any one of claims 1 to 6 is implemented.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the voice interaction method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Multi-modal interaction method and device, electronic equipment and storage medium
CN118782044A