Voice interaction method and device, electronic equipment, storage medium and program product

By instantly obtaining the collected data of the corresponding timestamp when the voice front end point is detected in the voice interaction system, the problem of low voice interaction accuracy caused by the lag of modal data in the prior art is solved, and a high-accuracy multimodal interaction is achieved.

CN120164465AActive Publication Date: 2025-06-17IFLYTEK CO LTD

Patent Information

Application Number
CN202510637530.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

Other modal data used in the prior art to determine the answer results have lag, resulting in a decrease in the accuracy of voice interactions and affecting the user interaction experience.

Method used

When the voice front endpoint of the audio data is detected, the collected data consistent with the voice front endpoint timestamp is obtained to avoid lag in the acquisition data and ensure that the answer results corresponding to the user's voice data are accurately determined based on the user's voice data and real-time acquisition data.

Benefits of technology

It improves the accuracy and response speed of voice interaction, improves the user interaction experience, and ensures the real-time and accuracy of multimodal interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164465A_ABST
    Figure CN120164465A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of human-computer interaction. The method comprises the following steps: under the condition that a voice front end point of audio data is detected, acquiring acquisition data consistent with the time stamp of the voice front end point; the voice front end point represents an initial input moment of the user voice data, and the collected data is used for assisting in understanding problems represented by the user voice data; and determining an answer result corresponding to the user voice data based on the user voice data and the collected data. According to the invention, when the voice front end point of the audio data is detected, the acquisition data consistent with the time stamp of the voice front end point can be obtained immediately, so that lagging of the acquisition data is avoided, and the accuracy of voice interaction is improved; and when the voice front end point of the audio data is detected, the acquisition data consistent with the timestamp of the voice front end point is acquired, so that the acquisition data is prepared in time, and the response speed of voice interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction technology, and in particular to a voice interaction method, device, electronic equipment, storage medium and program product. Background Art

[0002] Voice interaction is a convenient human-computer interaction method in smart cars. Currently, users' questions are becoming more and more diverse, no longer limited to simple vehicle control and navigation functions. For example, users ask what brand or model the car outside the car is, or what the building outside the car is. Therefore, in addition to analyzing and understanding the user's voice data, other modal data such as images outside the car are also needed to understand the answer. Therefore, multimodal interaction is needed.

[0003] At present, the user's voice data is first obtained, and then other modal data to be collected is determined based on the user's voice data, and then the answer result is determined based on the user's voice data and other modal data, so as to achieve multimodal interaction. However, in the prior art, other modal data used to determine the answer result has a lag, that is, the other modal data is not accurate, resulting in a decrease in the accuracy of voice interaction, which in turn affects the user's interactive experience; for example, the transcribed text of the user's voice data is "What scenic spot did you just pass by?" If the other modal data to be obtained is an image outside the car, the image outside the car collected at this time is collected after the user says the complete "What scenic spot did you just pass by?" and after voice recognition and voice understanding of the user's voice data, so the image outside the car collected may no longer cover the scenic spot, resulting in an inability to accurately obtain the answer result. Summary of the invention

[0004] The present invention provides a voice interaction method, device, electronic device, storage medium and program product, which are used to solve the defect of low accuracy of voice interaction in the prior art and realize a high-accuracy multimodal interaction solution.

[0005] The present invention provides a voice interaction method, comprising: When a voice front-end point of the audio data is detected, acquisition data consistent with the timestamp of the voice front-end point is acquired; the voice front-end point indicates the start input time of the user voice data, and the acquisition data is used to assist in understanding the problem represented by the user voice data; Based on the user voice data and the collected data, an answer result corresponding to the user voice data is determined.

[0006] According to a voice interaction method provided by the present invention, when a voice front-end point of audio data is detected, acquiring collected data consistent with a timestamp of the voice front-end point includes: While detecting the speech front-end point of the real-time input audio data, obtain the first collected data that is consistent with the time stamp of the speech front-end point from the already collected data, and / or, control the collection device to collect the second collected data that is consistent with the time stamp of the speech front-end point, and obtain the second collected data; Wherein, in the case of obtaining the first collected data, the collected data includes the first collected data, and in the case of obtaining the second collected data, the collected data includes the second collected data.

[0007] According to a voice interaction method provided by the present invention, determining the response result corresponding to the user voice data based on the user voice data and the collected data includes: In the case where the intent recognition result of the user voice data is a preset intent recognition result, comprehensively determine the response result corresponding to the user voice data based on the user voice data and the collected data; After obtaining the collected data that is consistent with the time stamp of the speech front-end point, it further includes: In the case where the intent recognition result of the user voice data is not a preset intent recognition result, determine the response result corresponding to the user voice data based on the user voice data.

[0008] According to a voice interaction method provided by the present invention, after obtaining the collected data that is consistent with the time stamp of the speech front-end point, it further includes: In the case where it is determined based on the intent recognition result of the user voice data that the collected data is not required, perform post-processing on the collected data; Wherein, the post-processing of the data includes at least one of the following: In the case where there is data that needs to be preprocessed in the collected data, cancel the data preprocessing process of the collected data; Delete the data that can be deleted in the collected data.

[0009] According to a voice interaction method provided by the present invention, the voice interaction method is applied to a processor in an automobile; Before obtaining the collected data that is consistent with the time stamp of the speech front-end point in the case of detecting the speech front-end point of the audio data, it further includes: When in the voice wake-up state and the vehicle is in the driving state, collect audio data in real time; Perform speech front-end point detection on the audio data.

[0010] According to a voice interaction method provided by the present invention, the user voice data is determined based on the following method: In the case of detecting the speech front-end point of the audio data, use the speech front-end point as the starting acquisition moment, and acquire the input speech data frames until the speech back-end point is detected and the acquisition stops; Based on each of the acquired speech data frames, determine the user speech data; Wherein, the speech back-end point represents the termination input moment of the user speech data.

[0011] The present invention also provides a voice interaction device, including: A data acquisition module, configured to acquire acquisition data with a time stamp consistent with the speech front-end point in the case of detecting the speech front-end point of the audio data; the speech front-end point represents the starting input moment of the user speech data, and the acquisition data is used to assist in understanding the problem represented by the user speech data; An answer determination module, configured to determine an answer result corresponding to the user speech data based on the user speech data and the acquisition data.

[0012] The present invention also provides a movable device, including: An audio acquisition device, which is configured to acquire audio data; A processor, which is configured to execute the voice interaction method as described in any one of the above.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the voice interaction method as described in any one of the above.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the voice interaction method as described in any one of the above.

[0015] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the voice interaction method as described in any one of the above.

[0016] The voice interaction method, device, electronic device, storage medium, and program product provided by the present invention, when detecting the voice front end point of the audio data, obtain the acquisition data consistent with the time stamp of the voice front end point, and the voice front end point represents the starting input moment of the user voice data, so as to ensure that when detecting the voice front end point of the audio data, the acquisition data consistent with the time stamp of the voice front end point is immediately obtained, thus avoiding the lag of the acquisition data, ensuring that based on the user voice data and the real-time acquisition data, the response result corresponding to the user voice data is accurately determined, thereby improving the accuracy of voice interaction and further enhancing the user interaction experience; and when detecting the voice front end point of the audio data, the acquisition data consistent with the time stamp of the voice front end point is obtained, so as to timely prepare the acquisition data for subsequently determining the response result in a timely manner based on the acquisition data, thereby enhancing the response speed of voice interaction; at the same time, the acquisition data is used to assist in understanding the problem represented by the user voice data, so based on the user voice data and the acquisition data, the response result corresponding to the user voice data can be determined more accurately. In summary, the present invention can output the response result corresponding to the user voice data in a timely and accurate manner, thereby realizing a multi-modal interaction solution with high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is one of the flow diagrams of the voice interaction method provided by the present invention.

[0019] Figure 2 It is the second flow diagram of the voice interaction method provided by the present invention.

[0020] Figure 3 It is the third flow diagram of the voice interaction method provided by the present invention.

[0021] Figure 4 It is the fourth flow diagram of the voice interaction method provided by the present invention.

[0022] Figure 5 It is the fifth flow diagram of the voice interaction method provided by the present invention.

[0023] Figure 6 It is the sixth flow diagram of the voice interaction method provided by the present invention.

[0024] Figure 7 It is the structural diagram of the voice interaction device provided by the present invention.

[0025] Figure 8 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments

[0026] In order to make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0027] Most current voice interaction solutions determine other modality data to be collected based on complete user voice data and then collect it. However, there is a lag in the collected other modality data, resulting in a decrease in the accuracy of voice interaction and a poor user interaction experience.

[0028] In view of the poor accuracy of the current voice interaction solutions and the poor user interaction experience, the present invention has conducted research. Initially, the idea was to immediately collect other modality data after obtaining complete user voice data. However, the present invention's research on this idea found that although this idea can reduce the determination time of other modality data to be collected based on complete user voice data, it still needs to wait until complete user voice data is obtained or a voice question-and-answer instruction is received before collecting other modality data, resulting in a lag in the other modality data and still poor accuracy of voice interaction, unable to meet the real-time problem of "What intersection was just passed by".

[0029] To address the problems existing in the above idea, the present invention has continuously conducted research and finally proposed a voice interaction method. In this method, when the voice front end point of audio data is detected, the collected data (other modality data) with the same time stamp as the voice front end point is obtained, so that there is no lag in the other modality data, ensuring that based on the user voice data and the collected data, the answer result corresponding to the user voice data can be accurately determined, thereby improving the accuracy of voice interaction and further enhancing the user interaction experience.

[0030] Next, the voice interaction method provided by the present invention will be introduced through the following embodiments. The following combines Figures 1-6 Describe the voice interaction method of the present invention.

[0031] The voice interaction method provided by the present invention can be applied to an electronic device, which can be set according to actual needs. For example, the electronic device can be a movable vehicle or a movable robot, etc. In a specific embodiment, the voice interaction method can be applied to the processor of a movable device. For example, the movable device is a car to implement the voice interaction method of an intelligent car.

[0032] Figure 1 is one of the schematic flowcharts of the voice interaction method provided by the present invention. As Figure 1 shown, the voice interaction method may include: step 110 and step 120.

[0033] Step 110, when detecting the voice front end point of the audio data, obtain the acquisition data consistent with the time stamp of the voice front end point.

[0034] Here, the audio data is the data of the voice front end point to be detected. The audio data includes multiple frames of audio data frames, and the number of frames is not specifically limited here. When the voice front end point of the audio data can be detected, the audio data includes non-voice data frames and voice data frames; when the voice front end point of the audio data is not detected, the audio data only includes non-voice data frames; the non-voice data frames can be silent data frames (no sound) or noise data frames (background sound), and the last frame of the audio data is usually a voice data frame.

[0035] The audio data is acquired. In an embodiment, the audio data can be acquired by an audio acquisition device (such as a microphone) and forwarded to the execution entity of the voice interaction method provided by the present invention. It should be understood that when the voice front end point has been detected but the voice back end point has not been detected, audio data may not be acquired for voice front end point detection.

[0036] In a specific embodiment, the audio data is input in real time. In other words, the audio data is acquired in real time, that is, while acquiring the audio data, voice front end point detection is performed on the audio data, so as to more real-time ensure that when the voice front end point of the audio data is detected, the acquisition data consistent with the time stamp of the voice front end point is immediately obtained, thereby avoiding lag, ensuring accurate determination of the response result corresponding to the user voice data based on the user voice data and the acquisition data, thereby improving the accuracy of voice interaction and further enhancing the user interaction experience. For example, the duration of the audio data is 3 seconds. In the previous nearly 3-second duration, the user did not ask questions until near the third second when the user started to ask questions. For example, when the user said "just now", the acquisition data consistent with the time stamp of the voice front end point was immediately obtained.

[0037] In one embodiment, when in the voice wake-up state, audio data is collected in real time, and voice front-end point detection is performed on the audio data, so that voice front-end point detection is only executed when in the voice wake-up state, avoiding continuously detecting the voice front-end point, thereby saving computing resources. That is to say, when the user wakes up the voice interaction system, voice front-end point detection is performed on the input audio data.

[0038] In one embodiment, if the voice interaction method is applied to a processor in a vehicle, and when the vehicle is in a driving state, audio data is collected in real time, and voice front-end point detection is performed on the audio data, so that voice front-end point detection is only executed when in the driving state, avoiding continuously detecting the voice front-end point, thereby saving computing resources; and considering that the real-time requirement is not high when the vehicle is in a stationary state, the voice interaction method can be executed when the vehicle is in a driving state.

[0039] It should be understood that the last frame included in the audio data is the currently input or currently collected audio data frame, and as for how long the audio data frames before need to be retained, it is not limited here. It should be noted that retaining the previous audio data frames is for facilitating the detection of the voice front-end point.

[0040] Among them, the voice front-end point represents the starting input moment of the user's voice data. In other words, the voice front-end point is the moment when the audio data transitions from a non-voice data frame to a voice data frame. Therefore, the voice front-end point is usually the input moment or the collection moment of the last frame in the audio data, so as to ensure that the voice front-end point is the current moment, thereby improving the real-time acquisition of the collected data and improving the accuracy of voice interaction; of course, the voice front-end point may not be the input moment or the collection moment of the last frame in the audio data. For example, if the transcribed text of the user's voice data is "What scenic spot did you just pass by", the voice front-end point is the input moment of the first "just". That is to say, when the user starts speaking, the collected data consistent with the voice front-end point timestamp is obtained.

[0041] The user voice data is the voice data spoken by the user, and the user voice data may include multiple voice data frames. For example, the transcribed text of the user voice data may be "What's the name of the intersection just now", "What brand is the car in front", "What model is the car in front", "What's the building just passed by" or "What scenic spot just passed by", etc. It should be noted that most of the user questions in the embodiments of the present invention are real-time questions, that is, questions asked in real time. For such real-time questions, the present invention can accurately perform voice interaction and respond in a timely manner.

[0042] The starting input moment of the user voice data is the input moment (acquisition moment) corresponding to the first character in the transcribed text of the user voice data. For example, if the transcribed text of the user voice data is "What's the name of the intersection just now", then the starting input moment of the user voice data is the input moment of the first "gang". Of course, it is also possible to determine the speaking moment corresponding to the first character in the transcribed text of the user voice data based on the starting input moment, so as to determine the voice front-end point based on the speaking moment, so that the voice front-end point represents the starting speaking moment of the user voice data, thereby improving the real-time acquisition of the collected data and thus improving the accuracy of voice interaction.

[0043] The voice front-end point can be detected by a voice endpoint detection method. The voice endpoint detection method can include, but is not limited to, the following two embodiments.

[0044] For the voice endpoint detection method, in one embodiment, the spectral features of the audio data are extracted frame by frame. Then, based on the extracted spectral features and a pre-constructed endpoint detection model, the posterior probabilities of each frame being voice and non-voice are determined. Finally, based on the posterior probability of each frame of voice, the voice front-end point detection result is output. The endpoint detection model is trained based on sample audio data and its corresponding voice front-end point detection result labels.

[0045] For the voice endpoint detection method, in another embodiment, the energy value of each frame in the audio data is determined. Then, based on the energy value of each frame, the voice front-end point detection result is output. For example, the time point corresponding to the audio data frame with an energy value greater than a preset threshold is the voice front-end point.

[0046] Here, the collected data is the data collected by the collection device. For example, the collected data is an image collected by an image collection device, or an audio collected by an audio collection device, or a speed collected by a speed sensor, or a heart rate collected by a heart rate sensor, etc. The number of the collected data can be one or more. For example, multiple collected data can include an external image of the vehicle and an external sound of the vehicle.

[0047] The acquisition time of the acquired data that is consistent with the voice front-end point timestamp is the voice front-end point. Further, the acquired data may include the data acquired at the voice front-end point and the data acquired within a preset duration before the voice front-end point. For example, if the voice front-end point is the 3rd second, the data from the 2nd second to the 3rd second can be acquired, so as to provide more acquired data for assisting in understanding the problem represented by the user's voice data, thereby improving the accuracy of voice interaction; and it can avoid missing useful acquired data. For example, the user's voice data is "What was the intersection just now". Since the mobile device (such as a car) is moving, the intersection just now is quickly missed. Therefore, it is necessary to acquire the data from the 2nd second to the 3rd second to improve the accuracy of voice interaction. Of course, the acquired data may also include the data acquired at the voice front-end point, the data acquired within a first preset duration before the voice front-end point, and the data acquired within a second preset duration after the voice front-end point. The embodiments of the present invention do not make specific limitations thereto. Among them, the first preset duration and the second preset duration can be set according to the actual situation. For example, both of them are 1 second.

[0048] The acquired data may include first acquired data and / or second acquired data.

[0049] The first acquired data is obtained from the acquired data that has been acquired; specifically, in the case of detecting the voice front-end point of the audio data, the first acquired data that is consistent with the voice front-end point timestamp is obtained from the acquired data that has been acquired. Among them, the acquired data that has been acquired is the data that can be acquired by itself without controlling the acquisition. For example, in the case of needing to acquire an image outside the vehicle, if the vehicle has turned on a driving recorder, an in-vehicle camera, or an in-vehicle 360-degree panoramic imaging system, since it has already been acquired, there is no need to control the acquisition of the image outside the vehicle.

[0050] The second acquired data is acquired by controlling an acquisition device; specifically, in the case of detecting the voice front-end point of the audio data, the acquisition device is controlled to acquire the second acquired data that is consistent with the voice front-end point timestamp, and the second acquired data is obtained. For example, in the case of needing to acquire an image outside the vehicle, the in-vehicle camera is controlled to take a real-time photo of the environment outside the vehicle to acquire the image outside the vehicle.

[0051] Further, in the case of detecting the voice front end point of the audio data, obtain the acquisition data set that is consistent with the voice front end point timestamp. Based on the intent recognition result of the user voice data, determine several pieces of acquisition data that match the intent recognition result from the acquisition data set, so as to determine the response result corresponding to the user voice data based on the user voice data and the several pieces of acquisition data. For example, the user voice data is "What scenic spots did I just pass by", and the acquisition data set includes the external vehicle image and the external vehicle sound. Therefore, only the external vehicle image needs to be determined from the two to determine the response result. Based on this, only the acquisition data matching the user intent is needed to determine the response result, without determining the response result based on other irrelevant data, thereby improving the voice interaction efficiency and the accuracy of the voice interaction. Of course, the embodiments of the present invention can be applied to only one scenario, so that it is known what data needs to be collected. Instead of determining several pieces of acquisition data from the acquisition data set, the response result corresponding to the user voice data is determined based on the user voice data and the acquisition data set. For example, the embodiments of the present invention are only applied to the situation where the user asks questions about the external vehicle environment and the questions are all related to the external vehicle image. Then it is known that the external vehicle image is collected, so there is no need to perform a screening operation.

[0052] Among them, the acquisition data is used to assist in understanding the problem represented by the user voice data. Based on this, it is convenient to more accurately determine the response result corresponding to the user voice data, thereby improving the accuracy of the voice interaction and further enhancing the user interaction experience.

[0053] Exemplarily, if the transcribed text of the user voice data is "What is the scenic spot I just passed by", the acquisition data is the external vehicle image; if the transcribed text of the user voice data is "What sound was outside the vehicle just now", the acquisition data is the external vehicle sound; if the transcribed text of the user voice data is "What was my expression just now", the acquisition data is the internal vehicle image; if the transcribed text of the user voice data is "What is the current vehicle speed", the acquisition data is the vehicle speed; if the transcribed text of the user voice data is "What was my heart rate just now", the acquisition data is the user heart rate.

[0054] In a specific embodiment, while detecting the voice front-end point of the audio data, the acquisition data consistent with the time stamp of the voice front-end point is obtained, so as to more real-time ensure that while detecting the voice front-end point of the audio data, the acquisition data consistent with the time stamp of the voice front-end point is immediately obtained, thereby avoiding lag, ensuring that based on the user voice data and the acquisition data, the response result corresponding to the user voice data is accurately determined, thereby improving the accuracy of voice interaction, and further enhancing the user interaction experience; and while detecting the voice front-end point of the audio data, the acquisition data consistent with the time stamp of the voice front-end point is obtained, so as to timely prepare the acquisition data for subsequent timely determining the response result based on the acquisition data, that is, improving the response speed of voice interaction. For example, the duration of the audio data is 3 seconds, and the user did not ask questions in the previous nearly 3 seconds. It was not until near the third second that the user began to ask questions. For example, when the user said "just now", the acquisition data consistent with the time stamp of the voice front-end point was immediately obtained.

[0055] Furthermore, time stamp annotation is performed on the acquisition data to ensure the synchronization between the acquisition data and the user voice data, ensuring that the response result corresponding to the user voice data is obtained more accurately and improving the accuracy of voice interaction.

[0056] Step 120: Based on the user voice data and the acquisition data, determine the response result corresponding to the user voice data.

[0057] Here, the response result is the answer to the question represented by the user voice data. For example, if the transcribed text of the user voice data is "What is the scenic spot I just passed by", the response result can be "A certain scenic spot".

[0058] In a specific embodiment, the user voice data and the acquisition data are input into the voice interaction model together to obtain the response result output by the voice interaction model. Furthermore, the voice interaction model can be constructed based on the large language model. Furthermore, the transcribed text of the user voice data and the acquisition data are input into the voice interaction model together to obtain the response result output by the voice interaction model.

[0059] The reason for using the large language model (Large Language model, LLM) to construct the voice interaction model is that the large language model has richer prior knowledge and reasoning ability compared with traditional small models or pre-trained models. At the same time, compared with implicit information extraction and modeling, the large language model can explicitly output steps such as thinking, analysis, and reasoning, so as to enhance the correctness of the final reasoning result at the semantic level. In this way, the robustness of the voice interaction model can be improved, thereby improving the accuracy of voice interaction.

[0060] Further, convert the user voice data into a transcribed text, and based on the transcribed text of the user voice data and the collected data, determine the response result corresponding to the user voice data. The transcribed text can better assist in understanding the problem represented by the user voice data, thereby further improving the accuracy of voice interaction.

[0061] It should be understood that by comprehensively understanding and analyzing based on the user voice data and the collected data, the response result corresponding to the user voice data can be determined more accurately, that is, by performing multimodal understanding and analysis, the accuracy of voice interaction can be improved.

[0062] It can be understood that considering that other modal data is collected after obtaining the complete user voice data, it is mostly applied in the vehicle stationary state. In the driving state, due to the movement of the vehicle, there will be problems of inconsistency and mismatch between the user voice data and the actual scenario. Therefore, in the embodiment of the present invention, when the voice front end point of the audio data is detected, the collected data with the same time stamp as the voice front end point is immediately obtained, avoiding the lag of the collected data, so that the user voice data and the collected data are consistent and matched, thereby improving the accuracy of voice interaction in the driving state. That is, the embodiment of the present invention can provide a voice interaction method that can accurately respond to user needs in the driving state and improve the real-time performance and accuracy of multimodal interaction; and the embodiment of the present invention can avoid restricting the usage scenario of voice interaction, that is, it can be used regardless of the driving state or the stationary state.

[0063] The voice interaction method provided by the embodiment of the present invention obtains the collected data with the same time stamp as the voice front end point when the voice front end point of the audio data is detected, and the voice front end point represents the starting input moment of the user voice data, so as to ensure that when the voice front end point of the audio data is detected, the collected data with the same time stamp as the voice front end point is immediately obtained, thereby avoiding the lag of the collected data, ensuring that based on the user voice data and the real-time collected data, the response result corresponding to the user voice data is accurately determined, thereby improving the accuracy of voice interaction and further enhancing the user interaction experience; and when the voice front end point of the audio data is detected, the collected data with the same time stamp as the voice front end point is obtained, so as to prepare the collected data in time for subsequent determination of the response result based on the collected data in time, thereby enhancing the response speed of voice interaction; at the same time, the collected data is used to assist in understanding the problem represented by the user voice data. Therefore, based on the user voice data and the collected data, the response result corresponding to the user voice data can be determined more accurately. In summary, the present invention can output the response result corresponding to the user voice data in a timely and accurate manner, thereby implementing a multimodal interaction solution with high accuracy.

[0064] Based on any of the above embodiments, a specific embodiment of the voice interaction method is given next. Figure 2This is the second schematic flowchart of the voice interaction method provided by the present invention. As Figure 2 shown, the voice interaction method includes: step 111 and step 120.

[0065] Step 111, while detecting the voice front end point of the real-time input audio data, obtain the first acquisition data with the same time stamp as the voice front end point from the already acquired data, and / or, control the acquisition device to acquire the second acquisition data with the same time stamp as the voice front end point, and obtain the second acquisition data.

[0066] Here, the audio data is input in real time, so that the voice front end point detection is performed while the audio data is being input in real time, and the voice front end point detection is performed as long as there is audio data input, rather than first obtaining a complete audio data and then performing post-processing for voice front end point detection, thereby avoiding the lag in the acquisition data and improving the accuracy of voice interaction.

[0067] That is to say, the last frame of the audio data is the currently real-time input, that is, the last frame included in the audio data is the currently input or currently acquired audio data frame. As for how long the audio data frames before need to be retained, there is no limitation here. It should be noted that retaining the previous audio data frames is for facilitating the detection of the voice front end point. For example, the transcribed text of the user's voice data is "Just passed by what scenic spots", the voice front end point is the input moment of the first "Just", and the input moment of the last frame of the corresponding audio data is the voice front end point, so that the acquisition data with the same time stamp as the voice front end point can be obtained when the user starts speaking.

[0068] In other words, the audio data is acquired in real time, that is, the voice front end point detection is performed on the audio data while the audio data is being acquired, so as to more real-time ensure that while detecting the voice front end point of the audio data, the acquisition data with the same time stamp as the voice front end point is immediately obtained, thereby avoiding lag.

[0069] It should be noted that while detecting the voice front-end point of the real-time input audio data, the acquisition data consistent with the time stamp of the voice front-end point is immediately obtained, so as to avoid the lag of the acquisition data, ensure that based on the user voice data and the acquisition data, the response result corresponding to the user voice data is accurately determined, thereby improving the accuracy of voice interaction, and further enhancing the user interaction experience. Moreover, while detecting the voice front-end point of the audio data, the acquisition data consistent with the time stamp of the voice front-end point is obtained, so as to promptly prepare the acquisition data for subsequently determining the response result in a timely manner based on the acquisition data, that is, to improve the response speed of voice interaction. For example, the duration of the audio data is 3 seconds. In the previous nearly 3-second duration, the user did not ask any questions. It was not until near the third second that the user started to ask questions. For example, when the user said "just now", the acquisition data consistent with the time stamp of the voice front-end point was immediately obtained.

[0070] Here, the acquired data is the data acquired by the acquisition device. For example, the acquired data is an image acquired by an image acquisition device, or an audio acquired by an audio acquisition device, or a speed acquired by a speed sensor, or a heart rate acquired by a heart rate sensor, and so on. The number of the acquired data can be one or more. For example, multiple acquired data can include an external vehicle image and an external vehicle sound. The acquired data is the data that can be acquired automatically without control of acquisition. For example, in the case of needing to acquire an external vehicle image, if the vehicle has turned on a driving recorder, an in-vehicle camera or an in-vehicle 360-degree panoramic imaging system, since it has already been acquired, there is no need to control the acquisition of the external vehicle image.

[0071] The acquisition moment of the first acquisition data consistent with the time stamp of the voice front-end point is the voice front-end point. Further, the first acquisition data can include the data acquired at the voice front-end point and the data acquired within a preset duration before the voice front-end point. For example, if the voice front-end point is the 3rd second, the data from the 2nd second to the 3rd second can be acquired, so as to provide more acquisition data for assisting in understanding the problem represented by the user voice data, and further improving the accuracy of voice interaction. And it can avoid missing useful acquisition data. For example, the user voice data is "What was the intersection just now". Since the movable device (such as a car) is moving, the intersection just now will soon be missed. Therefore, the data from the 2nd second to the 3rd second needs to be acquired to improve the accuracy of voice interaction. Of course, the first acquisition data can also include the data acquired at the voice front-end point, the data acquired within the first preset duration before the voice front-end point and the data acquired within the second preset duration after the voice front-end point. The embodiments of the present invention do not make specific limitations in this regard. Correspondingly, the acquired data needs to include data of the same type as the first acquisition data.

[0072] Here, the acquisition device is used to acquire second acquisition data, and the acquisition device can be set according to the type of the second acquisition data, which is not specifically limited here. For example, if the second acquisition data is an image, the acquisition device is an image acquisition device; if the second acquisition data is audio, the acquisition device is an audio acquisition device; if the second acquisition data is speed, the acquisition device is a speed sensor; if the second acquisition data is heart rate, the acquisition device is a heart rate sensor.

[0073] Here, the second acquisition data is acquired by controlling the acquisition device. For example, in the case of needing to acquire an image outside the vehicle, the in-vehicle camera is controlled to take a real-time photo of the external environment of the vehicle to acquire the image outside the vehicle.

[0074] The acquisition moment of the second acquisition data that is consistent with the voice front-end point timestamp is the voice front-end point. Further, the second acquisition data may include the data acquired at the voice front-end point and the data acquired within a preset duration before the voice front-end point. For example, if the voice front-end point is the 3rd second, the data from the 2nd second to the 3rd second can be acquired, so as to provide more acquisition data for assisting in understanding the problem represented by the user's voice data, thereby improving the accuracy of voice interaction; and it can avoid missing useful acquisition data. For example, the user's voice data is "What was the intersection just now". Since the car is moving, the intersection just now is quickly missed, so the data from the 2nd second to the 3rd second needs to be acquired to improve the accuracy of voice interaction. Of course, the second acquisition data may also include the data acquired at the voice front-end point, the data acquired within a first preset duration before the voice front-end point, and the data acquired within a second preset duration after the voice front-end point. The embodiments of the present invention do not make specific limitations on this.

[0075] It should be noted that since the second acquisition data is acquired by the acquisition device, the execution subject of the voice interaction method provided by the present invention needs to obtain the second acquisition data acquired by the acquisition device, that is, the acquisition device forwards the second acquisition data to the execution subject.

[0076] Among them, in the case of obtaining the first acquisition data, the acquisition data includes the first acquisition data, and in the case of obtaining the second acquisition data, the acquisition data includes the second acquisition data. It should be understood that the acquisition data may include the first acquisition data and / or the second acquisition data.

[0077] Exemplarily, if the transcribed text of the user voice data is "What is the scenic spot I just passed by", the collected data is the external vehicle image, and it can be the first collected data or the second collected data; if the transcribed text of the user voice data is "What was the sound outside the vehicle just now", the collected data is the external vehicle sound, and it can be the first collected data or the second collected data; if the transcribed text of the user voice data is "What was my expression just now", the collected data is the internal vehicle image, and it can be the first collected data or the second collected data; if the transcribed text of the user voice data is "What is the current vehicle speed", the collected data is the vehicle speed, and it can be the first collected data or the second collected data; if the transcribed text of the user voice data is "What was my heart rate just now", the collected data is the user heart rate, and it can be the first collected data or the second collected data.

[0078] Step 120: Based on the user voice data and the collected data, determine the answer result corresponding to the user voice data.

[0079] Specifically, based on the user voice data, and the first collected data and / or the second collected data, determine the answer result.

[0080] In the voice interaction method provided by the embodiments of the present invention, while detecting the voice front end point of the real-time input audio data, obtain the first collected data with the same time stamp as the voice front end point from the already collected data, and / or, control the collection device to collect the second collected data with the same time stamp as the voice front end point, and the voice front end point represents the starting input moment of the user voice data, so as to ensure that while detecting the voice front end point of the real-time input audio data, immediately obtain the collected data with the same time stamp as the voice front end point, thereby avoiding the lag of the collected data, ensuring that based on the user voice data and the real-time collected data, accurately determine the answer result corresponding to the user voice data, thereby improving the accuracy of voice interaction, and further enhancing the user interaction experience; and while detecting the voice front end point of the real-time input audio data, obtain the collected data with the same time stamp as the voice front end point, so as to timely prepare the collected data for subsequent timely determining the answer result based on the collected data, thereby enhancing the response speed of voice interaction.

[0081] Based on any of the above embodiments, a specific embodiment of the voice interaction method is given next. Figure 3 It is the third flow chart of the voice interaction method provided by the present invention, as Figure 3 shown, this voice interaction method includes: Step 110, Step 121 and Step 130. This Step 130 is after Step 110.

[0082] Step 110, when the voice front-end point of the audio data is detected, obtain the acquisition data that is consistent with the time stamp of the voice front-end point.

[0083] Specifically, when the voice front-end point of the audio data is detected, start collecting user voice data. The user voice data is the voice data spoken by the user, and the user voice data may include multiple voice data frames. Based on the user voice data, the user's complete question can be determined to ensure an accurate answer result subsequently.

[0084] Step 121, when the intent recognition result of the user voice data is a preset intent recognition result, comprehensively determine the answer result corresponding to the user voice data based on the user voice data and the acquisition data.

[0085] Here, the intent recognition result is obtained by performing intent recognition on the user voice data. Further, the intent recognition result is obtained by performing intent recognition on the transcribed text of the user voice data.

[0086] In one embodiment, perform speech recognition on the user voice data to obtain the transcribed text of the user voice data, and perform semantic understanding on the transcribed text to obtain the intent recognition result. Among them, both speech recognition and semantic understanding can be implemented through an artificial intelligence model.

[0087] In another embodiment, input the user voice data into an intent recognition model to obtain the intent recognition result output by the intent recognition model. The intent recognition model can be constructed based on a large language model or obtained by training an initial model constructed based on sample user voice data and their corresponding intent recognition result labels.

[0088] Here, the preset intent recognition result is preset. For example, it is the intent to inquire about information on buildings outside the vehicle, the intent to inquire about information on scenic spots outside the vehicle, the intent to inquire about information on cars outside the vehicle, and so on. The preset intent recognition results are all intents corresponding to multi-modal interactions, that is, intents that require acquisition data (other modal data) to jointly determine the answer result. The number of preset intent recognition results can be one or more. For example, if there are multiple intents that require multi-modal interactions, the number of preset intent recognition results is multiple.

[0089] Further, considering that not all acquisition data is related to the user's question, that is, not all acquisition data is related to the intent recognition result, in the embodiments of the present invention, based on the intent recognition result of the user voice data, target data that matches the question represented by the user voice data is screened out from the acquisition data; based on the user voice data and the target data, the answer result corresponding to the user voice data is comprehensively determined. Based on this, it is possible to avoid determining the answer result based on irrelevant acquisition data, thereby improving the efficiency of voice interaction and the accuracy of voice interaction.

[0090] It should be understood that based on the user voice data and the collected data, the response result corresponding to the user voice data is comprehensively determined, that is, through comprehensive understanding and analysis, the response result corresponding to the user voice data can be more accurately determined. That is, through multimodal understanding and analysis, the accuracy of voice interaction can be improved.

[0091] Step 130, in the case where the intent recognition result of the user voice data is not the preset intent recognition result, based on the user voice data, determine the response result corresponding to the user voice data.

[0092] The preset intent recognition results are all the intents corresponding to the required multimodal interaction. If the intent recognition result of the user voice data is not the preset intent recognition result, it indicates that no other modal data (collected data) is required. Therefore, it is only necessary to determine the response result corresponding to the user voice data based on the user voice data, thereby improving the efficiency of voice interaction and the accuracy of voice interaction.

[0093] For example, if the transcribed text of the user voice data is "What time is it now", then its intent recognition result is not the preset intent recognition result, and thus it is only necessary to determine the response result based on the user voice data.

[0094] In a specific embodiment, the user voice data is input into a voice interaction model to obtain the response result output by the voice interaction model. Further, the voice interaction model can be constructed based on a large language model. Further, the transcribed text of the user voice data is input into the voice interaction model to obtain the response result output by the voice interaction model.

[0095] The reason for constructing the voice interaction model using a large language model is that the large language model has richer prior knowledge and reasoning capabilities compared to traditional small models or pre-trained models. At the same time, compared with implicit information extraction and modeling, the large language model can explicitly output steps such as thinking, analysis, and reasoning, thereby being able to enhance the correctness of the final reasoning result at the semantic level. In this way, the robustness of the voice interaction model can be improved, and thus the accuracy of voice interaction can be improved.

[0096] Further, the user voice data is converted into a transcribed text, and based on the transcribed text of the user voice data, the response result corresponding to the user voice data is determined. The transcribed text can better assist in understanding the problem represented by the user voice data, thereby further improving the accuracy of voice interaction.

[0097] In the voice interaction method provided by an embodiment of the present invention, when the intent recognition result of the user voice data is a preset intent recognition result, based on the user voice data and the acquisition data, the response result corresponding to the user voice data is comprehensively determined, which can more accurately determine the response result corresponding to the user voice data, that is, perform multimodal understanding and analysis, and can improve the accuracy of voice interaction; while when the intent recognition result of the user voice data is not a preset intent recognition result, based on the user voice data, the response result corresponding to the user voice data is determined, so that without other modal data (acquisition data), the response result corresponding to the user voice data can be determined based on the user voice data alone, thereby improving the efficiency of voice interaction and the accuracy of voice interaction.

[0098] Based on any of the above embodiments, another embodiment of the voice interaction method is given next. Figure 4 It is the fourth flowchart of the voice interaction method provided by the present invention, as Figure 4 shown, the voice interaction method includes: step 110, step 120, and step 140. This step 140 is after step 110.

[0099] Step 110, when the voice front end point of the audio data is detected, obtain the acquisition data with the same time stamp as the voice front end point.

[0100] Here, the acquisition data is the data collected by the acquisition device. For example, the acquisition data is an image collected by an image acquisition device, and this acquisition data usually needs to be preprocessed; or an audio collected by an audio acquisition device, and this acquisition data usually needs to be preprocessed; or the speed collected by a speed sensor, or the heart rate collected by a heart rate sensor, etc. Further, to improve the real-time performance of voice interaction, after obtaining the acquisition data with the same time stamp as the voice front end point, data preprocessing is immediately performed, so that the response result can be quickly and directly determined based on the preprocessed acquisition data subsequently.

[0101] It should be understood that after the acquisition data is obtained, it usually needs to be cached or stored in the storage space for use in subsequent steps.

[0102] Step 120, based on the user voice data and the acquisition data, determine the response result corresponding to the user voice data.

[0103] In a specific embodiment, if the acquisition data requires data preprocessing, then based on the user voice data and the preprocessed acquisition data, determine the response result corresponding to the user voice data.

[0104] Step 140, when it is determined based on the intent recognition result of the user voice data that the acquisition data is not required, perform data post-processing on the acquisition data.

[0105] Here, the intention recognition result is obtained by performing intention recognition on the user's voice data. Further, the intention recognition result is obtained by performing intention recognition on the transcribed text of the user's voice data.

[0106] Considering that not all the collected data is related to the user's question, that is, not all the collected data is related to the intention recognition result, some intention recognition results do not require collected data, so it is necessary to perform post-processing on the data. The reason for performing post-processing on the data is that when the voice front-end point of the audio data is detected, the complete user voice data has not been obtained yet, so it is impossible to determine whether collected data is needed. That is, whether it is needed or not, the collected data consistent with the time stamp of the voice front-end point is obtained first.

[0107] In a specific embodiment, when the intention recognition result of the user's voice data is not the preset intention recognition result, it is determined that no collected data is needed.

[0108] The preset intention recognition result is preset. For example, it is the intention of asking for information about buildings outside the vehicle, the intention of asking for information about scenic spots outside the vehicle, the intention of asking for information about cars outside the vehicle, and so on. The preset intention recognition results are all intentions corresponding to multi-modal interactions, that is, the intentions that require collected data (other modal data) to jointly determine the answer result. Since the preset intention recognition results are all intentions corresponding to multi-modal interactions, when the intention recognition result of the user's voice data is not the preset intention recognition result, it indicates that no other modal data (collected data) is needed.

[0109] For example, the preset intention recognition result is the intention of asking for information about buildings outside the vehicle, and the collected data includes images outside the vehicle. When the intention recognition result of the user's voice data is not the intention of asking for information about buildings outside the vehicle, post-processing is performed on the collected data (such as deleting the images outside the vehicle and suspending the image preprocessing process).

[0110] Wherein, the data post-processing includes at least one of the following: when there is data in the collected data that needs to be preprocessed, cancel the data preprocessing process of the collected data; delete the data that can be deleted in the collected data.

[0111] Considering that some of the collected data needs to be preprocessed before the subsequent answer result can be determined, while some of the collected data does not need to be preprocessed. Therefore, when there is data in the collected data that needs to be preprocessed, cancel the data preprocessing process of the collected data to avoid wasting computing resources, thereby improving the efficiency of voice interaction.

[0112] For example, if the collected data is image data, it may need to be preprocessed; if the collected data is speed data, it may not need to be preprocessed.

[0113] The specific method of this data preprocessing is related to the data type of the collected data, and no specific limitation is made here. For example, if the collected data is image data, the corresponding data preprocessing methods may include but are not limited to: image cropping, denoising, etc.

[0114] Canceling the data preprocessing process of the collected data can be understood as canceling the data preprocessing process when the collected data is about to start the data preprocessing process, that is, not performing the data preprocessing process; or canceling the data preprocessing process when the collected data is already in the data preprocessing process, that is, interrupting the data preprocessing process and not performing the data preprocessing process subsequently.

[0115] Considering that not all collected data can be deleted, only the deletable data in the collected data is deleted. For example, a driving recorder and a vehicle-mounted 360-degree panoramic imaging system do not need to be deleted due to their own functional requirements. Based on this, deleting the deletable data in the collected data can reduce the occupation of storage space and avoid waste of storage resources. It can be understood that the deletable data is only the data collected and stored for the needs of this voice interaction method.

[0116] In addition, the user can also perform post-processing on the collected data through voice. For example, although the user asks a question, the user may immediately cancel the question.

[0117] The voice interaction method provided by the embodiments of the present invention cancels the data preprocessing process of the collected data when it is determined based on the intention recognition result of the user voice data that the collected data is not required and there is data in the collected data that needs to be preprocessed, thereby avoiding waste of computing resources and improving the efficiency of voice interaction; at the same time, when it is determined based on the intention recognition result of the user voice data that the collected data is not required, deleting the deletable data in the collected data can reduce the occupation of storage space, thereby avoiding waste of storage resources and providing more resources to where the voice interaction method really needs, thereby ensuring the stability of voice interaction.

[0118] Based on any of the above embodiments, another embodiment of the voice interaction method is given next. Figure 5 It is the fifth flow chart of the voice interaction method provided by the present invention, as Figure 5 shown, the voice interaction method includes: step 510, step 520, step 110 and step 120.

[0119] This voice interaction method is applied to a processor in an automobile. Based on this, voice interaction capabilities and multi-modal interaction capabilities can be realized on intelligent vehicles.

[0120] Step 510, when in the voice wake-up state and the vehicle is in the driving state, collect audio data in real time.

[0121] Here, the voice wake-up state means that the user wakes up the voice interaction system. For example, the user says a preset wake-up word to wake up the voice interaction system. The preset wake-up word is a pre-defined wake-up word, such as "Xiaofei, Xiaofei". It should be understood that after waking up the voice interaction, it is in the voice wake-up state, and at this time, user voice data can be received for voice interaction.

[0122] It can be understood that audio data is collected in real time only when in the voice wake-up state, that is, this voice interaction method is executed only then, avoiding continuously collecting audio data and continuously performing voice front-end point detection, thereby avoiding wasting resources and reducing device power consumption. In other words, voice front-end point detection is performed only when in the voice wake-up state, avoiding continuously detecting the voice front-end point, thus saving computing resources; that is, when the user wakes up the voice interaction system, voice front-end point detection is performed on the input audio data.

[0123] Here, the driving state means that the vehicle is in the driving state (moving state). Based on this, audio data is collected in real time only when the vehicle is in the driving state, that is, this voice interaction method is executed only then, so that when the vehicle is in the stationary state, the answer result can be directly determined based on the user voice data, thereby improving the voice interaction efficiency in the stationary state. And voice front-end point detection is performed only when in the driving state, which can avoid continuously detecting the voice front-end point, thus saving computing resources; and it is also considered that when the vehicle is in the stationary state, the requirement for real-time performance is not high, so this voice interaction method can be executed only when the vehicle is in the driving state.

[0124] Here, the audio data is collected in real time, so voice front-end point detection is performed while the audio data is being collected in real time, and voice front-end point detection is performed as long as audio data is collected, rather than first obtaining a complete audio data and then performing voice front-end point detection in post-processing, thereby avoiding lag in data collection and improving the accuracy of voice interaction.

[0125] That is to say, the last frame of the audio data is the currently collected in real time, that is, the last frame included in the audio data is the currently collected audio data frame, and as for how long the audio data frames from before need to be retained, there is no limit here. It should be noted that retaining the previous audio data frames is for facilitating the detection of the voice front-end point. For example, the transcribed text of the user voice data is "What scenic spots did you just pass by", the voice front-end point is the input moment of the first "gang", and the collection moment of the last frame of the corresponding audio data is the voice front-end point, so that the collection data consistent with the voice front-end point timestamp can be obtained when the user starts speaking.

[0126] In other words, the audio data is input in real time, so that voice front-end point detection is performed on the audio data while the audio data is being collected, thereby ensuring more real-time that when the voice front-end point of the audio data is detected, the collected data consistent with the time stamp of the voice front-end point is immediately obtained, thereby avoiding lag in the collected data and further improving the accuracy of voice interaction.

[0127] Step 520: Perform voice front-end point detection on the audio data.

[0128] The voice front-end point can be detected by a voice endpoint detection method. This voice endpoint detection method can refer to the above two embodiments.

[0129] Step 110: When the voice front-end point of the audio data is detected, obtain the collected data consistent with the time stamp of the voice front-end point.

[0130] After performing voice front-end point detection on the audio data as described above, when the voice front-end point of the audio data is detected, the collected data consistent with the time stamp of the voice front-end point is obtained, thereby ensuring in real time that when the voice front-end point of the audio data is detected, the collected data consistent with the time stamp of the voice front-end point is immediately obtained, thereby avoiding lag in the collected data, ensuring that based on the user voice data and the collected data, the answer result corresponding to the user voice data is accurately determined, thereby improving the accuracy of voice interaction and further enhancing the user interaction experience; and when the voice front-end point of the audio data is detected, the collected data consistent with the time stamp of the voice front-end point is obtained, thereby timely preparing the collected data for subsequent timely determining the answer result based on the collected data, that is, enhancing the response speed of voice interaction.

[0131] Step 120: Based on the user voice data and the collected data, determine the answer result corresponding to the user voice data.

[0132] Specifically, based on the user voice data and the collected data, determine the answer result when in the voice wake-up state and the vehicle is in the driving state.

[0133] The voice interaction method provided by the embodiments of the present invention collects audio data and performs voice front-end point detection only when in the voice wake-up state, avoiding continuously detecting the voice front-end point, thereby saving computing resources; collects audio data and performs voice front-end point detection only when in the driving state, avoiding continuously detecting the voice front-end point, thereby saving computing resources.

[0134] Based on any of the above embodiments, another embodiment of the voice interaction method is given next. Figure 6 It is the sixth flow diagram of the voice interaction method provided by the present invention, as Figure 6As shown, the user voice data is determined based on the following steps.

[0135] Step 610: When the voice front end point of the audio data is detected, use the voice front end point as the starting acquisition moment, and acquire the input voice data frames until the voice back end point is detected and the acquisition stops.

[0136] Here, the voice front end point represents the starting input moment of the user voice data. In other words, the voice front end point is the moment when the audio data transitions from non-voice data frames to voice data frames. Therefore, using the voice front end point as the starting acquisition moment, the voice data frames are acquired until the voice back end point is detected.

[0137] This starting acquisition moment is the starting acquisition moment of the user voice data, that is, the voice front end point represents the starting acquisition moment of the user voice data. It can be understood that the starting acquisition moment of the user voice data is the acquisition moment corresponding to the first character in the transcribed text of the user voice data. For example, if the transcribed text of the user voice data is "What's the name of the intersection just now", then the starting acquisition moment of the user voice data is the acquisition moment of the first "gang".

[0138] Among them, the voice back end point represents the termination input moment of the user voice data. In other words, the voice back end point is the moment when the voice data transitions to non-voice data frames. Therefore, the voice back end point is usually the input moment or acquisition moment of the last frame in the user voice data, so as to ensure the accuracy of the user voice data and further improve the accuracy of voice interaction. For example, if the transcribed text of the user voice data is "What scenic spots did you just pass by", then the voice back end point is the input moment of "qu". That is to say, the moment when the user finishes speaking is the voice back end point.

[0139] The voice back end point can be detected by a voice end point detection method. This voice end point detection method can refer to the above two embodiments.

[0140] Step 620: Based on each of the acquired voice data frames, determine the user voice data.

[0141] Here, the user voice data is the voice data spoken by the user, and the user voice data includes each voice data frame. For example, the transcribed text of the user voice data can be "What's the name of the intersection just now", "What brand is the car in front", "What model is the car in front", "What's the name of the building just passed by", or "What's the name of the scenic spot just passed by", etc.

[0142] In the voice interaction method provided by the embodiment of the present invention, when the voice front end point of the audio data is detected, the voice front end point is immediately used as the starting acquisition moment to immediately acquire the input voice data frames until the voice back end point is detected to stop the acquisition. Thus, when the last frame of the user voice data is acquired, the user voice data can be obtained (for example, the user voice data can be obtained immediately after the user finishes speaking the question), so that the answer result can be quickly determined based on the user voice data subsequently. Compared with the prior art where the complete audio data is first obtained and then the end points of the complete audio data are detected to obtain the user voice data, the present invention has no lag, can improve the real-time performance of voice interaction, and can improve the response speed of voice interaction.

[0143] The voice interaction device provided by the present invention will be described below. The voice interaction device described below can be correspondingly referred to the voice interaction method described above.

[0144] Figure 7 is a schematic structural diagram of the voice interaction device provided by the present invention, as Figure 7 shown, the voice interaction device includes a data acquisition module 710 and an answer determination module 720.

[0145] The data acquisition module 710 is configured to acquire the acquisition data with the same time stamp as the voice front end point when the voice front end point of the audio data is detected; the voice front end point represents the starting input moment of the user voice data, and the acquisition data is used to assist in understanding the problem represented by the user voice data.

[0146] The answer determination module 720 is configured to determine the answer result corresponding to the user voice data based on the user voice data and the acquisition data.

[0147] The voice interaction device provided by the embodiment of the present invention, when detecting the voice front end point of the audio data, obtains the acquisition data consistent with the time stamp of the voice front end point, and the voice front end point represents the starting input moment of the user voice data, so as to ensure that when detecting the voice front end point of the audio data, the acquisition data consistent with the time stamp of the voice front end point is immediately obtained, thereby avoiding the lag of the acquisition data, ensuring that based on the user voice data and the real-time acquisition data, the answer result corresponding to the user voice data is accurately determined, thereby improving the accuracy of voice interaction, and further enhancing the user interaction experience; and when detecting the voice front end point of the audio data, the acquisition data consistent with the time stamp of the voice front end point is obtained, so as to prepare the acquisition data in time for subsequent timely determination of the answer result based on the acquisition data, thereby enhancing the response speed of voice interaction; at the same time, the acquisition data is used to assist in understanding the problem represented by the user voice data, so based on the user voice data and the acquisition data, the answer result corresponding to the user voice data can be determined more accurately. In summary, the present invention can output the answer result corresponding to the user voice data in a timely and accurate manner.

[0148] Based on any of the above embodiments, the data acquisition module 710 is specifically configured to: When detecting the voice front end point of the real-time input audio data, obtain the first acquisition data consistent with the time stamp of the voice front end point from the acquired data, and / or, control the acquisition device to acquire the second acquisition data consistent with the time stamp of the voice front end point, and obtain the second acquisition data; Wherein, when the first acquisition data is obtained, the acquisition data includes the first acquisition data, and when the second acquisition data is obtained, the acquisition data includes the second acquisition data.

[0149] Based on any of the above embodiments, the answer determination module 720 is specifically configured to: when the intention recognition result of the user voice data is a preset intention recognition result, comprehensively determine the answer result corresponding to the user voice data based on the user voice data and the acquisition data.

[0150] The answer determination module 720 is further configured to: when the intention recognition result of the user voice data is not a preset intention recognition result, determine the answer result corresponding to the user voice data based on the user voice data.

[0151] Based on any of the above embodiments, the device further includes a data post-processing module, and the data post-processing module is configured to: When it is determined based on the intention recognition result of the user voice data that the acquisition data is not required, perform data post-processing on the acquisition data; Wherein, the data post-processing includes at least one of the following: In the case that there is data that needs to be pre - processed in the collected data, cancel the data pre - processing process of the collected data; Delete the data that can be deleted in the collected data.

[0152] Based on any of the above - mentioned embodiments, the voice interaction method is applied to a processor in an automobile; the device further includes: A data acquisition module, configured to collect audio data in real - time when in a voice wake - up state and the automobile is in a driving state; An endpoint detection module, configured to perform voice front - endpoint detection on the audio data.

[0153] Based on any of the above - mentioned embodiments, the user voice data is determined based on the following data determination module, and the data determination module is used for: When detecting the voice front - endpoint of the audio data, use the voice front - endpoint as the starting acquisition moment, collect the input voice data frames until detecting the voice back - endpoint and stop collecting; Determine the user voice data based on each of the collected voice data frames; Wherein, the voice back - endpoint represents the termination input moment of the user voice data.

[0154] The movable device provided by the present invention is described below, and the movable device described below can be correspondingly referred to the voice interaction method described above.

[0155] The present invention provides a movable device, which includes an audio acquisition device and a processor. The audio acquisition device is used to collect audio data; the processor is used to execute the voice interaction method described in any of the above - mentioned embodiments. For example, the movable device is a movable device such as an automobile or a mobile robot.

[0156] Figure 8 Illustrates a schematic physical structure diagram of an electronic device, such as Figure 8As shown in the figure, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call logical instructions in the memory 830 to execute a voice interaction method, which includes: when detecting a voice front end point of audio data, obtaining acquisition data that is consistent with the time stamp of the voice front end point; the voice front end point represents the starting input moment of user voice data, and the acquisition data is used to assist in understanding the problem represented by the user voice data; based on the user voice data and the acquisition data, determining a response result corresponding to the user voice data.

[0157] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0158] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the voice interaction method provided by the above-mentioned various methods. The method includes: when detecting a voice front end point of audio data, obtaining acquisition data that is consistent with the time stamp of the voice front end point; the voice front end point represents the starting input moment of user voice data, and the acquisition data is used to assist in understanding the problem represented by the user voice data; based on the user voice data and the acquisition data, determining a response result corresponding to the user voice data.

[0159] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the voice interaction method provided by the above-mentioned various methods. The method includes: when detecting the voice front end point of audio data, acquiring acquisition data consistent with the time stamp of the voice front end point; the voice front end point represents the starting input moment of user voice data, and the acquisition data is used to assist in understanding the problem characterized by the user voice data; based on the user voice data and the acquisition data, determining the answer result corresponding to the user voice data.

[0160] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0161] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A voice interaction method, characterized in that: include: When a voice front-end point of the audio data is detected, acquiring collected data consistent with a timestamp of the voice front-end point; The voice front-end point represents the starting input time of the user's voice data, and the collected data is used to assist in understanding the problem represented by the user's voice data; Based on the user voice data and the collected data, an answer result corresponding to the user voice data is determined.

2. The voice interaction method according to claim 1, characterized in that: When a voice front-end point of the audio data is detected, acquiring collected data consistent with the timestamp of the voice front-end point includes: When a voice front-end point of the audio data input in real time is detected, first collected data consistent with the timestamp of the voice front-end point is obtained from the collected data, and / or a collection device is controlled to collect second collected data consistent with the timestamp of the voice front-end point, and the second collected data is obtained; Wherein, when the first collected data is obtained, the collected data includes the first collected data, and when the second collected data is obtained, the collected data includes the second collected data.

3. The voice interaction method according to claim 1, characterized in that: The determining, based on the user voice data and the collected data, an answer result corresponding to the user voice data includes: When the intention recognition result of the user voice data is a preset intention recognition result, comprehensively determining an answer result corresponding to the user voice data based on the user voice data and the collected data; After acquiring the collected data consistent with the timestamp of the voice front-end point, the method further includes: In a case where the intention recognition result of the user voice data is not a preset intention recognition result, an answer result corresponding to the user voice data is determined based on the user voice data.

4. The voice interaction method according to any one of claims 1 to 3, characterized in that: After acquiring the collected data consistent with the timestamp of the voice front-end point, the method further includes: When it is determined that the collected data is not needed based on the intention recognition result of the user voice data, performing data post-processing on the collected data; The data post-processing includes at least one of the following: In the case that there is data that needs to be preprocessed in the collected data, canceling the data preprocessing process of the collected data; Delete the deletable data in the collected data.

5. The voice interaction method according to any one of claims 1 to 3, characterized in that: The voice interaction method is applied to a processor in a car; In the case where the voice front-end point of the audio data is detected, before acquiring the collected data consistent with the timestamp of the voice front-end point, the method further includes: When the vehicle is in a voice awakening state and the vehicle is in a driving state, collecting audio data in real time; Perform speech front-end detection on the audio data.

6. The voice interaction method according to any one of claims 1 to 3, characterized in that: The user voice data is determined based on the following method: When a voice front end point of the audio data is detected, the voice front end point is used as the starting acquisition time to acquire input voice data frames until the endpoint stops acquiring after the voice is detected; Determining the user voice data based on the collected voice data frames; The voice back endpoint indicates the time when the user voice data input is terminated.

7. A voice interaction device, characterized in that: include: A data acquisition module, for acquiring, when a voice front-end point of audio data is detected, collected data consistent with a timestamp of the voice front-end point; The voice front-end point represents the starting input time of the user's voice data, and the collected data is used to assist in understanding the problem represented by the user's voice data; The answer determination module is used to determine the answer result corresponding to the user voice data based on the user voice data and the collected data.

8. A movable device, characterized in that: include: An audio acquisition device, wherein the audio acquisition device is used to acquire audio data; A processor, wherein the processor is used to execute the voice interaction method as described in any one of claims 1 to 6.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the voice interaction method according to any one of claims 1 to 6 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the voice interaction method according to any one of claims 1 to 6 is implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the voice interaction method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Speech signal processing method, device, terminal device and medium

    CN108053822A

  • Multi-mode voice endpoint detection method and device, vehicle-mounted terminal and storage medium

    CN113255556A

  • Voice processing method and device, terminal and program product

    CN115440205A

  • Tiredness detection early warning method and system and computing device

    CN116013373A

  • Sample audio data acquisition method, speech recognition method and related device

    CN117894300A

Cited By

  • Voice reply method and device, equipment, storage medium and program product

    CN121366574A