A voice collection method, system, vehicle and storage medium
Patent Information
- Application Number
- CN202310361310.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-06
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-04-06
AI Technical Summary
[0004]鉴于上述问题,本申请实施例提供了一种语音采集方法、系统、车辆以及存储介质,以解决现有技术中在采集用户的语音时,会将背景杂音等干扰音一同采集,从而导致对用户的语音识别的准确度较低的问题
[0086]本申请实施例提供了一种语音采集方法,所述方法包括:在启动语音采集页面时,分别检测音频焦点的控制权和麦克风焦点的控制权是否被其他场景占用;在检测到所述音频焦点的控制权被其他场景占用的情况下,判断占用所述音频焦点的控制权的场景的重要级;在判定占用所述音频焦点的控制权的场景的重要级不高于预设重要级的情况下,将所述音频焦点的控制权移交给所述语音采集页面;在检测到所述麦克风焦点的控制权被所述其他场景占用的情况下,判断占用所述麦克风焦点的控制权的场景的重要级;在判定占用所述麦克风焦点的控制权的场景的重要级不高于所述预设重要级的情况下,将所述麦克风焦点的控制权移交给所述语音采集页面;在所述语音采集页面占用所述音频焦点的控制权和所述麦克风焦点的控制权的情况下,采集用户发出的语音。本申请实施例提供了一种语音采集方法,通过在采集用户的语音时,占用音频焦点的控制权和麦克风焦点的控制权,从而使得在语音采集页面下采集用户的语音时,减少采集过程中杂音的干扰,确保了采集的用户的语音的纯净度,进而有助于提高对用户的语音识别的准确性。
Smart Images

Figure CN116389975B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, and in particular to a speech acquisition method, system, vehicle, and storage medium. Background Technology
[0002] With the continuous development of speech recognition technology, more and more cars are using it, especially after the addition of speech recognition technology to the vehicle's in-vehicle multimedia host. By collecting user feedback information and performing speech recognition on the feedback information, it is possible to quickly analyze the current problems. However, when collecting the audio of user voice feedback, there are various background noises and irrelevant sounds in the surrounding environment, which are collected together with the audio of user feedback. This makes it impossible for ASR (Automatic Speech Recognition) to accurately recognize the user's feedback content.
[0003] To address the aforementioned issues, this application proposes a voice acquisition method. Summary of the Invention
[0004] In view of the above problems, this application provides a voice acquisition method, system, vehicle, and storage medium to solve the problem that in the prior art, when acquiring user voice, background noise and other interference noise are also acquired, resulting in low accuracy of user voice recognition.
[0005] A first aspect of this application provides a voice acquisition method, which is applied to an in-vehicle multimedia host in a vehicle, the method comprising:
[0006] When the voice capture page is launched, it is checked whether the control of the audio focus and the control of the microphone focus are occupied by other scenes.
[0007] If it is detected that the control of the audio focus is occupied by another scene, the importance level of the scene occupying the control of the audio focus is determined;
[0008] If the importance level of the scene occupying the control of the audio focus is determined to be no higher than the preset importance level, the control of the audio focus will be transferred to the voice acquisition page;
[0009] If it is detected that control of the microphone focus is occupied by other scenarios, the importance level of the scenario occupying control of the microphone focus is determined;
[0010] If the importance level of the scenario occupying the control of the microphone focus is determined to be no higher than the preset importance level, the control of the microphone focus is transferred to the voice acquisition page;
[0011] When the voice capture page occupies control of both the audio focus and the microphone focus, it captures the user's voice.
[0012] Optionally, it also includes:
[0013] If the importance level of the scene occupying the control of the audio focus is determined to be higher than the preset importance level, the audio volume output of the scene occupying the control of the audio focus is weakened;
[0014] If the importance level of the scenario that occupies the control of the microphone focus is determined to be higher than the preset importance level, it is detected whether the control of the microphone focus has been released;
[0015] If control of the microphone focus is released, control of the microphone focus is transferred to the voice acquisition page.
[0016] Optionally, it also includes:
[0017] If the voice acquisition page occupies control of the audio focus and the microphone focus, and if a request is detected that a scenario with an importance level higher than the preset importance level occupies control of the audio focus and / or the microphone focus, then control of the audio focus and / or the microphone focus will be transferred to the scenario with the higher importance level.
[0018] Optionally, it also includes:
[0019] When the voice acquisition page is activated, responding to and executing voice commands issued by the user that include execution commands or voice commands to switch to other scenarios is prohibited.
[0020] Optionally, before capturing the user's voice when the voice capture page occupies control of both the audio focus and the microphone focus, the method further includes:
[0021] The audio focus and the microphone focus are fitted together to form a communication-type focus;
[0022] When the audio focus and the microphone focus are fitted together as the communication focus, the incoming call notification is transferred to the mobile terminal upon receipt.
[0023] Optionally, it also includes:
[0024] During the acquisition of the user's voice, if a request is detected that a scenario with an importance level higher than the preset importance level is occupying control of the audio focus and / or the microphone focus, the control of the audio focus and / or the microphone focus will be transferred from the voice acquisition page to the scenario with the importance level higher than the preset importance level.
[0025] After releasing control of the audio focus and / or the microphone focus in a scenario where the importance level is higher than the preset importance level, the control of the audio focus and / or the microphone focus is transferred to the voice acquisition page.
[0026] Optionally, a completion feedback button is displayed at a first preset position on the voice acquisition page; after the voice acquisition page is launched, it further includes:
[0027] In response to the detected trigger operation of the completion feedback button, the system jumps from the voice acquisition page to the voice recognition page and re-allows the response to and execution of the user's voice commands containing execution commands and voice commands to switch to other scenarios.
[0028] On the speech recognition page, the speech is recognized to obtain the corresponding text information.
[0029] The voice and text information are uploaded to the backend server.
[0030] Optionally, when the voice acquisition page occupies control of both the audio focus and the microphone focus, acquiring the user's voice includes:
[0031] Based on the trained speech recognition model, the recognition of human voices in environmental audio is enhanced and noise in the environmental audio is filtered out.
[0032] The user's voice is extracted from the ambient audio.
[0033] Optionally, it also includes:
[0034] The area that the vehicle's microphone can radiate is divided into multiple sound zones, and these multiple sound zones are prioritized.
[0035] When speech exists in at least two of the multiple speech regions, the speech of the higher-priority speech region among the at least two speech regions is collected.
[0036] Optionally, it also includes:
[0037] During the acquisition of the user's voice, in response to the detection of a command manually triggered by the user to exit the voice acquisition page, the user exits the voice acquisition page and the acquired voice is deleted.
[0038] Optionally, it also includes:
[0039] If the current network is weak and the voice cannot be uploaded during the uploading of the user's voice, the voice will be saved locally.
[0040] Upon the next launch of the voice acquisition page, in response to the detected instruction to continue uploading, the page will be accessed for voice recognition, and the locally saved voice will be recognized again.
[0041] Optionally, it also includes:
[0042] During the collection of the user's voice, in response to the detected background running instruction, the voice collection page is run in the background, and the collection of the user's voice is terminated after a preset time.
[0043] If the user returns to the voice capture page again, the user is prompted to provide feedback again, and in response to the detected return instruction, the user's voice is captured again.
[0044] A second aspect of this application provides a voice acquisition system, the system comprising:
[0045] The detection module is used to detect whether the control of audio focus and microphone focus is occupied by other scenes when the voice acquisition page is launched;
[0046] The first judgment module is used to determine the importance level of the scene occupying the control of the audio focus when it is detected that the control of the audio focus is occupied by another scene.
[0047] The first handover module is used to hand over the control of the audio focus to the voice acquisition page when it is determined that the importance level of the scene occupying the control of the audio focus is not higher than the preset importance level.
[0048] The second judgment module is used to determine the importance level of the scene occupying the control of the microphone focus when it is detected that the control of the microphone focus is occupied by the other scene;
[0049] The second handover module is used to hand over control of the microphone focus to the voice acquisition page when the importance level of the scene occupying the control of the microphone focus is not higher than the preset importance level.
[0050] The acquisition module is used to acquire the user's voice when the voice acquisition page occupies control of the audio focus and the microphone focus.
[0051] Optionally, it also includes:
[0052] The first determination submodule is used to weaken the audio volume output of the scene that occupies the control of the audio focus when the importance level of the scene that occupies the control of the audio focus is higher than the preset importance level.
[0053] The second determination submodule is used to detect whether the control of the microphone focus has been released when the importance level of the scene occupying the control of the microphone focus is higher than the preset importance level.
[0054] The first handover submodule is used to hand over control of the microphone focus to the voice acquisition page when it is detected that the control of the microphone focus has been released.
[0055] Optionally, it also includes:
[0056] The second handover submodule is used to transfer the control of the audio focus and / or the microphone focus to the scenario whose importance level is higher than the preset importance level when the voice acquisition page occupies the control of the audio focus and / or the microphone focus.
[0057] Optionally, it also includes:
[0058] The disable submodule is used to disable the response to and execution of voice commands issued by the user that include execution commands and voice commands that switch to other scenarios when the voice acquisition page is started.
[0059] Optionally, before capturing the user's voice when the voice capture page occupies control of both the audio focus and the microphone focus, the method further includes:
[0060] A fitting submodule is used to fit the audio focus and the microphone focus into a communication-type focus;
[0061] The third handover submodule is used to hand over the incoming call reminder to the mobile terminal after receiving the call reminder, when the audio focus and the microphone focus are fitted to the communication focus.
[0062] Optionally, it also includes:
[0063] The fourth handover submodule, during the acquisition of the user's voice, if it detects a request from a scene with an importance level higher than the preset importance level to occupy the control of the audio focus and / or the control of the microphone focus, it will transfer the control of the audio focus and / or the control of the microphone focus from the voice acquisition page to the scene with the importance level higher than the preset importance level.
[0064] The fifth handover submodule is used to transfer control of the microphone focus to the voice acquisition page after releasing control of the microphone focus in a scenario where the importance level is higher than the preset importance level.
[0065] Optionally, a completion feedback button is displayed at a first preset position on the voice acquisition page; after the voice acquisition page is launched, it further includes:
[0066] The first jump rotor module is used to jump from the voice acquisition page to the voice recognition page in response to the detected trigger operation of the completion feedback button, and to re-allow the response to and execution of the voice commands issued by the user that include execution commands and voice commands to switch to other scenarios.
[0067] The first recognition submodule is used to recognize the speech on the speech recognition page and obtain the text information corresponding to the speech.
[0068] The upload submodule is used to upload the voice and text information to the backend server.
[0069] Optionally, when the voice acquisition page occupies control of both the audio focus and the microphone focus, the voice output by the user is acquired. The acquisition module includes:
[0070] The filtering submodule is used to enhance the recognition of human voices in environmental audio and filter noise in the environmental audio based on the trained speech recognition model.
[0071] The separation submodule is used to separate the user's voice from the ambient audio.
[0072] Optionally, it also includes:
[0073] The segmentation submodule is used to divide the area that the vehicle's microphone can radiate into multiple sound zones and prioritize these multiple sound zones.
[0074] The first acquisition submodule is used to acquire the speech of the higher-priority speech region among the at least two speech regions when speech exists in at least two of the multiple speech regions.
[0075] Optionally, it also includes:
[0076] The exit submodule is used to exit the voice acquisition page and delete the acquired voice data in response to a user-triggered command to exit the voice acquisition page during the acquisition of the user's voice data.
[0077] Optionally, it also includes:
[0078] The storage submodule is used to store the voice recording locally if the current network is weak and the voice recording cannot be uploaded during the uploading of the user's voice recording.
[0079] The second recognition submodule is used to, upon the next launch of the voice acquisition page, respond to the detected instruction to continue uploading, enter the voice recognition page, and re-recognize the locally saved voice.
[0080] Optionally, it also includes:
[0081] The termination submodule is used to run the voice collection page in the background in response to a detected background running instruction during the collection of the user's voice, and terminate the collection of the user's voice after a preset time.
[0082] The second acquisition submodule is used to prompt the user to provide feedback again if the user returns to the voice acquisition page, and to re-acquire the user's voice in response to the detected return instruction.
[0083] A third aspect of this application provides a vehicle including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the voice acquisition method as described in the first aspect of this application.
[0084] A fourth aspect of this application provides a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the voice acquisition method as described in the first aspect of this application.
[0085] This application has the following advantages:
[0086] This application provides a voice acquisition method, comprising: upon activating a voice acquisition page, detecting whether control of the audio focus and microphone focus are occupied by other scenes; if control of the audio focus is detected to be occupied by other scenes, determining the importance level of the scene occupying the audio focus; if the importance level of the scene occupying the audio focus is determined to be no higher than a preset importance level, transferring control of the audio focus to the voice acquisition page; if control of the microphone focus is detected to be occupied by other scenes, determining the importance level of the scene occupying the microphone focus; if the importance level of the scene occupying the microphone focus is determined to be no higher than the preset importance level, transferring control of the microphone focus to the voice acquisition page; and acquiring the user's voice while the voice acquisition page occupies control of both the audio focus and microphone focus. This application provides a voice acquisition method that, by occupying control of the audio focus and microphone focus when acquiring user voice, reduces interference from background noise during the acquisition process, ensuring the purity of the acquired user voice, and thus helping to improve the accuracy of user voice recognition. Attached Figure Description
[0087] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0088] Figure 1 This is a flowchart illustrating the steps of a voice acquisition method provided in an embodiment of this application;
[0089] Figure 2 This is a schematic diagram of a voice acquisition page provided in an embodiment of this application;
[0090] Figure 3 This is a schematic diagram of a voice acquisition system provided in an embodiment of this application;
[0091] Figure 4 This is a schematic diagram of a vehicle provided in an embodiment of this application. Detailed Implementation
[0092] Exemplary embodiments of this application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0093] A first aspect of this application provides a voice acquisition method, the method being applied to a vehicle's in-vehicle multimedia host, referring to... Figure 1 The following is a flowchart of a voice acquisition method provided in an embodiment of this application, the method comprising:
[0094] Step S101: When starting the voice acquisition page, detect whether the control of the audio focus and the control of the microphone focus are occupied by other scenes.
[0095] Step S102: If it is detected that the control of the audio focus is occupied by other scenes, determine the importance level of the scene occupying the control of the audio focus;
[0096] Step S103: If the importance level of the scene occupying the control of the audio focus is not higher than the preset importance level, the control of the audio focus is transferred to the voice acquisition page.
[0097] Step S104: If it is detected that the control of the microphone focus is occupied by other scenes, determine the importance level of the scene occupying the control of the microphone focus;
[0098] Step S105: If the importance level of the scene occupying the control of the microphone focus is not higher than the preset importance level, the control of the microphone focus is transferred to the voice acquisition page.
[0099] Step S106: When the voice acquisition page occupies control of both the audio focus and the microphone focus, the voice emitted by the user is acquired.
[0100] In this embodiment, the voice assistant in the in-vehicle multimedia host needs to be activated in advance. Specifically, the voice assistant can be activated by manually clicking the physical button on the steering wheel, or by voice with the microphone on. Voice activation can involve the user saying a wake-up phrase such as "Hello, Xiao Ou" or "Hello, Euler." After the voice assistant is activated, the user can continue to open the feedback application via voice. Specifically, the user can open the feedback application by saying the application's name preceded by "open." For example, if the application is called "Smart Euler," the user can say: "Open Smart Euler." Furthermore, considering the application's relevance to feedback, a generalized term is associated with the application. Therefore, when the user says "I want to provide feedback" or "This application is so laggy, I want to provide feedback," the "Smart Euler" application can also be opened.
[0101] In this embodiment of the application, in the wake-free mode, that is, the voice assistant is in a state of listening in the background for a long time, there is no need to wake up the voice assistant in advance. The user can open the problem feedback application by saying "Open Smart Euler" or "I want to give feedback" or "Why is this application so laggy? I want to give feedback", etc., which are directly related to the problem feedback application.
[0102] Furthermore, after opening the feedback application, the system can determine whether the user genuinely intends to provide feedback based on their voice. Specifically, this can be done by setting a target intent in a feedback-related vocabulary and simultaneously using existing data generalization techniques to obtain generalized words corresponding to the target intent. In practice, the target intent includes target keywords. When the user's voice matches a target keyword or its corresponding generalized word, it is determined that the user intends to provide feedback. For example, using "difficult to use" as a preset target keyword, generalized words for "difficult to use" could include "very difficult to use," "not easy to use," "too difficult to use," etc. If the user's voice is "This application is difficult to use today," then the user's voice matches the keyword "difficult to use," indicating an intention to provide feedback. Alternatively, if the user's voice is "This application is too difficult to use today," then the user's voice matches the generalized word corresponding to the target keyword "too difficult to use," indicating an intention to provide feedback.
[0103] Furthermore, when it is inferred that the user may intend to provide feedback, semantic clarification is needed to further determine whether the user truly needs feedback in order to ensure the accuracy of the service. After receiving confirmation from the user, the voice feedback page is activated. In practical applications, the semantic clarification information represents a clear expression to confirm with the user again whether they intend to provide feedback, specifically by outputting semantic clarification information such as "Do you need to provide feedback? If so, please say yes."
[0104] Furthermore, after launching the voice feedback page, prompt text is displayed on the voice feedback page to encourage users to describe the problem from multiple perspectives.
[0105] Preferably, a "Start Feedback" button is displayed on the voice feedback page. When a trigger operation on the "Start Feedback" button is detected, the current page will jump from the voice feedback page to the voice acquisition page.
[0106] Furthermore, upon startup of the voice capture page, it checks whether control of the audio focus and microphone focus is already occupied by other scenarios. If audio focus control is detected as being occupied, the importance level of the scenario occupying the audio focus is determined. If the importance level of the scenario occupying the audio focus is not higher than a preset importance level, control of the audio focus is transferred to the voice capture page. Similarly, if microphone focus control is detected as being occupied by other scenarios, the importance level of the scenario occupying the microphone focus is determined. If the importance level of the scenario occupying the microphone focus is not higher than a preset importance level, control of the microphone focus is transferred to the voice capture page. In practical applications, a scenario can consist of one or more applications. For example, map applications belong to navigation scenarios; audio and video applications belong to multimedia scenarios; vehicle alarm systems and E-CALL (Emergency Call) belong to security scenarios, etc.
[0107] It is important to note that during the user's voice capture, the voice capture page must simultaneously control both the audio focus and the microphone focus. The audio focus is used to control the audio output, while the microphone focus is used to control the voice input. This reduces interference from audio output from other scenarios during the user's voice capture, thereby improving the purity of the captured user's voice.
[0108] In a preferred embodiment, if it is determined that the importance level of the scene occupying the control of the audio focus is higher than the preset importance level, the audio volume output by the scene occupying the control of the audio focus is weakened;
[0109] If the importance level of the scenario that occupies the control of the microphone focus is determined to be higher than the preset importance level, it is detected whether the control of the microphone focus has been released;
[0110] If control of the microphone focus is released, control of the microphone focus is transferred to the voice acquisition page.
[0111] Specifically, if a user wakes up the voice function and needs to provide feedback and start the voice collection page, or if a user needs to provide feedback and start the voice collection page in wake-free mode, when it is detected that the control of the audio focus is occupied by another scene, if the importance level of the scene occupying the control of the audio focus is higher than the preset importance level, the volume of the audio being output by the scene occupying the control of the audio focus will be weakened.
[0112] When it is detected that the microphone focus control is occupied by another scene, if the importance level of the scene currently occupying the microphone focus control is higher than the preset importance level, then even if the user needs to provide feedback, the voice capture page can only take over the microphone focus control after the scene currently occupying the microphone focus releases the microphone focus control.
[0113] For example, in some important scenarios, such as when the vehicle's alarm system is of higher importance than the preset importance level, if the vehicle's alarm system occupies the audio focus and microphone focus after the voice capture page is activated, the voice capture page will not be able to directly take control of the audio focus and microphone focus. Instead, it will need to reduce the volume of the audio being output by the vehicle's alarm system. The control of the audio focus and microphone focus can only be transferred to the voice capture page after the alarm ends and the control of the audio focus and microphone focus is released.
[0114] Understandably, when a user provides feedback, if no scene has control over the audio focus and microphone focus, the voice capture page will directly take over control of the audio focus and microphone focus.
[0115] In another preferred embodiment, if the voice acquisition page occupies control of the audio focus and the microphone focus, and if a request is detected that a scenario with an importance level higher than the preset importance level occupies control of the audio focus and / or the microphone focus, then control of the audio focus and / or the microphone focus is transferred to the scenario with the higher importance level.
[0116] Specifically, when the voice acquisition page is in control of both the audio focus and the microphone focus, if a request is detected that a scenario with a higher importance level than the preset importance level is taking over the control of the audio focus and / or the microphone focus, for example, a security scenario such as a vehicle alarm system taking over the control of the audio focus and / or the microphone focus, the control of the audio focus and / or the microphone focus needs to be transferred to the security scenario such as the vehicle alarm system, which has a higher importance level than the preset importance level. At the same time, a prompt message such as "Voice acquisition has been terminated" should be displayed on the voice acquisition page.
[0117] In another preferred embodiment, when the voice acquisition page is activated, responding to and executing voice commands issued by the user that include execution commands and voice commands that switch to other scene pages is prohibited.
[0118] Specifically, when the voice capture page is active, responding to and executing user-issued voice commands containing execution commands is prohibited. Examples include voice commands to open or close car windows, open or close the air conditioning, or activate or deactivate certain functions within certain applications. Furthermore, it is impossible to switch to other applications via voice commands while on the current voice capture page. In practice, a manual switch to another application can be used to force a transition. After manually switching the current voice capture page to another application, voice capture will terminate, and the captured audio will not be saved.
[0119] In another preferred embodiment, a completion feedback button is displayed at a first preset position on the voice acquisition page; after the voice acquisition page is launched, in response to the detected trigger operation on the completion feedback button, the user is redirected from the voice acquisition page to the voice recognition page, and the user is allowed to respond to and execute voice commands issued by the user that include execution commands and voice commands to switch to other scenarios.
[0120] like Figure 2 The image shown is a schematic diagram of a voice acquisition page provided in this application. Figure 2As shown, a completion feedback button is displayed in the first preset position on the voice acquisition page. When a trigger operation on the completion feedback button is detected, the system jumps from the voice acquisition page to the voice recognition page in response to the detected trigger operation. Alternatively, a preset countdown is displayed on the voice acquisition page, and after the preset countdown ends, the system jumps from the voice acquisition page to the voice recognition page. At this point, the user's feedback has been collected, and therefore, the system is allowed to respond to and execute user-issued voice commands containing execution commands and voice commands to switch to other scenarios for subsequent voice control. On the voice recognition page, the voice is recognized to obtain the corresponding text information. In practical applications, it is necessary to confirm and verify the text information corresponding to the user's voice; and then upload the voice and the confirmed and verified text information to the backend server so that backend staff can analyze the user's feedback.
[0121] In one optional embodiment, while uploading the user's voice to the server, the vehicle's condition information and the user's account information at the time of the user's feedback are also uploaded to the server. The vehicle condition information includes at least one of the following: vehicle speed, mileage data, and geographical location information. This allows the backend to perform detailed analysis of the user's feedback. Specifically, it can attribute the problems in the user's voice to specific locations. For example, based on geographical location, it can be analyzed where the problem occurs; for instance, if a user reports a certain problem, it may be found that it only occurs in mountainous areas in the south, but not in plains areas in the north. Based on mileage, it can be analyzed whether certain problems occur or intensify when the user reaches or exceeds a certain mileage range. Based on vehicle speed, it can be analyzed whether network problems or noise problems become more apparent above a certain threshold or are evenly distributed across different speed ranges.
[0122] Furthermore, based on the user's account information, the system queries the user's historical feedback information and / or pushes solutions corresponding to the user's reported issues to the user's terminal. Specifically, based on the user's account information, the system can find the record of the person who raised the issue. Each vehicle's user account is unique and contains the user's identity information (name, mobile phone number, etc.). If staff need to contact the user, they can find the user through the account in the background. The system can also find the historical feedback information under the account and send the corresponding solutions to the corresponding user terminal based on the user's contact information under the account.
[0123] In another preferred embodiment, before capturing the user's voice when the voice capture page occupies control of both the audio focus and the microphone focus, the method further includes:
[0124] The audio focus and the microphone focus are fitted together as a communication focus; when the audio focus and the microphone focus are fitted together as the communication focus, the incoming call reminder is transferred to the mobile terminal after receiving the incoming call reminder.
[0125] Specifically, when the voice capture page gains control of both the audio and microphone focus, before capturing the user's voice, the audio and microphone focus are matched as a communication-type focus. Once the audio and microphone focus are matched as a communication-type focus, upon receiving an incoming call notification, the notification is handed over to the mobile terminal. If the voice capture process is interrupted by an incoming call notification, voice data loss will occur. Matching the audio and microphone focus as a communication-type focus in this situation avoids interference from the incoming call notification, thus ensuring the integrity and accuracy of the voice data.
[0126] In another preferred embodiment, during the acquisition of the user's voice, if a request is detected that a scenario with an importance level higher than the preset importance level is occupying control of the audio focus and / or the microphone focus, the control of the audio focus and / or the microphone focus is transferred from the voice acquisition page to the scenario with the importance level higher than the preset importance level; after the scenario with the importance level higher than the preset importance level releases the control of the audio focus and / or the microphone focus, the control of the audio focus and / or the microphone focus is transferred back to the voice acquisition page.
[0127] Specifically, when capturing user voice on the voice capture page, the page controls both the audio focus and the microphone focus. Controlling the audio focus allows for timely output of voice prompts when needed, and also restricts audio output from other scenarios, minimizing its impact on the captured user voice. Furthermore, controlling the microphone focus ensures clear voice capture, effectively improving the quality and reliability of the captured voice.
[0128] In this embodiment, when another scene attempts to occupy control of the audio focus and / or microphone focus, it is first determined whether the importance level of that scene is higher than a preset importance level. If the importance level of that scene is not higher than the preset importance level, then that scene cannot occupy control of the audio focus and / or microphone focus, meaning the voice acquisition process will not be affected. When a scene with an importance level higher than the preset importance level attempts to occupy control of the audio focus and / or microphone focus, the current voice acquisition page will release the occupied control of the audio focus and / or microphone focus and transfer control of the audio focus and / or microphone focus to the scene with an importance level higher than the preset importance level, to ensure... Scenarios with an importance level higher than the preset importance level can promptly acquire control of the audio focus and / or microphone focus. It should be noted that during voice acquisition, if the control of the audio focus and / or microphone focus on the voice acquisition page is occupied by a scenario with an importance level higher than the preset importance level, the voice acquisition process will be terminated, and a "Voice acquisition has been terminated" message will be displayed on the voice acquisition page. After the control of the audio focus and / or microphone focus is released, the control of the audio focus and / or microphone focus will be transferred back to the voice acquisition page, and the voice acquisition page will reacquire the user's voice after regaining control of the audio focus and / or microphone focus.
[0129] In another preferred embodiment, based on the trained speech recognition model, the recognition of human voices in the ambient audio is enhanced and noise in the ambient audio is filtered out; from the filtered ambient audio, the speech emitted by the user is separated.
[0130] Specifically, by continuously training on the training set, various background noises and interfering sound sources are identified, such as wind sounds, window opening sounds, music sounds, key press sounds, coughing sounds, etc. In this way, these noises can be filtered out from the environmental audio containing noises such as wind sounds, window opening sounds, music sounds, key press sounds, coughing sounds, etc., so as to better separate and extract human voices from various interference sounds, and provide high-quality audio materials for subsequent speech recognition.
[0131] In another preferred embodiment, the area that the vehicle's microphone can radiate is divided into multiple sound zones, and the multiple sound zones are assigned priorities; when there is speech in at least two of the multiple sound zones, the speech in the sound zone with the higher priority among the at least two sound zones is collected first.
[0132] Specifically, the area covered by the vehicle's microphone is divided into multiple sound zones. These zones could be designated as the driver's cab, front passenger cab, right rear passenger seat, and left rear passenger seat. The priority order for each zone is: driver's cab > front passenger cab > left rear passenger seat > right rear passenger seat. During voice capture, if voice is present in two zones, the higher-priority zone is captured; if more than two zones are present, the highest-priority zone is captured. This ensures that when multiple zones are present, only the highest-priority zone is captured, reducing interference from other zones and improving the accuracy of voice recognition.
[0133] In another preferred embodiment, during the acquisition of the user's voice, in response to detecting a command manually triggered by the user to exit the voice acquisition page, the user exits the voice acquisition page and the acquired voice is deleted.
[0134] In another preferred embodiment, if the current network is weak and the voice cannot be uploaded during the uploading of the user's voice, the voice is saved locally. Upon the next startup of the voice acquisition page, in response to the detected instruction to continue uploading, the voice recognition page is accessed, and the locally saved voice is re-recognized.
[0135] If the current network condition is poor or weak during the upload of the user's voice, and the user's voice cannot be uploaded, the collected voice will be saved locally. When the user starts the voice collection page again, the user will be prompted whether to continue uploading. If the user's instruction to continue uploading is detected, the system will respond to the instruction, enter the voice recognition page, recognize the saved voice, confirm and verify the recognized text information, and then upload it to the backend server.
[0136] Optionally, upon detecting a user's instruction to stop uploading, the saved audio is deleted in response to the instruction.
[0137] In another preferred embodiment, during the collection of the user's voice, in response to a detected background running instruction, the voice collection page is run in the background, and the collection of the user's voice is terminated after a preset duration.
[0138] If the user returns to the voice capture page again, the user is prompted to provide feedback again, and in response to the detected return instruction, the user's voice is captured again.
[0139] Specifically, when a background running instruction is detected, the voice collection page continues to run in the background in response to the detected background running instruction. However, after a preset time, the collection of the user's voice will be terminated, thereby avoiding the continuous collection of irrelevant voice information of the user in the background.
[0140] Furthermore, if the user returns to the voice collection page, they will be prompted that the feedback has been terminated and they need to submit feedback again. Upon receiving the user's return instruction, the system will respond to the user's return instruction, return to the initial interface of the voice collection page, and restart the feedback process.
[0141] This application provides a voice acquisition method, comprising: upon activating a voice acquisition page, detecting whether control of the audio focus and microphone focus are occupied by other scenes; if control of the audio focus is detected to be occupied by other scenes, determining the importance level of the scene occupying the audio focus; if the importance level of the scene occupying the audio focus is determined to be no higher than a preset importance level, transferring control of the audio focus to the voice acquisition page; if control of the microphone focus is detected to be occupied by other scenes, determining the importance level of the scene occupying the microphone focus; if the importance level of the scene occupying the microphone focus is determined to be no higher than the preset importance level, transferring control of the microphone focus to the voice acquisition page; and acquiring the user's voice while the voice acquisition page occupies control of both the audio focus and microphone focus. This application provides a voice acquisition method that, by occupying control of the audio focus and microphone focus when acquiring user voice, reduces interference from background noise during the acquisition process, ensuring the purity of the acquired user voice, and thus helping to improve the accuracy of user voice recognition.
[0142] Based on the same inventive concept, a second aspect of the embodiments of this application provides a voice acquisition system, such as... Figure 3 As shown, the system includes:
[0143] The detection module 201 is used to detect whether the control of the audio focus and the control of the microphone focus are occupied by other scenes when the voice acquisition page is started.
[0144] The first judgment module 202 is used to determine the importance level of the scene occupying the control of the audio focus when it is detected that the control of the audio focus is occupied by other scenes;
[0145] The first handover module 203 is used to hand over the control of the audio focus to the voice acquisition page when it is determined that the importance level of the scene occupying the control of the audio focus is not higher than the preset importance level.
[0146] The second judgment module 204 is used to determine the importance level of the scene occupying the control of the microphone focus when it is detected that the control of the microphone focus is occupied by the other scene;
[0147] The second handover module 205 is used to hand over the control of the microphone focus to the voice acquisition page when it is determined that the importance level of the scene occupying the control of the microphone focus is not higher than the preset importance level.
[0148] The acquisition module 206 is used to acquire the user's voice when the voice acquisition page occupies control of the audio focus and the microphone focus.
[0149] Optionally, it also includes:
[0150] The first determination submodule is used to weaken the audio volume output of the scene that occupies the control of the audio focus when the importance level of the scene that occupies the control of the audio focus is higher than the preset importance level.
[0151] The second determination submodule is used to detect whether the control of the microphone focus has been released when the importance level of the scene occupying the control of the microphone focus is higher than the preset importance level.
[0152] The first handover submodule is used to hand over control of the microphone focus to the voice acquisition page when it is detected that the control of the microphone focus has been released.
[0153] Optionally, it also includes:
[0154] The second handover submodule is used to transfer the control of the audio focus and / or the microphone focus to the scenario whose importance level is higher than the preset importance level when the voice acquisition page occupies the control of the audio focus and / or the microphone focus.
[0155] Optionally, it also includes:
[0156] The disable submodule is used to disable the response to and execution of voice commands issued by the user that include execution commands and voice commands that switch to other scenarios when the voice acquisition page is started.
[0157] Optionally, before capturing the user's voice when the voice capture page occupies control of both the audio focus and the microphone focus, the method further includes:
[0158] A fitting submodule is used to fit the audio focus and the microphone focus into a communication-type focus;
[0159] The third handover submodule is used to hand over the incoming call reminder to the mobile terminal after receiving the call reminder, when the audio focus and the microphone focus are fitted to the communication focus.
[0160] Optionally, it also includes:
[0161] The fourth handover submodule, during the acquisition of the user's voice, if it detects a request from a scene with an importance level higher than the preset importance level to occupy the control of the audio focus and / or the control of the microphone focus, it will transfer the control of the audio focus and / or the control of the microphone focus from the voice acquisition page to the scene with the importance level higher than the preset importance level.
[0162] The fifth handover submodule is used to transfer control of the microphone focus to the voice acquisition page after releasing control of the microphone focus in a scenario where the importance level is higher than the preset importance level.
[0163] Optionally, a completion feedback button is displayed at a first preset position on the voice acquisition page; after the voice acquisition page is launched, it further includes:
[0164] The first jump rotor module is used to jump from the voice acquisition page to the voice recognition page in response to the detected trigger operation of the completion feedback button, and to re-allow the response to and execution of the voice commands issued by the user that include execution commands and voice commands to switch to other scenarios.
[0165] The first recognition submodule is used to recognize the speech on the speech recognition page and obtain the text information corresponding to the speech.
[0166] The upload submodule is used to upload the voice and text information to the backend server.
[0167] Optionally, when the voice acquisition page occupies control of both the audio focus and the microphone focus, the voice output by the user is acquired. The acquisition module 206 includes:
[0168] The filtering submodule is used to enhance the recognition of human voices in environmental audio and filter noise in the environmental audio based on the trained speech recognition model.
[0169] The separation submodule is used to separate the user's voice from the ambient audio.
[0170] Optionally, it also includes:
[0171] The segmentation submodule is used to divide the area that the vehicle's microphone can radiate into multiple sound zones and prioritize these multiple sound zones.
[0172] The first acquisition submodule is used to acquire the speech of the higher-priority speech region among the at least two speech regions when speech exists in at least two of the multiple speech regions.
[0173] Optionally, it also includes:
[0174] The exit submodule is used to exit the voice acquisition page and delete the acquired voice data in response to a user-triggered command to exit the voice acquisition page during the acquisition of the user's voice data.
[0175] Optionally, it also includes:
[0176] The storage submodule is used to store the voice recording locally if the current network is weak and the voice recording cannot be uploaded during the uploading of the user's voice recording.
[0177] The second recognition submodule is used to, upon the next launch of the voice acquisition page, respond to the detected instruction to continue uploading, enter the voice recognition page, and re-recognize the locally saved voice.
[0178] Optionally, it also includes:
[0179] The termination submodule is used to run the voice collection page in the background in response to a detected background running instruction during the collection of the user's voice, and terminate the collection of the user's voice after a preset time.
[0180] The second acquisition submodule is used to prompt the user to provide feedback again if the user returns to the voice acquisition page, and to re-acquire the user's voice in response to the detected return instruction.
[0181] Based on the same inventive concept, a third aspect of the embodiments of this application provides a vehicle 100, such as... Figure 4 As shown, it includes a memory 110, a processor 120, and a computer program stored on the memory 110. The processor 120 executes the computer program to implement the voice acquisition method described in the first aspect of the embodiments of this application.
[0182] Based on the same inventive concept, in a fourth aspect of this application, a computer-readable storage medium is provided, on which a computer program / instruction is stored, which, when executed by a processor, implements the voice acquisition method as described in the first aspect of this application.
[0183] Each embodiment in this specification focuses on the differences from other embodiments. For the same or similar parts between the embodiments, please refer to each other.
[0184] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0185] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0186] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0187] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1The steps of the function specified in one or more boxes.
[0188] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0189] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0190] The above provides a detailed description of the provided voice acquisition method, system, vehicle, and storage medium. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A voice acquisition method, characterized in that, The method includes: When the voice capture page is launched, it is checked whether the control of the audio focus and the control of the microphone focus are occupied by other scenes. If it is detected that the control of the audio focus is occupied by another scene, the importance level of the scene occupying the control of the audio focus is determined; If the importance level of the scene occupying the control of the audio focus is determined to be no higher than the preset importance level, the control of the audio focus will be transferred to the voice acquisition page; If it is detected that control of the microphone focus is occupied by other scenarios, the importance level of the scenario occupying control of the microphone focus is determined; If the importance level of the scenario occupying the control of the microphone focus is determined to be no higher than the preset importance level, the control of the microphone focus is transferred to the voice acquisition page; When the voice acquisition page occupies control of both the audio focus and the microphone focus, it acquires the user's voice. The voice can be uploaded to a backend server. When uploading the voice to the backend server, the vehicle's condition information at the time of the user's feedback and the user's account information are also uploaded to the backend server. The vehicle condition information includes at least one of the following: vehicle speed, mileage data, and geographical location information. The backend server analyzes the user's feedback based on the vehicle condition information and attributes the cause of the user's feedback. The method further includes: If the importance level of the scene occupying the control of the audio focus is determined to be higher than the preset importance level, the audio volume output of the scene occupying the control of the audio focus is weakened; If the importance level of the scenario that occupies the control of the microphone focus is determined to be higher than the preset importance level, it is detected whether the control of the microphone focus has been released; If control of the microphone focus is released, control of the microphone focus is transferred to the voice acquisition page.
2. The voice acquisition method according to claim 1, characterized in that, Also includes: If the voice acquisition page occupies control of the audio focus and the microphone focus, and if a request is detected that a scenario with an importance level higher than the preset importance level occupies control of the audio focus and / or the microphone focus, then control of the audio focus and / or the microphone focus will be transferred to the scenario with the higher importance level.
3. The voice acquisition method according to claim 1, characterized in that, Also includes: When the voice acquisition page is activated, responding to and executing voice commands issued by the user that include execution commands or voice commands to switch to other scenarios is prohibited.
4. The voice acquisition method according to claim 1, characterized in that, Before capturing the user's voice when the voice capture page occupies control of both the audio focus and the microphone focus, the method further includes: The audio focus and the microphone focus are fitted together to form a communication-type focus; When the audio focus and the microphone focus are fitted together as the communication focus, the incoming call notification is transferred to the mobile terminal upon receipt.
5. The voice acquisition method according to claim 1, characterized in that, Also includes: During the acquisition of the user's voice, if a request is detected that a scenario with an importance level higher than the preset importance level is occupying control of the audio focus and / or the microphone focus, the control of the audio focus and / or the microphone focus will be transferred from the voice acquisition page to the scenario with the importance level higher than the preset importance level. After releasing control of the audio focus and / or the microphone focus in a scenario where the importance level is higher than the preset importance level, the control of the audio focus and / or the microphone focus is transferred to the voice acquisition page.
6. The voice acquisition method according to claim 3, characterized in that, The voice acquisition page displays a completion feedback button in a first preset position; after the voice acquisition page is launched, it also includes: In response to the detected trigger operation of the completion feedback button, the system jumps from the voice acquisition page to the voice recognition page and re-allows the response to and execution of the user's voice commands containing execution commands and voice commands to switch to other scenarios. On the speech recognition page, the speech is recognized to obtain the corresponding text information. The voice and text information are uploaded to the backend server.
7. The voice acquisition method according to claim 1, characterized in that, When the voice acquisition page occupies control of both the audio focus and the microphone focus, the process of acquiring the user's voice includes: Based on the trained speech recognition model, the recognition of human voices in environmental audio is enhanced and noise in the environmental audio is filtered out. The user's voice is extracted from the ambient audio.
8. The voice acquisition method according to claim 1, characterized in that, Also includes: The area that the vehicle's microphone can radiate is divided into multiple sound zones, and these multiple sound zones are prioritized. When speech exists in at least two of the multiple speech regions, the speech of the higher-priority speech region among the at least two speech regions is collected.
9. The voice acquisition method according to any one of claims 1-8, characterized in that, Also includes: During the acquisition of the user's voice, in response to the detection of a command manually triggered by the user to exit the voice acquisition page, the user exits the voice acquisition page and the acquired voice is deleted.
10. The voice acquisition method according to any one of claims 1-8, characterized in that, Also includes: If the current network is weak and the voice cannot be uploaded during the uploading of the user's voice, the voice will be saved locally. Upon the next launch of the voice acquisition page, in response to the detected instruction to continue uploading, the page will be accessed for voice recognition, and the locally saved voice will be recognized again.
11. The voice acquisition method according to any one of claims 1-8, characterized in that, Also includes: During the collection of the user's voice, in response to the detected background running instruction, the voice collection page is run in the background, and the collection of the user's voice is terminated after a preset time. If the user returns to the voice capture page again, the user is prompted to provide feedback again, and in response to the detected return instruction, the user's voice is captured again.
12. A voice acquisition system, characterized in that, The system includes: The detection module is used to detect whether the control of audio focus and microphone focus is occupied by other scenes when the voice acquisition page is launched; The first judgment module is used to determine the importance level of the scene occupying the control of the audio focus when it is detected that the control of the audio focus is occupied by another scene. The first handover module is used to hand over the control of the audio focus to the voice acquisition page when it is determined that the importance level of the scene occupying the control of the audio focus is not higher than the preset importance level. The second judgment module is used to determine the importance level of the scene occupying the control of the microphone focus when it is detected that the control of the microphone focus is occupied by the other scene; The second handover module is used to hand over control of the microphone focus to the voice acquisition page when the importance level of the scene occupying the control of the microphone focus is not higher than the preset importance level. The acquisition module is used to acquire the user's voice when the voice acquisition page occupies control of both the audio focus and the microphone focus; wherein, the voice can be uploaded to a backend server, and when uploading the voice to the backend server, the vehicle's condition information at the time of the user's feedback and the user's account information are also uploaded to the backend server, the vehicle condition information including at least one of the following: vehicle speed, mileage data, and geographical location information; the backend server analyzes the user's feedback based on the vehicle condition information and attributes the cause of the user's feedback; The system also includes: The first determination submodule is used to weaken the audio volume output of the scene that occupies the control of the audio focus when the importance level of the scene is determined to be higher than the preset importance level. The second determination submodule is used to detect whether the control of the microphone focus has been released when the importance level of the scene occupying the control of the microphone focus is higher than the preset importance level. The first handover submodule is used to hand over control of the microphone focus to the voice acquisition page when it is detected that the control of the microphone focus has been released.
13. A vehicle comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the voice acquisition method as described in any one of claims 1 to 11.
14. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the voice acquisition method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
In-vehicle multi-sound-zone pickup method and device, electronic equipment and storage medium
CN111599357A
Audio focus management method, vehicle device and computer readable storage medium
CN115878067A