Voice processing method and device and storage medium

By installing a dual-mode Bluetooth module on smart helmets and vehicles, and using the HFP protocol to establish voice channels, the driver solves the inconvenience and safety risks of operating the mobile phone during driving, and realizes wake-up voice interaction and hand-free call, improving driving convenience and safety.

CN120602908APending Publication Date: 2025-09-05CHINA AUTOMOTIVE INNOVATION CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510625114.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

When driving a motorcycle or electric vehicle, the driver needs to operate his mobile phone when making or receiving calls, which leads to poor experience and safety risks, and is cumbersome to control the vehicle.

Method used

By installing a dual-mode Bluetooth module in the smart helmet and the target vehicle, the HFP protocol of classic Bluetooth is used to establish a voice channel for wake-up-free voice interaction, and switch to the voice channel of the mobile terminal when needed, realizing the transmission of voice signals and call audio.

Benefits of technology

It enables drivers to make phone calls and control the vehicle without manual operation while driving, improving the convenience and safety of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602908A_ABST
    Figure CN120602908A_ABST
Patent Text Reader

Abstract

The invention discloses a voice processing method and device and a storage medium, and the method comprises the steps: an intelligent helmet employs a voice communication protocol of a dual-mode Bluetooth module to build a first voice channel between the intelligent helmet and a target vehicle, and transmits a captured voice signal to the target vehicle through the first voice channel; the target vehicle carries out instruction type analysis on the voice signal to obtain an instruction type result, and if the instruction type result represents that an operation corresponding to the voice signal is an operation of controlling the mobile terminal to execute a telephone service, channel switching indication information is sent to the intelligent helmet; the intelligent helmet disconnects the first voice channel according to the channel switching indication information and establishes a second voice channel with the mobile terminal; according to the invention, the voice transmission among the target vehicle, the intelligent helmet and the mobile terminal is realized, the voice interaction convenience is effectively improved, and the voice interaction method and the intelligent helmet have the advantages that the voice transmission among the target vehicle, the intelligent helmet and the mobile terminal is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a speech processing method, device, and storage medium. Background Art

[0002] Drivers are required to wear helmets when operating vehicles such as motorcycles and electric vehicles. The primary purpose of a helmet is to protect the rider's head from impact, preventing or mitigating injuries. However, if the driver needs to make or receive a call, they often need to operate their phone. This requires the driver to stop the vehicle to operate the phone, which is a poor user experience. Furthermore, it poses significant safety risks and impacts road safety. Furthermore, when the driver needs to control the vehicle, manual operation is often required, which is cumbersome.

[0003] Therefore, how to achieve voice transmission between helmets, vehicles and mobile terminals and improve the convenience of voice interaction is a technical problem that needs to be solved urgently. Summary of the Invention

[0004] The present application provides a voice processing method, device and storage medium, which can realize voice transmission among a target vehicle, a smart helmet and a mobile terminal, effectively improving the convenience of voice interaction.

[0005] On the one hand, an embodiment of the present application provides a voice processing method applied to a smart helmet equipped with a dual-mode Bluetooth module, the method comprising:

[0006] Establishing a first voice channel between the smart helmet and the target vehicle using the voice communication protocol of the dual-mode Bluetooth module;

[0007] Acquiring a voice signal captured by the sound pickup module of the smart helmet, and sending the voice signal to the target vehicle through the first voice channel so that the target vehicle performs command type analysis on the voice signal to obtain a command type result;

[0008] receiving channel switching instruction information sent by the target vehicle when the instruction type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform a telephone service operation;

[0009] disconnecting the first voice channel and establishing a second voice channel between the smart helmet and the mobile terminal according to the channel switching instruction information;

[0010] Acquire the call audio data captured by the sound pickup module, and transmit the call audio data to the mobile terminal through the second voice channel.

[0011] In one aspect, an embodiment of the present application provides a voice processing method, which is applied to a target vehicle equipped with a dual-mode Bluetooth module. The method includes:

[0012] Establishing a first voice channel between the target vehicle and the smart helmet using the voice communication protocol of the dual-mode Bluetooth module;

[0013] If a voice signal sent by the smart helmet is received in the first voice channel, the voice signal is parsed for a command type to obtain a command type result; the voice signal is captured by a sound pickup module of the smart helmet;

[0014] If the instruction type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform telephone service operations, a channel switching indication message is sent to the smart helmet so that the smart helmet disconnects the first voice channel and establishes a second voice channel with the mobile terminal according to the channel switching indication message, and obtains the call audio data captured by the sound pickup module, and transmits the call audio data to the mobile terminal through the second voice channel.

[0015] On the other hand, the present application also provides a voice processing device for use in a smart helmet equipped with a dual-mode Bluetooth module, the device comprising:

[0016] A first voice channel establishing module on the helmet side is used to establish a first voice channel between the smart helmet and the target vehicle using the voice communication protocol of the dual-mode Bluetooth module;

[0017] a voice signal sending module, configured to obtain the voice signal captured by the sound pickup module of the smart helmet, and send the voice signal to the target vehicle through the first voice channel so that the target vehicle performs command type analysis on the voice signal to obtain a command type result;

[0018] a channel switching instruction information receiving module, configured to receive channel switching instruction information sent by the target vehicle when the instruction type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform a telephone service operation;

[0019] A second voice channel establishing module, configured to disconnect the first voice channel and establish a second voice channel between the smart helmet and the mobile terminal according to the channel switching indication information;

[0020] The call audio transmission module is used to obtain the call audio data captured by the sound pickup module and transmit the call audio data to the mobile terminal through the second voice channel.

[0021] On the other hand, the present application also provides a voice processing device, which is applied to a target vehicle equipped with a dual-mode Bluetooth module, and the device includes:

[0022] A vehicle-side first voice channel establishing module, configured to establish a first voice channel between the target vehicle and the smart helmet using the voice communication protocol of the dual-mode Bluetooth module;

[0023] a command type parsing module, configured to, upon receiving a voice signal sent by the smart helmet in the first voice channel, parse the voice signal for a command type to obtain a command type result; the voice signal is captured by the sound pickup module of the smart helmet;

[0024] A channel switching indication information sending module is used to send channel switching indication information to the smart helmet if the instruction type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform a telephone service operation, so that the smart helmet disconnects the first voice channel and establishes a second voice channel with the mobile terminal according to the channel switching indication information, and obtains the call audio data captured by the sound pickup module, and transmits the call audio data to the mobile terminal through the second voice channel.

[0025] On the other hand, the present application also provides an electronic device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the speech processing method as described above.

[0026] On the other hand, the present application also provides a computer storage medium, which stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the speech processing method as described above.

[0027] In another aspect, the present application further provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to implement the above-described speech processing method.

[0028] The speech processing method provided in this application has the following technical effects:

[0029] In an embodiment of the present application, the smart helmet uses the voice communication protocol of the dual-mode Bluetooth module to establish a first voice channel between the smart helmet and the target vehicle; in the traditional entertainment system, the car-side sound pickup module is used to provide wake-up-free voice services, and no additional communication connection is required for data transmission, and the established classic Bluetooth connection is generally only used for telephone services. In this application, multiple connections can be maintained at the same time by using the dual-mode Bluetooth module in both directions, and the voice communication protocol connection of classic Bluetooth is used for real-time voice interaction and switching of telephone services. The smart helmet establishes a voice channel in the absence of telephone services, thereby sending the captured voice signal to the target vehicle through the first voice channel; by continuously monitoring the first voice channel, wake-up-free voice interaction is achieved. Specifically, if a voice signal sent by a smart helmet is received in the first voice channel, the target vehicle performs a command type analysis on the voice signal to obtain a command type result. If the command type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform a telephone service operation, the target vehicle sends a channel switching instruction message to the smart helmet; the smart helmet disconnects the first voice channel according to the channel switching instruction message and establishes a second voice channel between the smart helmet and the mobile terminal; the smart helmet obtains the call audio data captured by the sound pickup module, and transmits the call audio data to the mobile terminal through the second voice channel. This application uses the voice channel of classic Bluetooth to complete the multiplexing of telephone voice and voice assistant interaction, realize voice transmission among the target vehicle, smart helmet and mobile terminal, and effectively improve the convenience of voice interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0031] Figure 1 This is a diagram of the application environment of the speech processing method provided in the embodiment of the present application;

[0032] Figure 2 This is a flow chart of the speech processing method provided in the embodiment of the present application;

[0033] Figure 3 1 is a flow chart of a method for parsing a voice signal into command types according to an embodiment of the present application;

[0034] Figure 4 1 is a flow chart of a method for transmitting media audio provided in an embodiment of the present application;

[0035] Figure 5Schematic diagram of the data transmission process of the voice processing method provided in the embodiment of the present application;

[0036] Figure 6 This is a schematic diagram of the processing flow of the call service provided in the embodiment of the present application;

[0037] Figure 7 This is a schematic diagram of the voice processing architecture of the smart helmet provided in an embodiment of the present application;

[0038] Figure 8 Schematic diagram of the structure of the speech processing device provided in an embodiment of the present application;

[0039] Figure 9 Schematic diagram of the structure of the speech processing device provided in an embodiment of the present application;

[0040] Figure 10 This is a hardware structure block diagram of a server of a voice processing method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0042] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.

[0043] It is understandable that in the specific implementation of this application, data related to voice signals is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0044] Figure 1 This is an application environment diagram of the speech processing method provided in an embodiment of the present application.

[0045] like Figure 1 As shown, the application environment may include at least a target vehicle 01, a smart helmet 02 and a mobile terminal 03.

[0046] In an optional embodiment, the target vehicle 01 can be a vehicle such as a motorcycle or electric car that is used in scenarios where a helmet needs to be worn. In a traditional car cockpit, the intelligent voice interaction function generally uses the car's built-in microphone to capture the user's voice, which is directly processed by the car entertainment system or interacts with the cloud to process commands to complete the voice dialogue interaction function. The voice processing method in this application is applied to the target vehicle 01. Taking a motorcycle as an example, in addition to the microphone and power amplifier equipment on the motorcycle body itself, the driving helmet will serve as another audio input and output device with a higher frequency of use. Therefore, the voice processing method of this application introduces a smart helmet 02 as a device to capture user voice signals, and realizes voice dialogue interaction functions including voice control of the vehicle and hands-free mobile terminal phone calls during vehicle driving through the smart helmet 02.

[0047] In an optional embodiment, the mobile terminal 03 can be used to implement call services. Specifically, the mobile terminal 03 may include but is not limited to smart phones, desktop computers, tablet computers, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices and other types of electronic devices that implement call services; it may also be software running on the above electronic devices, such as applications, applets, etc. The operating system running on the electronic device in the embodiment of the present application may include but is not limited to Android, IOS, Linux, Windows, etc. It should be noted that the target vehicle 01, the smart helmet 02 and the mobile terminal 03 can achieve communication connection through Bluetooth technology.

[0048] In an optional embodiment, the application environment in this application may further include a server 04. Server 04 may include a standalone server, a distributed server, or a server cluster consisting of multiple servers. Server 04 may include a network communication unit, a processor, and a memory. Specifically, server 04 may be used to provide background services for target vehicle 01, such as voice recognition services.

[0049] In the prior art, when a motorcycle, electric vehicle or other vehicle is in motion, if the driver needs to make or answer a call, he or she often needs to operate a mobile phone to do so. On the one hand, the driver needs to stop the vehicle to operate the mobile phone, which is a poor experience. On the other hand, it also brings greater safety risks and affects road safety. In addition, when the driver needs to control the vehicle itself, he or she often needs to operate the vehicle manually. For example, when the vehicle itself has an integrated navigation system, if the driver needs to use the navigation function, he or she needs to manually enter the destination and select the navigation route, etc., which is a relatively cumbersome operation. Therefore, the present application proposes a voice processing method that can realize voice transmission among the target vehicle, smart helmet and mobile terminal, and realize user wake-up-free voice interaction and hands-free calling functions.

[0050] Figure 2 : It is a flow chart of the voice processing method provided in the embodiment of the present application. Among them, the smart helmet and the target vehicle are both equipped with dual-mode Bluetooth modules. The dual-mode Bluetooth module can support classic Bluetooth and low-power Bluetooth (Bluetooth Low Energy, BLE) at the same time. In the embodiment of the present application, the hands-free protocol (HFP) of classic Bluetooth is used to realize wake-up-free voice interaction and call services, and the Bluetooth low-power audio protocol (LowEnergy Audio) of low-power Bluetooth is used to realize the transmission of media audio data. HFP is the core protocol for realizing hands-free calls in Bluetooth technology. It mainly supports the establishment of voice and control channels between Bluetooth devices (such as headphones, car systems) and mobile phones, enabling the device to act as a remote speaker and microphone for the mobile phone, and provide call control functions (such as answering / hanging up, volume adjustment, etc.). LE Audio is a new technical standard launched by the Bluetooth Special Interest Group (SIG), which aims to provide audio transmission functions through low-power Bluetooth BLE. It not only improves the audio quality, but also saves more energy, which is conducive to improving the user's voice experience. In the prior art, voice interaction data transmission can be performed based on traditional Wi-Fi communication technology, but this method consumes large power and has poor portability. This application proposes a voice processing method for voice interaction data transmission based on the short-range real-time transmission capability of Bluetooth technology.

[0051] Specifically, the method provided in this application includes:

[0052] S201: The smart helmet establishes a first voice channel between the smart helmet and the target vehicle using the voice communication protocol of the dual-mode Bluetooth module;

[0053] The dual-mode Bluetooth module's voice communication protocol can be the HFP protocol of classic Bluetooth. The HFP connection of classic Bluetooth is used for real-time voice interaction and telephone service switching. In one embodiment, Bluetooth pairing can first be performed between the smart helmet and the target vehicle. After successful Bluetooth pairing, a first voice channel is established between the smart helmet and the target vehicle.

[0054] S203: The smart helmet obtains the voice signal captured by the sound pickup module of the smart helmet and sends the voice signal to the target vehicle through the first voice channel;

[0055] It should be noted that the smart helmet has a built-in sound pickup module, such as a microphone, that can capture the user's voice signals. The voice signals can be control commands issued by the user, such as "Call contact A," "Play music," "Navigate to mall B," etc. This application does not limit the specific content of the voice signals.

[0056] When the smart helmet captures the voice signal emitted by the user, it can be transmitted to the Bluetooth receiver of the target vehicle through the pre-established first voice channel, an air interface, so that the target vehicle can perform subsequent processing.

[0057] S205: If the voice signal sent by the smart helmet is received in the first voice channel, the target vehicle performs command type analysis on the voice signal to obtain a command type result;

[0058] Command type parsing, that is, parsing the content of the voice signal to obtain the user's corresponding operation for the voice signal. In the embodiment of the present application, the command type results can be divided into vehicle control type commands and telephone service operation type commands. Vehicle control type commands are commands for controlling the target vehicle, such as "navigate to mall B"; telephone service operation type commands are commands for controlling calls on the mobile terminal, such as "call contact A", "answer the call", etc.

[0059] S207: If the instruction type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform a telephone service operation, the target vehicle sends a channel switching instruction message to the smart helmet;

[0060] S209: The smart helmet disconnects the first voice channel according to the channel switching instruction information and establishes a second voice channel between the smart helmet and the mobile terminal;

[0061] S211: The smart helmet obtains the call audio data captured by the sound pickup module, and transmits the call audio data to the mobile terminal through the second voice channel.

[0062] Since a first voice channel is established between the target vehicle and the smart helmet at this time, the first voice channel can realize the transmission of voice signals. If the command type of the voice signal is a vehicle control command, the target vehicle performs the corresponding control operation; if the command type of the voice signal is a telephone service operation command, the HFP connection is switched for the phone, and the target vehicle sends a channel switching instruction information to the smart helmet, so that the smart helmet establishes an HFP connection with the mobile terminal. The user's call audio data is encoded by the Bluetooth module and transmitted back to the mobile phone for two-way data transmission of the phone.

[0063] In an embodiment of the present application, the smart helmet uses the voice communication protocol of the dual-mode Bluetooth module to establish a first voice channel between the smart helmet and the target vehicle; while in traditional entertainment systems, the car-side sound pickup module is used for wake-up-free voice services, and no additional communication connection is required for data transmission, and the established classic Bluetooth HFP connection is generally only used for telephone services. In this application, multiple connections can be maintained simultaneously by using the dual-mode Bluetooth module in both directions, and the HFP connection of classic Bluetooth is used for real-time voice interaction and switching of telephone services. The smart helmet establishes an HFP transmission path, i.e., the first voice channel, in the absence of telephone services, so that the captured voice signal is sent to the target vehicle through the first voice channel; by continuously monitoring the first voice channel, HFP wake-up-free voice interaction is achieved. Specifically, if a voice signal sent by a smart helmet is received in the first voice channel, the target vehicle performs a command type analysis on the voice signal to obtain a command type result. If the command type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform a telephone service operation, the target vehicle sends a channel switching instruction message to the smart helmet; the smart helmet disconnects the first voice channel according to the channel switching instruction message and establishes a second voice channel between the smart helmet and the mobile terminal; the smart helmet obtains the call audio data captured by the sound pickup module, and transmits the call audio data to the mobile terminal through the second voice channel. This application uses the HFP channel of classic Bluetooth to complete the multiplexing of telephone voice and voice assistant interaction, realize voice transmission among the target vehicle, smart helmet and mobile terminal, and effectively improve the convenience of voice interaction.

[0064] In one embodiment, after the smart helmet transmits the call audio data to the mobile terminal through the second voice channel, the method of the present application further includes:

[0065] If the smart helmet detects the call end signal sent by the mobile terminal, it disconnects the second voice channel and sends a reconnection request to the target vehicle; the target vehicle responds to the reconnection request by generating a reconnection confirmation message and sending the reconnection confirmation message to the smart helmet; since the second voice channel is also established based on the HFP of Classic Bluetooth, according to the HFP protocol, after the call ends, the mobile terminal will send a call end signal to the smart helmet, and then the smart helmet will disconnect the second voice channel between it and the mobile terminal and attempt to reconnect with the target vehicle to continue detecting the user's voice signal;

[0066] The smart helmet receives the reconnection confirmation information sent by the target vehicle, and re-establishes the first voice channel between the smart helmet and the target vehicle, thereby continuously monitoring whether the pickup module captures the user's voice command again, thereby realizing wake-up-free voice control.

[0067] In an embodiment of the present application, since the first voice channel remains connected, the user's voice signal can be continuously detected without the user having to say the wake-up word and then capturing the user's voice command, thereby realizing a wake-up-free voice interaction function and improving the convenience of user voice interaction.

[0068] Figure 3 FIG. 1 is a flow chart of a method for parsing a voice signal into a command type according to an embodiment of the present application. Figure 3 As shown, the method includes:

[0069] S301: The smart helmet obtains a voice signal captured by a sound pickup module and performs noise reduction processing on the voice signal to obtain a noise-reduced voice signal;

[0070] The sound pickup module can be a built-in sound pickup microphone in the smart helmet. In one implementation, the speech signal can be denoised using traditional methods, such as spectral subtraction, Wiener filtering, wavelet transform, etc., which will not be described in detail here. In another implementation, noise reduction can also be performed using deep learning methods. Specifically, a paired dataset of noisy speech samples and clean speech samples is constructed; a preset noise reduction model is constructed, where the noise reduction model may include a generator and a discriminator. The generator is used to process the noisy speech samples to generate noise-reduced speech. The discriminator is used to distinguish between the generated noise-reduced speech (labeled 0) and the clean speech samples (labeled 1), and outputs a probability value ranging from 0 to 1, indicating the discriminator's confidence that the input data is a true clean speech sample. The discriminator loss is calculated based on the difference between the probability value and the label, and the discriminator parameters are back-propagated to improve the discrimination ability. Then, the discriminator parameters are fixed, and the noise-reduced speech is generated using the generator and input into the discriminator. The generator loss is calculated, and the generator parameters are updated based on the generator loss through backpropagation to generate more realistic data. It should be noted that the above-mentioned discriminator and generator are trained in a repeated alternating process until the generator and discriminator reach a dynamic equilibrium. The generator at this equilibrium is then determined as the target noise reduction model. In application, the speech signal is input into the target noise reduction model for signal generation processing to obtain the denoised speech signal.

[0071] S303: The smart helmet encodes the noise-reduced voice signal according to a preset data encoding strategy to obtain a voice transmission signal; the preset data encoding strategy is determined according to the voice communication protocol;

[0072] The voice communication protocol may be the HFP protocol. The preset data encoding strategy is the data encoding strategy corresponding to the HFP protocol. The encoding process is a prior art and will not be described in detail here.

[0073] S305 The smart helmet sends the voice transmission signal to the target vehicle through the first voice channel;

[0074] S307: The target vehicle obtains the voice transmission signal received in the first voice channel, and decodes the voice transmission signal to obtain the voice signal;

[0075] It should be noted that the target vehicle is decoded according to a preset data decoding strategy, which is a data decoding strategy corresponding to the HFP protocol. The decoding process is a prior art and will not be described in detail here.

[0076] S309: The target vehicle performs speech recognition processing on the speech signal to obtain text information;

[0077] In one implementation, speech recognition processing is performed using a speech recognition model. The speech signal is input into a pre-trained speech recognition model to obtain text information corresponding to the speech signal. It should be noted that during speech recognition processing, voiceprint detection can also be performed on the speech signal to obtain voiceprint features. These features are then matched with preset voiceprint features in a voiceprint database. If the match is successful, the speech signal is input into the pre-trained speech recognition model. If the match fails, voiceprint feedback information is generated and sent to the smart helmet. The voiceprint database stores the voiceprint features of the authorized driver of the target vehicle. Voiceprint detection can improve vehicle control safety.

[0078] S311: The target vehicle performs semantic recognition processing on the text information to obtain operation intention information and operation parameters;

[0079] It should be noted that the present application does not limit the specific process of semantic recognition processing. For example, semantic recognition can be performed using a pre-trained semantic recognition model.

[0080] The action intention information is the action word in the text message, and the action parameter is the execution parameter corresponding to the action word. In one example, if the text message corresponding to the voice signal is "Play song C", the action intention information can be to open the music player to play music, and the action parameter can be song C. For another example, if the text message corresponding to the voice signal is "Call contact A", the action intention information can be to open the mobile terminal's phone app to make a call, and the action parameter can be contact A's phone number.

[0081] S313: The target vehicle determines the instruction type result according to the operation intention information and the operation parameters.

[0082] In an embodiment of the present application, the smart helmet obtains a voice transmission signal after noise reduction and encoding, and transmits the voice transmission signal to the target vehicle. The target vehicle performs voice recognition and semantic recognition processing to accurately obtain the instruction type result, so that the target vehicle can perform different subsequent processing according to the accurate instruction type result, thereby realizing the convenience of voice interaction between the smart helmet, the target vehicle and the mobile terminal.

[0083] It should be noted that if the command type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform a telephone service operation, the target vehicle sends a channel switching instruction to the smart helmet. For details, see step S207. In another embodiment, if the command type result indicates that the operation corresponding to the voice signal is a control operation other than controlling the mobile terminal to perform a telephone service operation, the target vehicle performs the control operation corresponding to the voice signal; if the target vehicle detects that the control operation has been completed, it sends command feedback information to the smart helmet; the command feedback information indicates whether the control operation was successfully executed.

[0084] In an embodiment of the present application, after the target vehicle detects that the control operation is completed, command feedback information is sent to the smart helmet, so that the user can clearly know the execution result of the voice command, thereby improving the user's voice interaction experience.

[0085] In one embodiment, the target vehicle establishes a communication connection with the mobile terminal. The method provided by the present application further includes: while the first voice channel is in a connected state, the target vehicle receives the incoming call indication information sent by the mobile terminal; then, the target vehicle sends the incoming call indication information to the smart helmet through the first voice channel; the incoming call indication information is output through the earphone module of the smart helmet, so that the smart helmet can continuously monitor the voice signal sent by the user. For example, in this scenario, the user sends a voice "confirm answer", and the voice signal is output through the above Figure 3 The illustrated embodiment is sent to the target vehicle, causing the target vehicle to perform command type analysis and obtain a command type result. Since the command type result in this scenario indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform a telephone service operation, the target vehicle sends a channel switching instruction to the smart helmet, causing the smart helmet to disconnect the first voice channel and establish a second voice channel with the mobile terminal based on the channel switching instruction, enabling hands-free calling. The specific steps of this process are described above in steps S207 to S211 and will not be repeated here.

[0086] In an embodiment of the present application, while the first voice channel remains connected, if an incoming call event is detected, the vehicle-mounted terminal sends the incoming call indication information to the smart helmet. After the user confirms the answer, the smart helmet disconnects the first voice channel with the target vehicle and establishes a second voice channel with the mobile terminal, realizing a hands-free call service, which effectively improves the convenience and safety of the call.

[0087] In one embodiment, since the target vehicle and the smart helmet in the embodiment of the present application are both equipped with dual-mode Bluetooth modules, the method of the present application can also transmit media audio. Figure 4: is a flow chart of a method for transmitting media audio provided by an embodiment of the present application. Specifically, the method may include:

[0088] S401: The target vehicle establishes a media data channel between the smart helmet and the target vehicle using the media communication protocol of the dual-mode Bluetooth module; the media data channel is independent of the first voice channel;

[0089] The media data channel can utilize the data transmission channel established by the LE Audio protocol of BLE.

[0090] S403: If the instruction type result is to play media audio, the target vehicle sends the media audio data to the smart helmet through the media data channel;

[0091] For example, when the voice signal is "Play song C", the target vehicle parses the voice signal and finds that the instruction type corresponding to the voice signal is to play media audio, and specifically song C. The target vehicle then sends the media audio data corresponding to song C to the smart helmet through the media data channel.

[0092] S405: The smart helmet outputs the media audio data through the earphone module.

[0093] In the embodiment of the present application, voice interaction is maintained through the HFP of the classic Bluetooth of the dual-mode Bluetooth module, while normal media audio transmission and playback or other audio services are maintained through the LE Audio of BLE, thereby improving driving safety on the one hand and achieving the convenience of voice interaction on the other hand, thereby improving the voice interaction experience.

[0094] It should be noted that in order to enable the user to clearly hear the caller's voice during a call and improve the user's call experience, in this embodiment of the application, while the second voice channel remains connected, the smart helmet can pause receiving media audio data sent by the target vehicle and send a media audio interruption indication to the target vehicle. After the second voice channel is disconnected, the smart helmet sends a media audio playback indication to the target vehicle and resumes receiving media audio data sent by the target vehicle.

[0095] In one embodiment, the smart helmet can also intelligently adjust the playback volume of media audio and call audio. If the media data channel and the second voice channel exist at the same time, the playback volume of the media audio data is reduced and the playback volume of the call audio data is increased, thereby improving the user's call experience.

[0096] Figure 5 It is a data transmission flow diagram of the voice processing method provided in an embodiment of the present application.

[0097] Both the smart helmet and the target vehicle are equipped with dual-mode Bluetooth modules. The microphone of the smart helmet captures the voice commands issued by the user, and after processing through the noise reduction and encoding modules, the command data is sent to the vehicle side using the classic Bluetooth HFP channel of the dual-mode Bluetooth module. After the dual-mode Bluetooth module on the vehicle side receives the voice command, it parses the voice command through the voice module and performs subsequent processing based on the parsing result of the command type. It should be noted that the vehicle side can send feedback data of the command execution to the smart helmet through the HFP channel. In addition, the vehicle side and the smart helmet can also establish an LE Audio channel. The vehicle side sends media audio data to the smart helmet, and the smart helmet outputs the media audio data to the headphones. Through the above-mentioned dual-mode Bluetooth module, it can support audio services of classic Bluetooth and BLE Bluetooth at the same time, providing a higher quality and lower latency audio experience.

[0098] Figure 6 It is a schematic diagram of the processing flow of the call service provided in the embodiment of the present application.

[0099] like Figure 6 As shown, the traditional car entertainment system uses the car microphone for wake-up-free voice service, does not require an additional communication connection for data transmission, and the established classic Bluetooth HFP connection is generally only used for telephone services. In this application, the smart helmet and the target vehicle use dual-mode Bluetooth modules in both directions, can maintain multiple connections at the same time, and use the classic Bluetooth HFP connection for real-time voice interaction and switching of telephone services. Specifically, first, when there is no telephone service, the smart helmet establishes an HFP transmission path after the Bluetooth pairing connection, realizing two-way transmission of voice signals between the smart helmet and the target vehicle; at the same time, if there is an incoming call or an outgoing call event, the HFP connection is switched for the phone, and the call audio data is transmitted in both directions of the phone in the same way as the traditional Bluetooth phone processing. Through the above method, HFP is used to achieve wake-up-free, and hands-free phone calls are also achieved, which improves the convenience of voice interaction and improves driving safety.

[0100] Figure 7 Schematic diagram of the voice processing architecture of the smart helmet provided in an embodiment of the present application.

[0101] like Figure 7As shown, when the smart helmet's microphone pickup module receives audio data, i.e., a voice signal, the first voice channel (HFP for Classic Bluetooth) remains connected. Therefore, the target vehicle's Bluetooth protocol stack can determine whether the audio data is a telephone service operation. If it is a voice interaction command, the target vehicle's voice module performs voice recognition, allowing the vehicle-side to complete the command control operation. If it is a telephone service, the HFP connection is switched for the call. The target vehicle sends a channel switching instruction to the smart helmet, establishing an HFP connection between the smart helmet and the mobile terminal. The user's call audio data is encoded by the Bluetooth module and transmitted back to the mobile phone, enabling two-way data transmission of the call.

[0102] Figure 8 It is a structural diagram of the speech processing device provided in an embodiment of the present application.

[0103] like Figure 8 As shown, the device 800 can be applied to a smart helmet equipped with a dual-mode Bluetooth module. The device 800 may include:

[0104] The first voice channel establishing module 801 on the helmet side is used to establish a first voice channel between the smart helmet and the target vehicle using the voice communication protocol of the dual-mode Bluetooth module;

[0105] The voice signal sending module 802 is used to obtain the voice signal captured by the sound pickup module of the smart helmet, and send the voice signal to the target vehicle through the first voice channel so that the target vehicle performs command type analysis on the voice signal to obtain a command type result;

[0106] The channel switching instruction information receiving module 803 is configured to receive the channel switching instruction information sent by the target vehicle when the instruction type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform a telephone service operation;

[0107] A second voice channel establishing module 804 is configured to disconnect the first voice channel and establish a second voice channel between the smart helmet and the mobile terminal according to the channel switching instruction information;

[0108] The call audio transmission module 805 is used to obtain the call audio data captured by the sound pickup module and transmit the call audio data to the mobile terminal through the second voice channel.

[0109] In some embodiments, the apparatus 800 may further include:

[0110] A helmet media data channel establishment module is used to establish a media data channel between the smart helmet and the target vehicle using the media communication protocol of the dual-mode Bluetooth module; the media data channel is independent of the first voice channel;

[0111] The media audio data output module is used to receive the media audio data sent by the target vehicle through the media data channel and output the media audio data through the earphone module of the smart helmet.

[0112] In some embodiments, the voice signal sending module may include:

[0113] A noise reduction submodule is used to obtain a voice signal captured by the sound pickup module of the smart helmet and perform noise reduction processing on the voice signal to obtain a noise-reduced voice signal;

[0114] an encoding submodule, configured to encode the noise-reduced voice signal according to a preset data encoding strategy to obtain a voice transmission signal; the preset data encoding strategy is determined according to the voice communication protocol;

[0115] The voice transmission signal sending submodule is used to send the voice transmission signal to the target vehicle through the first voice channel.

[0116] In some embodiments, the apparatus 800 may further include:

[0117] a reconnection request sending module, configured to disconnect the second voice channel and send a reconnection request to the target vehicle upon detecting a call end signal sent by the mobile terminal;

[0118] The reconnection confirmation information receiving module is used to receive the reconnection confirmation information sent by the target vehicle and re-establish the first voice channel between the smart helmet and the target vehicle.

[0119] Figure 9 It is a structural diagram of the speech processing device provided in an embodiment of the present application.

[0120] like Figure 9 As shown, the device 900 can be used in a target vehicle equipped with a dual-mode Bluetooth module. The device 900 may include:

[0121] The vehicle-side first voice channel establishing module 901 is used to establish a first voice channel between the target vehicle and the smart helmet using the voice communication protocol of the dual-mode Bluetooth module;

[0122] The instruction type parsing module 902 is configured to, upon receiving a voice signal sent by the smart helmet in the first voice channel, parse the voice signal for the instruction type to obtain an instruction type result; the voice signal is captured by the sound pickup module of the smart helmet;

[0123] The channel switching indication information sending module 903 is used to send channel switching indication information to the smart helmet if the instruction type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform telephone service operations, so that the smart helmet disconnects the first voice channel according to the channel switching indication information and establishes a second voice channel with the mobile terminal, and obtains the call audio data captured by the sound pickup module, and transmits the call audio data to the mobile terminal through the second voice channel.

[0124] In some embodiments, the instruction type parsing module may include:

[0125] a decoding submodule, configured to obtain a voice transmission signal received in the first voice channel, and decode the voice transmission signal to obtain the voice signal;

[0126] A speech recognition submodule, configured to perform speech recognition processing on the speech signal to obtain text information;

[0127] A semantic recognition submodule, configured to perform semantic recognition processing on the text information to obtain operation intention information and operation parameters;

[0128] The instruction type result determination submodule is used to determine the instruction type result according to the operation intention information and the operation parameters.

[0129] In some embodiments, the apparatus 900 may further include:

[0130] a control operation execution module, configured to execute the control operation corresponding to the voice signal if the instruction type result indicates that the operation corresponding to the voice signal is a control operation other than controlling the mobile terminal to perform a telephone service operation;

[0131] The feedback information sending module is used to send instruction feedback information to the smart helmet if it is detected that the control operation is completed; the instruction feedback information indicates whether the control operation is successfully executed.

[0132] In some embodiments, the apparatus 900 may further include:

[0133] A vehicle-side media data channel establishment module is used to establish a media data channel between the smart helmet and the target vehicle using the media communication protocol of the dual-mode Bluetooth module; the media data channel is independent of the first voice channel;

[0134] The media audio data sending module is used to send the media audio data to the smart helmet through the media data channel if the instruction type result is to play media audio, so that the headphone module of the smart helmet outputs the media audio.

[0135] In some embodiments, the target vehicle establishes a communication connection with the mobile terminal, and the apparatus 900 further includes:

[0136] an incoming call indication information receiving module, configured to receive incoming call indication information sent by the mobile terminal while the first voice channel is in a connected state;

[0137] The incoming call indication information sending module is used to send the incoming call indication information to the smart helmet through the first voice channel, so that the incoming call indication information is output through the earphone module of the smart helmet.

[0138] The device and method embodiments in the device embodiments are based on the same inventive concept.

[0139] An embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the method provided in the above method embodiment.

[0140] An embodiment of the present application also provides a computer storage medium, which can be set in a terminal to store at least one instruction or at least one program related to a method provided in the above method embodiment for implementing a method embodiment, and the at least one instruction or at least one program is loaded and executed by the processor to implement the method provided in the above method embodiment.

[0141] The embodiments of the present application further provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method provided in the above method embodiment.

[0142] Optionally, in an embodiment of the present application, the storage medium may be located in at least one of a plurality of network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0143] The memory described in the embodiment of the present application can be used to store software programs and modules, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory may also include a memory controller to provide the processor with access to the memory.

[0144] The method provided in the embodiment of the present application can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Taking running on a server as an example, Figure 10 This is a hardware structure diagram of a server of a voice processing method provided in an embodiment of the present application. Figure 10As shown, the server 1000 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 1010 (the central processing unit 1010 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1030 for storing data, and one or more storage media 1020 (such as one or more mass storage devices) for storing application programs 1023 or data 1022. Among them, the memory 1030 and the storage medium 1020 can be temporary storage or permanent storage. The program stored in the storage medium 1020 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1010 can be configured to communicate with the storage medium 1020 to execute a series of instruction operations in the storage medium 1020 on the server 1000. The server 1000 may also include one or more power supplies 1060, one or more wired or wireless network interfaces 1050, one or more input and output interfaces 1040, and / or one or more operating systems 1021, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0145] The input / output interface 1040 can be used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communication provider of the server 1000. In one embodiment, the input / output interface 1040 includes a network interface controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one embodiment, the input / output interface 1040 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0146] It can be understood by those skilled in the art that Figure 10 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 10 More or fewer components than shown, or with Figure 10 Different configurations shown.

[0147] It can be seen from the embodiments of the voice processing method, device and storage medium provided by the present application that in the embodiments of the present application, the smart helmet uses the voice communication protocol of the dual-mode Bluetooth module to establish a first voice channel between the smart helmet and the target vehicle; while in traditional entertainment systems, the car-side sound pickup module is used for wake-up-free voice services, and no additional communication connection is required for data transmission, and the established classic Bluetooth HFP connection is generally only used for telephone services. In this application, multiple connections can be maintained at the same time by using the dual-mode Bluetooth module in both directions, and the HFP connection of classic Bluetooth is used for real-time voice interaction and switching of telephone services. The smart helmet establishes an HFP transmission path, i.e., the first voice channel, in the absence of a telephone service, so that the captured voice signal is sent to the target vehicle through the first voice channel; by continuously monitoring the first voice channel, HFP wake-up-free voice interaction is achieved. Specifically, if a voice signal sent by a smart helmet is received in the first voice channel, the target vehicle performs a command type analysis on the voice signal to obtain a command type result. If the command type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform a telephone service operation, the target vehicle sends a channel switching instruction message to the smart helmet; the smart helmet disconnects the first voice channel according to the channel switching instruction message and establishes a second voice channel between the smart helmet and the mobile terminal; the smart helmet obtains the call audio data captured by the sound pickup module, and transmits the call audio data to the mobile terminal through the second voice channel. This application uses the HFP channel of classic Bluetooth to complete the multiplexing of telephone voice and voice assistant interaction, realize voice transmission among the target vehicle, smart helmet and mobile terminal, and effectively improve the convenience of voice interaction.

[0148] It should be noted that the order of the embodiments of the present application described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0149] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, equipment, and storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For relevant portions, refer to the descriptions of the method embodiments.

[0150] Those skilled in the art will understand that all or part of the steps of implementing the above embodiments may be accomplished by hardware, or by a program instructing the relevant hardware to accomplish the steps. The program may be stored in a computer storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.

[0151] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A speech processing method, characterized in that: Applied to a smart helmet equipped with a dual-mode Bluetooth module, the method includes: Establishing a first voice channel between the smart helmet and the target vehicle using the voice communication protocol of the dual-mode Bluetooth module; Acquiring a voice signal captured by the sound pickup module of the smart helmet, and sending the voice signal to the target vehicle through the first voice channel so that the target vehicle performs command type analysis on the voice signal to obtain a command type result; receiving channel switching instruction information sent by the target vehicle when the instruction type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform a telephone service operation; disconnecting the first voice channel and establishing a second voice channel between the smart helmet and the mobile terminal according to the channel switching instruction information; Acquire the call audio data captured by the sound pickup module, and transmit the call audio data to the mobile terminal through the second voice channel.

2. The method according to claim 1, characterized in that The method further comprises: A media data channel is established between the smart helmet and the target vehicle using the media communication protocol of the dual-mode Bluetooth module; the media data channel is independent of the first voice channel; The media audio data is received from the target vehicle through the media data channel, and the media audio data is output through the earphone module of the smart helmet.

3. The method according to claim 1, characterized in that The step of obtaining a voice signal captured by a sound pickup module of the smart helmet and sending the voice signal to the target vehicle through the first voice channel includes: Acquiring a voice signal captured by a sound pickup module of the smart helmet, and performing noise reduction processing on the voice signal to obtain a noise-reduced voice signal; Encoding the noise-reduced voice signal according to a preset data encoding strategy to obtain a voice transmission signal; the preset data encoding strategy is determined according to the voice communication protocol; The voice transmission signal is sent to the target vehicle through the first voice channel.

4. The method according to claim 1, wherein After transmitting the call audio data to the mobile terminal through the second voice channel, the method further includes: If a call end signal sent by the mobile terminal is detected, disconnecting the second voice channel and sending a reconnection request to the target vehicle; Receive the reconnection confirmation information sent by the target vehicle, and re-establish the first voice channel between the smart helmet and the target vehicle.

5. A speech processing method, characterized in that: Applied to target vehicle; The target vehicle is equipped with a dual-mode Bluetooth module, and the method includes: Establishing a first voice channel between the target vehicle and the smart helmet using the voice communication protocol of the dual-mode Bluetooth module; If a voice signal sent by the smart helmet is received in the first voice channel, the voice signal is parsed for a command type to obtain a command type result; the voice signal is captured by a sound pickup module of the smart helmet; If the instruction type result indicates that the operation corresponding to the voice signal is to control the mobile terminal to perform telephone service operations, a channel switching indication message is sent to the smart helmet so that the smart helmet disconnects the first voice channel and establishes a second voice channel with the mobile terminal according to the channel switching indication message, and obtains the call audio data captured by the sound pickup module, and transmits the call audio data to the mobile terminal through the second voice channel.

6. The method according to claim 5, characterized in that If a voice signal sent by the smart helmet is received in the first voice channel, the voice signal is parsed for a command type to obtain a command type result, including: Acquire a voice transmission signal received in the first voice channel, and decode the voice transmission signal to obtain the voice signal; Performing speech recognition processing on the speech signal to obtain text information; Performing semantic recognition processing on the text information to obtain operation intention information and operation parameters; The instruction type result is determined according to the operation intention information and the operation parameters.

7. The method according to claim 5, characterized in that The method further comprises: If the instruction type result indicates that the operation corresponding to the voice signal is a control operation other than controlling the mobile terminal to perform a telephone service operation, executing the control operation corresponding to the voice signal; If it is detected that the control operation is completed, instruction feedback information is sent to the smart helmet; the instruction feedback information indicates whether the control operation is successfully executed.

8. The method according to claim 5, characterized in that The method further comprises: A media data channel is established between the smart helmet and the target vehicle using the media communication protocol of the dual-mode Bluetooth module; the media data channel is independent of the first voice channel; If the instruction type result is to play media audio, the media audio data is sent to the smart helmet through the media data channel so that the earphone module of the smart helmet outputs the media audio.

9. The method according to claim 5, characterized in that The target vehicle establishes a communication connection with the mobile terminal, and the method further includes: While the first voice channel is in a connected state, receiving incoming call indication information sent by the mobile terminal; The incoming call indication information is sent to the smart helmet through the first voice channel, so that the incoming call indication information is output through the earphone module of the smart helmet.

10. A computer storage medium, characterized in that The computer storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the speech processing method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Multi-stage wake-up method and system based on low-power-consumption wake-up protocol, medium and product

    CN121054003A